Event detection method and device based on multiple modes, electronic equipment and storage medium

By introducing multimodal data and dynamic weight fusion technology into graphical image recognition technology, the problem of insufficient event recognition accuracy in complex scenarios in traditional technology is solved, and higher recognition accuracy and robustness are achieved, adapting to different environmental conditions and preventing data interference.

CN119989258APending Publication Date: 2025-05-13SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD +2
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202411982715.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The accuracy of event recognition in traditional graphic image recognition technology still needs to be improved in complex scenarios, and it is difficult to maintain stable recognition performance under different environmental conditions, and it is susceptible to the risk of interference or forgery of single modal data.

Method used

By introducing multimodal data, combining the corresponding features of data of multiple modes, using their respective advantages to complement each other, using dynamic weight fusion technology to convert feature data of different modes into a unified feature space, and determining dynamic weights based on environmental conditions, weighted additions are performed to obtain fusion features, and event detection is finally performed through a pre-trained event detection model.

Benefits of technology

It significantly improves the accuracy and robustness of event recognition during public security and stability, maintains stable identification performance under different environmental conditions, effectively prevents the risk of single modal data being interfered or forged, and improves the recognition accuracy of event detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989258A_ABST
    Figure CN119989258A_ABST
Patent Text Reader

Abstract

The invention provides an event detection method based on multiple modes. The method comprises the steps of obtaining to-be-processed data of multiple modes of a target scene; feature extraction is carried out on the to-be-processed data according to modalities, feature data of multiple modalities are obtained, and the feature data of the same modality correspond to the to-be-processed data; fusing the feature data according to first dynamic weights to obtain a first fused feature, each modal corresponding to one first dynamic weight, and the first dynamic weights being determined according to the environment of the target scene; and performing event detection on the fusion features through a pre-trained event detection model to obtain an event detection result of the target scene. According to the invention, by introducing the multi-modal data and combining the features corresponding to the multi-modal data, stable identification performance can be maintained under different environmental conditions, the risk that single-modal data is interfered or counterfeited is effectively prevented, and the identification accuracy of event detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a multimodal event detection method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of society, public security and stability maintenance work is facing increasingly complex challenges. As one of the important means of public security and stability maintenance, the improvement of the accuracy of graphic image recognition technology is of great significance to improving the overall prevention and control capabilities. However, the accuracy of event recognition in complex scenarios by traditional graphic image recognition technology still needs to be improved. Therefore, this scheme proposes a design scheme based on multimodality to continuously improve the accuracy of graphic image recognition for public security and stability maintenance. Summary of the invention

[0003] The embodiment of the present invention provides a multimodal-based event detection method, which aims to improve the recognition accuracy of public security and stability maintenance events in complex scenarios. By introducing multimodal data, combining the features corresponding to data of multiple modalities, and utilizing their respective complementary advantages, the accuracy and robustness of event recognition in public security and stability maintenance can be significantly improved, while maintaining stable recognition performance under different environmental conditions, effectively preventing the risk of single modal data being interfered with or forged, and improving the recognition accuracy of event detection.

[0004] In a first aspect, an embodiment of the present invention provides a multimodal event detection method, the method comprising the following steps:

[0005] Obtaining multiple modal data to be processed of the target scene;

[0006] Extracting features of the data to be processed respectively according to the modality to obtain feature data of the multiple modalities, wherein the feature data of the same modality corresponds to the data to be processed;

[0007] The feature data are fused according to a first dynamic weight to obtain a first fused feature, each modality corresponds to a first dynamic weight, and the first dynamic weight is determined according to the environment of the target scene;

[0008] Event detection is performed on the fused features using a pre-trained event detection model to obtain an event detection result for the target scene.

[0009] Optionally, the extracting features of the data to be processed respectively according to the modality to obtain feature data of the multiple modalities includes:

[0010] Determine the feature extraction networks corresponding to different modalities. Different modalities correspond to different feature extraction networks.

[0011] Based on the feature extraction network corresponding to each modality, feature extraction is performed on the to-be-processed data of each modality to obtain feature data of the multiple modalities.

[0012] Optionally, fusing the feature data according to a first dynamic weight to obtain a first fused feature includes:

[0013] Convert the feature data of each modality into a unified feature space to obtain a plurality of unified feature data, each of which corresponds to the feature data of one modality;

[0014] The unified feature data are fused according to a first dynamic weight to obtain a first fused feature.

[0015] Optionally, fusing the unified feature data according to a dynamic weight to obtain a first fused feature includes:

[0016] Based on the environmental conditions of the target scene, determining the first dynamic weights in each mode, where different environmental conditions correspond to different first dynamic weights in each mode;

[0017] The unified feature data are weightedly added using the first dynamic weight to obtain a first fused feature.

[0018] Optionally, after performing event detection on the fused feature through the pre-trained event detection model to obtain the event detection result of the target scene, the method further includes:

[0019] Based on the event detection result, positive and negative samples are labeled for each modality to obtain a labeled sample set, where the labeled sample set includes an equal number of positive samples and negative samples;

[0020] Based on the negative samples in the labeled sample set, adjusting the dynamic weights corresponding to each modality to obtain a second dynamic weight;

[0021] Based on the second dynamic weight, obtaining second fused features, each of the second fused features corresponds to a sample in the labeled sample set;

[0022] Based on the second fusion feature, a lightweight classification model is trained, and the model scale of the lightweight classification model is smaller than the model scale of the event detection model.

[0023] Optionally, after training a lightweight classification model, the method further includes:

[0024] Using the lightweight classification model, detecting the event detection result output by the pre-trained event detection model;

[0025] Based on the detection result, the first dynamic weight or the lightweight classification model is adjusted.

[0026] Optionally, adjusting the pre-trained event detection model or the lightweight classification model based on the detection result includes:

[0027] If the detection result is correct, modifying the first dynamic weight based on the second dynamic weight;

[0028] If the detection result is wrong, the lightweight classification model is adjusted.

[0029] In a second aspect, an embodiment of the present invention further provides a multimodal event detection device, the multimodal event detection device comprising:

[0030] An acquisition module, used to acquire the to-be-processed data of multiple modes of the target scene;

[0031] A feature extraction module, used to extract features from the data to be processed according to the modality, to obtain feature data of the multiple modalities, wherein the feature data of the same modality corresponds to the data to be processed;

[0032] A feature fusion module, used for fusing the feature data according to a first dynamic weight to obtain a first fusion feature, where each modality corresponds to a first dynamic weight;

[0033] The event detection module is used to perform event detection on the fusion features through a pre-trained event detection model to obtain an event detection result of the target scene.

[0034] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the multimodal event detection method provided in the embodiment of the present invention when executing the computer program.

[0035] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the multimodal-based event detection method provided in the embodiment of the invention are implemented.

[0036] In an embodiment of the present invention, multiple modal data to be processed of a target scene are obtained; feature extraction is performed on the data to be processed according to the modality to obtain feature data of multiple modalities, and the feature data of the same modality corresponds to the data to be processed; the feature data is fused according to the first dynamic weight to obtain a first fused feature, each modality corresponds to a first dynamic weight, and the first dynamic weight is determined according to the environment of the target scene; event detection is performed on the fused feature through a pre-trained event detection model to obtain an event detection result of the target scene. The present invention introduces multimodal data, combines the features corresponding to the data of multiple modalities, and utilizes their respective advantages to complement each other, thereby significantly improving the accuracy and robustness of event recognition in public security and stability maintenance, while maintaining stable recognition performance under different environmental conditions, which can effectively prevent the risk of single modal data being interfered with or forged, and improve the recognition accuracy of event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0038] Figure 1 is a flow chart of a multi-modal event detection method provided by an embodiment of the present invention;

[0039] Figure 2 is a structural schematic diagram of a multi-modal event detection device provided by an embodiment of the present invention;

[0040] Figure 3 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] like Figure 1 As shown, Figure 1 1 is a flowchart of a multimodal event detection method provided by an embodiment of the present invention, the multimodal event detection method comprising the steps of:

[0043] 101. Obtain data to be processed of multiple modes of a target scene.

[0044] In an embodiment of the present invention, the above-mentioned event detection method based on multimodality can be applied to an event detection platform, and the above-mentioned event detection platform can be constructed based on a server or a distributed server. The above-mentioned event detection platform includes a data interface (available for user upload), an event detection model, a lightweight classification model, and an event detection program based on multimodality. The above-mentioned data interface can be used to obtain data to be processed or training data, and the above-mentioned data to be processed can be data of multiple modes. The event detection program based on multimodality can be used to implement each step of the event detection method based on multimodality.

[0045] The above target scenes may be scenes where event detection is required, such as supermarkets, parks, traffic intersections, etc.

[0046] The above-mentioned multi-modal event detection method can be used in public security and public security and stability maintenance scenarios to detect public security and stability maintenance events. The above-mentioned multiple modalities of data to be processed can be data in modalities such as graphic image data, infrared data, thermal imaging data, and audio data. The data to be processed in different modalities can be obtained through different devices. The above-mentioned graphic image data can be obtained by shooting with a high-definition camera, the above-mentioned infrared data can be obtained by shooting with an infrared camera, the above-mentioned thermal imaging data can be obtained by a thermal imaging device, and the above-mentioned audio data can be obtained by an audio device (such as a sound sensor).

[0047] In a possible embodiment, after acquiring data of different modalities through different devices, the data of different modalities may be preprocessed to obtain data to be processed after preprocessing of data of different modalities. Data of different modalities may correspond to different preprocessing methods. For example, the graphic image data and audio data are subjected to denoising and enhancement processing, and the infrared data and thermal imaging data are subjected to denoising, elimination, and data correction processing to reduce the influence of various interference factors, and obtain preprocessed data of each modality.

[0048] 102. Perform feature extraction on the data to be processed according to the modality to obtain feature data of multiple modalities.

[0049] In an embodiment of the present invention, after obtaining the data to be processed corresponding to each modality, feature extraction can be performed on the data to be processed through the corresponding feature extraction network to obtain feature data corresponding to the data to be processed of different modalities, and the feature data of the same modality corresponds to the data to be processed.

[0050] The feature extraction network can be a convolutional neural network (CNN). One or more feature extraction networks are trained for each modality. When feature extraction is required, the feature extraction network under the same modality is matched according to the modality of the data to be processed, and the matched feature extraction network is used to extract features of the data to be processed to obtain feature data corresponding to the data to be processed. Convolutional neural networks (CNN) are used to extract features that need to be focused on in different modal data.

[0051] Specifically, the data to be processed may be data in modalities such as graphic image data, infrared data, thermal imaging data, and audio data. Feature extraction may be performed on graphic image data to obtain image features, feature extraction may be performed on infrared data to obtain infrared features, feature extraction may be performed on thermal imaging data to obtain thermal imaging features, and feature extraction may be performed on audio data to obtain audio features.

[0052] Among them, image features: extract visual features in video frames; infrared features: extract shape, contour and other features in infrared images; thermal imaging features: extract temperature distribution, hot spots and other features; audio features: extract time domain waveform, frequency domain features and time-frequency domain features.

[0053] In a possible embodiment, the feature data may also be normalized. Specifically, feature data of different modalities may be normalized. Data of different modalities have different feature dimensions and representations, and appropriate feature mapping is required based on actual conditions. At the same time, canonical correlation analysis (CCA) is used to transform features of different modalities into a shared space so that they can be aligned with each other.

[0054] 103. Fuse the feature data according to the first dynamic weight to obtain a first fused feature.

[0055] In the embodiment of the present invention, each modality corresponds to a first dynamic weight, which is determined according to the environment of the target scene.

[0056] Multimodal feature fusion can be performed by weighted summation or weighted average. Weights are assigned according to the importance of each modal feature, and then the feature vectors of each modality are multiplied by their corresponding weights, and then the results are added. The weight assignment of each modal feature adopts a dynamic assignment strategy. The light intensity features, thermal imaging and infrared features in the image are used to determine whether it is a high-light or low-light condition. Different environmental conditions can correspond to different first dynamic weights. For example, during the day or under high-light conditions, the image weight accounts for the largest proportion, and the initialization weights are set to image feature weight: 0.55, infrared feature weight: 0.15, thermal imaging feature weight: 0.15, and audio feature weight: 0.15. At night or under low-light conditions, the image effect is reduced, and infrared and thermal imaging features can make up for the lack of image information. The initialization weights can be set to image feature weight: 0.35, infrared feature weight: 0.30, thermal imaging feature weight: 0.30, and audio feature weight: 0.15.

[0057] In a possible embodiment, the optimal first dynamic weight may be modified and determined through model effect and cross-validation.

[0058] 104. Event detection is performed on the fused features through the pre-trained event detection model to obtain the event detection result of the target scene.

[0059] In an embodiment of the present invention, the above-mentioned event detection model is a pre-trained model, and therefore, can be used directly. The above-mentioned pre-trained event detection model can be obtained by training the event detection model to be trained through a sample data set. Specifically, multi-modal sample data can be collected, and the modality of the sample data can be the same as the modality of the data to be processed, that is, the sample data can be data of modalities such as graphic image data, infrared data, thermal imaging data, and audio data. Of course, the above-mentioned sample data can also be pre-processed sample data, and sample data of different modalities can correspond to different pre-processing methods. For example, the graphic image data and audio data are subjected to denoising and enhancement processing, and the infrared data and thermal imaging data are subjected to denoising, elimination, and data correction processing to reduce the influence of each interference factor, and obtain the pre-processed data of each modality. The sample data has a corresponding event label, and further, the sample data of each modality corresponds to its own event label. The output of the above-mentioned event detection model to be trained is an event detection result, which includes an event type and an event confidence. Each event type corresponds to an event confidence. The event confidence is used to indicate the reliability of the event type. The larger the event confidence, the more reliable the event type is. Conversely, the less reliable the event type is. The sample data in the sample data set are subjected to feature extraction by modality to obtain sample feature data of multiple modalities. The sample feature data are fused according to the first dynamic weight to obtain a first sample fusion feature. The first sample fusion feature is input into the event detection model to be trained, and the event detection processing is performed by the event detection model to be trained. The event detection result is output. The event detection result includes the predicted event type and the event confidence. The error loss between the predicted event type and the event label is calculated by the loss function. The error loss is minimized as the optimization goal. The model parameters of the event detection model to be trained are adjusted, and the adjustment process of the model parameters is iterated. After the training stop condition is reached, the training is stopped to obtain the pre-trained event detection model. The above-mentioned loss function can be a cross entropy loss function or an average error loss function.

[0060] The first fusion feature is processed for event detection through a pre-trained event detection model, and the event type and event confidence are output. The event type whose event confidence is greater than the confidence threshold is determined as the event type of the target scene, thereby obtaining the event detection result of the target scene. The event detection result of the target scene may include at least one event type whose event confidence is greater than the confidence threshold.

[0061] In an embodiment of the present invention, multiple modal data to be processed of a target scene are obtained; feature extraction is performed on the data to be processed according to the modality to obtain feature data of multiple modalities, and the feature data of the same modality corresponds to the data to be processed; the feature data is fused according to the first dynamic weight to obtain a first fused feature, each modality corresponds to a first dynamic weight, and the first dynamic weight is determined according to the environment of the target scene; event detection is performed on the fused feature through a pre-trained event detection model to obtain an event detection result of the target scene. The present invention introduces multimodal data, combines the features corresponding to the data of multiple modalities, and utilizes their respective advantages to complement each other, thereby significantly improving the accuracy and robustness of event recognition in public security and stability maintenance, while maintaining stable recognition performance under different environmental conditions, which can effectively prevent the risk of single modal data being interfered with or forged, and improve the recognition accuracy of event detection.

[0062] It is understandable that in the specific implementation of this application, multimodal data, scenario data, model data and other related data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data, as well as the training, deployment and calling of algorithm models, must comply with relevant laws, regulations and standards of relevant countries and regions.

[0063] Optionally, in the step of performing feature extraction on the data to be processed according to the modality to obtain feature data of multiple modalities, feature extraction networks corresponding to different modalities can be determined, and different modalities correspond to different feature extraction networks; based on the feature extraction networks corresponding to each modality, feature extraction is performed on the data to be processed of each modality to obtain feature data of multiple modalities.

[0064] In an embodiment of the present invention, after obtaining the data to be processed corresponding to each modality, feature extraction can be performed on the data to be processed through the corresponding feature extraction network to obtain feature data corresponding to the data to be processed of different modalities, and the feature data of the same modality corresponds to the data to be processed.

[0065] Convolutional neural networks can be used to extract features that need to be focused on in different modal data. The above-mentioned feature extraction network can be a convolutional neural network. One or more feature extraction networks are trained for each modality. When feature extraction is required, the feature extraction network under the same modality is matched according to the modality of the data to be processed, and the matched feature extraction network is used to extract features from the data to be processed to obtain feature data corresponding to the data to be processed.

[0066] Specifically, the data to be processed may be data in modalities such as graphic image data, infrared data, thermal imaging data, and audio data. Feature extraction may be performed on graphic image data to obtain image features, feature extraction may be performed on infrared data to obtain infrared features, feature extraction may be performed on thermal imaging data to obtain thermal imaging features, and feature extraction may be performed on audio data to obtain audio features.

[0067] Optionally, in the step of fusing the feature data according to the first dynamic weight to obtain the first fused feature, the feature data of each modality can be converted into a unified feature space to obtain multiple unified feature data, each unified feature data corresponds to the feature data of one modality; the unified feature data are fused according to the first dynamic weight to obtain the first fused feature.

[0068] In an embodiment of the present invention, feature data of different modalities can be converted into a unified feature space by normalization or vector encoding. Feature data of different modalities can be normalized. Data of different modalities have different feature dimensions and representations. Appropriate feature mapping needs to be performed according to actual conditions. At the same time, canonical correlation analysis (CCA) is used to convert features of different modalities into a shared unified feature space so that they can be aligned with each other. Feature data of different modalities can also be vector encoded to convert features of different modalities into a shared unified feature space so that they can be aligned with each other.

[0069] In the unified feature space, multimodal feature fusion is performed by weighted summation or weighted average, and weights are assigned according to the importance of each modal feature. Specifically, the weighted summation formula is expressed as: [\mathbf{F}{\text{first fusion feature}}=w_1\cdot\mathbf{F}{\text{image feature}}+w_2\cdot\mathbf{F}{\text{infrared feature}}+w_3\cdot\mathbf{F}{\text{thermal imaging feature}}+w_4\cdot\mathbf{F}_{\text{audio feature}}], where (w_1,w_2,w_3,w_4) is the first dynamic weight of each modal feature, and (\mathbf{F}{\text{image feature}},\mathbf{F}{\text{infrared feature}},\mathbf{F}{\text{thermal imaging feature}},\mathbf{F}{\text{audio feature}}) is the normalized feature vector.

[0070] Optionally, in the step of fusing the unified feature data according to the dynamic weight to obtain the first fused feature, the first dynamic weight under each modality can be determined based on the environmental conditions of the target scene, and different environmental conditions correspond to different first dynamic weights under each modality; the unified feature data is weightedly added using the first dynamic weight to obtain the first fused feature.

[0071] In an embodiment of the present invention, the weight distribution of each modal feature data adopts a dynamic distribution strategy. The distribution strategy may be that different environmental conditions correspond to different weights, that is, the determination of the above-mentioned first dynamic weight is based on the environmental conditions of the target scene. The above-mentioned environmental conditions may include multiple factors such as light intensity, noise level, and degree of occlusion. The environmental conditions have a direct impact on the availability and reliability of different modal data. For example, it is possible to determine whether it is a high-light or low-light condition through the light intensity characteristics, thermal imaging characteristics, and infrared characteristics in the image, and different weights can be defined for different lighting conditions. For example, in a low-light scene, the feature data of the image modality may be greatly affected, while the feature data of the thermal imaging modality, the feature data of the infrared modality, and the feature data of the audio modality may be relatively more reliable. Therefore, in this case, the feature data of the thermal imaging modality, the feature data of the infrared modality, and the feature data of the audio modality can be given a greater weight when fused.

[0072] For example, during the day or under high light conditions, the image weight accounts for the largest proportion, and the initial weights are set to image feature weight: 0.55, infrared feature weight: 0.15, thermal imaging feature weight: 0.15, and audio feature weight: 0.15. At night or in low light conditions, the image effect is reduced, and infrared and thermal imaging features can make up for the lack of image information. The initial weights can be set to image feature weight: 0.35, infrared feature weight: 0.30, thermal imaging feature weight: 0.30, and audio feature weight: 0.15.

[0073] Optionally, after the step of performing event detection on the fused features through a pre-trained event detection model to obtain the event detection results of the target scene, positive and negative samples of each modality can be labeled based on the event detection results to obtain a labeled sample set, where the labeled sample set includes an equal number of positive samples and negative samples; based on the negative samples in the labeled sample set, the dynamic weights corresponding to each modality are adjusted to obtain a second dynamic weight; based on the second dynamic weight, a second fused feature is obtained, each second fused feature corresponds to a sample in the labeled sample set; based on the second fused feature, a lightweight classification model is trained, and the model scale of the lightweight classification model is smaller than the model scale of the event detection model.

[0074] In an embodiment of the present invention, after obtaining the event detection result, the event detection result can be sent to a labeler, who labels the data of each modality with positive and negative samples based on the event detection result. The above positive samples refer to those samples that are correctly identified as events by the event detection model, while the above negative samples refer to those samples that are incorrectly identified or not identified as events.

[0075] The event detection model is used to detect events on the first fusion feature, and the event detection results are output. Positive and negative samples are labeled based on the output results. Negative samples mark the correctness and errors of modalities such as images, sounds, infrared or thermal imaging, and the same number of positive samples are selected to construct a labeled sample set. Since both positive and negative samples are multimodal sample data, multimodal feature data can be extracted.

[0076] For negative samples, the weights of different modal features are adjusted according to the labeled erroneous modal data. For example, if the event detection result corresponding to the audio modality is wrong, the first dynamic weight corresponding to the audio modality is reduced. If the event detection result corresponding to the thermal imaging modality is correct, the first dynamic weight corresponding to the thermal imaging modality is appropriately increased. At the same time, combined with the fusion weight of the positive sample, the first dynamic weight of the multimodal feature data is dynamically adjusted to obtain the second dynamic weight. The multimodal feature data is fused through the second dynamic weight to obtain a new fusion feature, namely the second fusion feature. For positive samples, all of which are labeled correct modal data, there is no need to adjust the first dynamic weight of the positive sample.

[0077] By annotating the second fusion features corresponding to the negative samples in the data set and the second fusion features corresponding to the positive samples, a lightweight classification model is trained. The lightweight classification model can be an SVM support vector machine or an SSD neural network model. The above lightweight classification model can also be called a small classification model. The model scale of the lightweight classification model is smaller than the model scale of the event detection model.

[0078] The above-mentioned annotated data set is annotated according to the event detection results output by the event detection model, and can be understood as a real-time or periodically updated annotated data set. After the accuracy of the event detection results output by the event detection model drops below an accuracy threshold, the real-time training of the lightweight classification model can be started. It should be noted that when the lightweight classification model is trained in real time, event detection is still required through the event detection model, but the event detection results are only used to obtain the annotated data set to collect more negative samples for the event detection model.

[0079] In a possible embodiment, the trained quantitative classification model can be directly used for event detection. At this time, the event detection model can also be trained and adjusted through a labeled data set to improve the accuracy of the event detection model. When the accuracy of the event detection model is improved to above the accuracy threshold, event detection is performed through the event detection model. At this time, the event detection results can be sent to the user.

[0080] Optionally, after the step of training a lightweight classification model, the event detection results output by the pre-trained event detection model can be detected by the lightweight classification model; based on the detection results, the first dynamic weight or the lightweight classification model is adjusted.

[0081] In an embodiment of the present invention, the above-mentioned lightweight classification model can also be used for verification testing. After the data to be processed passes through steps 102, 103, and 104, and the event detection result output by the event detection model is obtained, it can be re-detected by the lightweight classification model to output the final detection result. The event detection result is correct or wrong. If the event detection result is correct, the first dynamic weight can be adjusted, and if the event detection result is wrong, the lightweight classification model can be adjusted.

[0082] Optionally, in the step of adjusting the pre-trained event detection model or lightweight classification model based on the detection result, if the detection result is correct, the first dynamic weight can be corrected based on the second dynamic weight; if the detection result is wrong, the lightweight classification model can be adjusted.

[0083] In an embodiment of the present invention, if the event detection result is correct after the secondary classification test of the lightweight classification model, the samples originally identified as negative samples can be labeled as positive samples without affecting the original positive samples, and the first dynamic weight can be adjusted. Through the adjusted first dynamic weight, a new first fusion feature is obtained to train and adjust the event detection model. If the event detection result is wrong after the secondary classification test of the lightweight classification model, the lightweight classification model and the event detection model are continuously trained and fine-tuned again through the labeled data set, and finally an event detection model and a lightweight classification model that can identify correctly are obtained.

[0084] It is understandable that if the detection result of the lightweight classification model is correct, it means that the output of the event detection model is consistent with the judgment of the lightweight classification model. In this case, it can be considered that the current first dynamic weight configuration is effective, and the accuracy of the event detection model is also high, which can better reflect the contribution of different modal data in event detection.

[0085] Of course, in a possible embodiment, in order to further improve the performance of event detection, the first dynamic weight may be fine-tuned using additional information provided by the lightweight classification model.

[0086] In one possible embodiment, the performance of the lightweight classification model in processing the annotated sample set can be statistically analyzed, especially for those samples that were originally misclassified or missed by the event detection model. By analyzing the correct classification of the misclassified or missed samples in the lightweight classification model, the second dynamic weights of each modality can be fine-tuned to better reflect the actual importance of multi-modal feature data fusion in event detection. The above fine-tuning can be based on statistical data, for example, adjusting its weight according to the contribution of the modal data in the correct classification.

[0087] If the detection result of the lightweight classification model is wrong, it means that the output of the pre-trained event detection model is inconsistent with the judgment of the lightweight classification model. In this case, the lightweight classification model can be adjusted to improve the detection accuracy of the lightweight classification model.

[0088] The method of adjusting the lightweight classification model can be to increase the number and diversity of the labeled sample set to provide more learning samples and scenarios for the lightweight classification model. Online learning or incremental learning methods can also be used to enable the lightweight classification model to continuously learn and adapt from new data, thereby improving its generalization ability and detection accuracy.

[0089] In a possible embodiment, the event detection model and the lightweight model are deployed together in the event detection product. According to the event detection results of the event detection model, positive samples and negative samples are labeled to obtain a labeled data set, and the event detection model and the lightweight model are updated in real time through the labeled data set.

[0090] like Figure 2 As shown, an embodiment of the present invention provides a multi-modal event detection device, the multi-modal event detection device comprising:

[0091] An acquisition module 201 is used to acquire multiple modal data to be processed of a target scene;

[0092] A feature extraction module 202 is used to extract features from the data to be processed according to the modality to obtain feature data of the multiple modalities, wherein the feature data of the same modality corresponds to the data to be processed;

[0093] A feature fusion module 203 is used to fuse the feature data according to a first dynamic weight to obtain a first fusion feature, where each mode corresponds to a first dynamic weight;

[0094] The event detection module 204 is used to perform event detection on the fusion features through a pre-trained event detection model to obtain an event detection result of the target scene.

[0095] Optionally, the feature extraction module 202 is also used to determine feature extraction networks corresponding to different modalities, and different modalities correspond to different feature extraction networks; based on the feature extraction networks corresponding to each modality, feature extraction is performed on the to-be-processed data of each modality to obtain feature data of the multiple modalities.

[0096] Optionally, the feature fusion module 203 is also used to convert the feature data of each modality into a unified feature space to obtain multiple unified feature data, each unified feature data corresponds to the feature data of one modality; and fuse the unified feature data according to a first dynamic weight to obtain a first fused feature.

[0097] Optionally, the feature fusion module 203 is also used to determine the first dynamic weight under each modality based on the environmental conditions of the target scene, and different environmental conditions correspond to different first dynamic weights under each modality; the unified feature data is weightedly added through the first dynamic weight to obtain a first fusion feature.

[0098] Optionally, the device further comprises:

[0099] A labeling module, used to label positive and negative samples of each modality based on the event detection result to obtain a labeled sample set, wherein the labeled sample set includes the same number of positive samples and negative samples;

[0100] A first adjustment module, configured to adjust the dynamic weights corresponding to each modality based on the negative samples in the labeled sample set to obtain a second dynamic weight;

[0101] A processing module, configured to obtain second fused features based on the second dynamic weight, each of the second fused features corresponding to a sample in the labeled sample set;

[0102] A training module is used to train a lightweight classification model based on the second fusion feature, and the model scale of the lightweight classification model is smaller than the model scale of the event detection model.

[0103] Optionally, the device further comprises:

[0104] A detection module, configured to detect the event detection result output by the pre-trained event detection model through the lightweight classification model;

[0105] The second adjustment module is used to adjust the first dynamic weight or the lightweight classification model based on the detection result.

[0106] Optionally, the second adjustment module is further used to correct the first dynamic weight based on the second dynamic weight if the detection result is correct; and adjust the lightweight classification model if the detection result is wrong.

[0107] like Figure 3 As shown, an embodiment of the present invention further provides an electronic device, including a processor, and the processor can execute any of the above-mentioned multi-modal event detection methods.

[0108] Specifically, it includes a processor 301 and a memory 302, and a computer program for executing a multi-modal event detection method stored in the memory 302 and capable of running on the processor 301, wherein:

[0109] The processor 301 runs the computer program based on the multi-modal event detection method stored in the memory 302 to perform the following steps:

[0110] Obtaining multiple modal data to be processed of the target scene;

[0111] Extracting features of the data to be processed respectively according to the modality to obtain feature data of the multiple modalities, wherein the feature data of the same modality corresponds to the data to be processed;

[0112] The feature data are fused according to a first dynamic weight to obtain a first fused feature, each modality corresponds to a first dynamic weight, and the first dynamic weight is determined according to the environment of the target scene;

[0113] Event detection is performed on the fused features using a pre-trained event detection model to obtain an event detection result for the target scene.

[0114] Optionally, the processor 301 performs feature extraction on the data to be processed according to the modality to obtain feature data of the multiple modalities, including:

[0115] Determine the feature extraction networks corresponding to different modalities. Different modalities correspond to different feature extraction networks.

[0116] Based on the feature extraction network corresponding to each modality, feature extraction is performed on the to-be-processed data of each modality to obtain feature data of the multiple modalities.

[0117] Optionally, the processor 301 performs the step of fusing the feature data according to a first dynamic weight to obtain a first fused feature, including:

[0118] Convert the feature data of each modality into a unified feature space to obtain a plurality of unified feature data, each of which corresponds to the feature data of one modality;

[0119] The unified feature data are fused according to a first dynamic weight to obtain a first fused feature.

[0120] Optionally, the processor 301 performs the step of fusing the unified feature data according to a dynamic weight to obtain a first fused feature, including:

[0121] Based on the environmental conditions of the target scene, determining the first dynamic weights in each mode, where different environmental conditions correspond to different first dynamic weights in each mode;

[0122] The unified feature data are weightedly added using the first dynamic weight to obtain a first fused feature.

[0123] Optionally, after performing event detection on the fused feature through the pre-trained event detection model to obtain the event detection result of the target scene, the method executed by the processor 301 further includes:

[0124] Based on the event detection result, positive and negative samples are labeled for each modality to obtain a labeled sample set, where the labeled sample set includes an equal number of positive samples and negative samples;

[0125] Based on the negative samples in the labeled sample set, adjusting the dynamic weights corresponding to each modality to obtain a second dynamic weight;

[0126] Based on the second dynamic weight, obtaining second fused features, each of the second fused features corresponds to a sample in the labeled sample set;

[0127] Based on the second fusion feature, a lightweight classification model is trained, and the model scale of the lightweight classification model is smaller than the model scale of the event detection model.

[0128] Optionally, after training a lightweight classification model, the method executed by the processor 301 further includes:

[0129] Using the lightweight classification model, detecting the event detection result output by the pre-trained event detection model;

[0130] Based on the detection result, the first dynamic weight or the lightweight classification model is adjusted.

[0131] Optionally, the processor 301 performs the adjustment of the pre-trained event detection model or the lightweight classification model based on the detection result, including:

[0132] If the detection result is correct, modifying the first dynamic weight based on the second dynamic weight;

[0133] If the detection result is wrong, the lightweight classification model is adjusted.

[0134] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the multimodal event detection method provided by the embodiment of the present invention are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0135] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0136] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A multimodal event detection method, characterized in that: The method comprises the following steps: Obtaining multiple modal data to be processed of the target scene; Extracting features of the data to be processed respectively according to the modality to obtain feature data of the multiple modalities, wherein the feature data of the same modality corresponds to the data to be processed; The feature data are fused according to a first dynamic weight to obtain a first fused feature, each modality corresponds to a first dynamic weight, and the first dynamic weight is determined according to the environment of the target scene; Event detection is performed on the fused features using a pre-trained event detection model to obtain an event detection result for the target scene.

2. The multimodal event detection method according to claim 1, characterized in that: The extracting features of the data to be processed respectively according to the modality to obtain the feature data of the multiple modalities includes: Determine the feature extraction networks corresponding to different modalities. Different modalities correspond to different feature extraction networks. Based on the feature extraction network corresponding to each modality, feature extraction is performed on the to-be-processed data of each modality to obtain feature data of the multiple modalities.

3. The multimodal event detection method according to claim 2, characterized in that: The step of fusing the feature data according to a first dynamic weight to obtain a first fused feature includes: Convert the feature data of each modality into a unified feature space to obtain a plurality of unified feature data, each of which corresponds to the feature data of one modality; The unified feature data are fused according to a first dynamic weight to obtain a first fused feature.

4. The multimodal event detection method according to claim 3, characterized in that: The step of fusing the unified feature data according to the dynamic weight to obtain the first fused feature includes: Based on the environmental conditions of the target scene, determining the first dynamic weights in each mode, where different environmental conditions correspond to different first dynamic weights in each mode; The unified feature data are weightedly added using the first dynamic weight to obtain a first fused feature.

5. The multimodal event detection method according to any one of claims 1 to 4, characterized in that: After performing event detection on the fused features through the pre-trained event detection model to obtain the event detection result of the target scene, the method includes: Based on the event detection result, positive and negative samples are labeled for each modality to obtain a labeled sample set, where the labeled sample set includes an equal number of positive samples and negative samples; Based on the negative samples in the labeled sample set, adjusting the dynamic weights corresponding to each modality to obtain a second dynamic weight; Based on the second dynamic weight, obtaining second fused features, each of the second fused features corresponds to a sample in the labeled sample set; Based on the second fusion feature, a lightweight classification model is trained, and the model scale of the lightweight classification model is smaller than the model scale of the event detection model.

6. The multimodal event detection method according to claim 5, characterized in that: After training a lightweight classification model, the method further includes: Using the lightweight classification model, detecting the event detection result output by the pre-trained event detection model; Based on the detection result, the first dynamic weight or the lightweight classification model is adjusted.

7. The multimodal event detection method according to claim 6, characterized in that: The adjusting the pre-trained event detection model or the lightweight classification model based on the detection result includes: If the detection result is correct, modifying the first dynamic weight based on the second dynamic weight; If the detection result is wrong, the lightweight classification model is adjusted.

8. A multimodal event detection device, characterized in that: The multimodal event detection device comprises: An acquisition module, used to acquire the to-be-processed data of multiple modes of the target scene; A feature extraction module, used to extract features from the data to be processed according to the modality, to obtain feature data of the multiple modalities, wherein the feature data of the same modality corresponds to the data to be processed; A feature fusion module, used for fusing the feature data according to a first dynamic weight to obtain a first fusion feature, where each modality corresponds to a first dynamic weight; The event detection module is used to perform event detection on the fusion features through a pre-trained event detection model to obtain an event detection result of the target scene.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the multimodal-based event detection method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the multi-modality-based event detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Security behavior event identification method and system based on multi-modal analysis

    CN120544129A

  • Automatic driving system reliability self-diagnosis method and platform based on multi-modal sensor fusion

    CN120708310A

  • Non-motor vehicle illegal event detection method and device, electronic equipment and storage medium

    CN120724233A

  • Automatic monitoring method for infrastructure construction operation violation based on multi-modal data fusion

    CN120974243A

  • A method for automatically monitoring construction operation violations of infrastructure construction by multi-modal data fusion

    CN120974243B