Safety early warning method and system based on multi-modal data fusion and storage medium

Through multimodal data fusion, sensors, Internet of Things devices and audio data are used to extract features and calculate correlations, and the safety warning model is trained, which solves the problem of low accuracy caused by single data in traditional early warning methods, and achieves more efficient safety warning.

CN120236172APending Publication Date: 2025-07-01JIAXING JIEDAO INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510319613.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Traditional security early warning methods rely on a single data source, resulting in a single data dimension, incomplete information, low warning accuracy, and ineffective response to complex public safety threats.

Method used

Multimodal data fusion method is adopted to collect multimodal data by deploying sensors, IoT devices, shooting devices and audio acquisition devices, and extract features using time series analysis, machine learning and deep learning models, calculate modal feature correlations and adjust attention weights, generate training sets and test sets, train safety warning models and update parameters.

Benefits of technology

It improves the accuracy and reliability of the safety warning model, can more comprehensively identify and predict safety events, prevent modal overfitting or underfitting, and achieve faster and more accurate early warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236172A_ABST
    Figure CN120236172A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of safety early warning, and discloses a safety early warning method and system based on multi-modal data fusion and a storage medium. The method comprises the following steps: deploying data acquisition equipment in a target environment, acquiring multi-modal data, and preprocessing the multi-modal data; respectively extracting features of each modal data; calculating correlation among different modal features, calculating an attention weight based on a correlation score, and introducing a gating mechanism to adjust the attention weight in combination with historical data; multiplying the adjusted attention weight by each modal data to obtain weighted modal data, and splicing to obtain fused modal data; and generating a training set and a test set based on the fusion modal data, training a safety early warning model, and outputting an early warning result through the test set. And calculating a prediction error according to the early warning result, and updating parameters of the safety early warning model according to the prediction error. The accuracy and reliability of the safety early warning model can be improved, and the false alarm rate is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of security warning technologies, and particularly to a security warning method, system, and storage medium based on multi-modal data fusion. Background Art

[0002] With the development of society and the acceleration of the urbanization process, public security issues have become increasingly prominent. Emergencies such as fires, traffic accidents, and natural disasters occur frequently, posing a great threat to people's lives and property safety. Traditional security warning methods mainly rely on single data sources, such as temperature sensors, cameras, etc., and have problems such as single data dimension, incomplete information, and low warning accuracy. Therefore, there is an urgent need for a security warning method based on multi-modal data fusion to improve the accuracy and reliability of the warning model and better ensure public security.

[0003] A similar prior art Chinese patent application with the publication number CN116881850A provides a security warning system based on multi-modal data fusion, including: a multi-modal data fusion recognition module for obtaining the operation data during the multi-modal data fusion process, identifying the fusion integrity during the multi-modal data fusion process based on the operation data, and sending the multi-modal data fusion integrity to the server; a multi-modal data fusion feedback module receiving the signal of good multi-modal data fusion integrity transmitted by the server, obtaining the multi-modal data feedback end data according to the data corresponding to the good multi-modal data fusion integrity, identifying the working state of the multi-modal data feedback end based on the multi-modal data feedback end data, and sending the working state of the multi-modal data feedback end to the server; a multi-modal data warning module receiving the abnormal data signal of poor fusion integrity transmitted by the server and the feedback signal of abnormal working of the multi-modal data feedback end for warning. This method compares the benchmark value RH of multi-modal data fusion with a preset threshold Rh to judge the fusion integrity, and this judgment method based on a fixed formula and a preset threshold may not be able to fully and accurately reflect the complex situation in the actual fusion process.

[0004] A similar prior art is also the Chinese patent application with the publication number CN117852867A, which provides a security risk classification and early warning system based on multi-modal data, including a data collection module, an anomaly filtering module, a data conversion module, a text analysis module, and a classification and early warning module. Through the data collection module and the anomaly filtering module, it is possible to process various modal comment data such as images, videos, audios, and texts. The addition of the anomaly filtering module; through the data conversion module, it is possible to convert various modal comment data such as images, videos, and audios into text data; in the text analysis module, a special food name library is constructed, and in combination with the stop word library, food names similar to or containing keywords related to food spoilage are filtered. However, although this method sets various filtering strategies in the anomaly filtering module, such as screening abnormal comments based on time difference, negative review ratio, image similarity, and video key frame image similarity, these strategies may not be able to fully cover all types of malicious negative reviews or interference items.

[0005] Therefore, the present invention provides a security early warning method, system, and storage medium based on multi-modal data fusion. Summary of the Invention

[0006] This application provides a security early warning method based on multi-modal data fusion, which is used to improve the accuracy and reliability of the security early warning model and better ensure public safety.

[0007] In a first aspect, this application provides a security early warning method based on multi-modal data fusion, and the method includes:

[0008] Step S1: Deploy data collection devices in the target environment to collect multi-modal data of the target environment. The data collection devices include various sensors, Internet of Things devices, shooting devices, and audio collection devices. The multi-modal data includes sensor data, Internet of Things device data, video data, and audio data. According to the characteristics of the collected data and the early warning requirements, set the data collection frequency for different data collection devices, and preprocess the collected multi-modal data;

[0009] Step S2: Use a time series analysis model to extract the state change characteristics of the sensor data, use a machine learning algorithm to extract the device state characteristics of the Internet of Things devices from the Internet of Things device data, use a pre-trained convolutional neural network to extract the video characteristics of the video frames, and use a deep learning model to extract the audio characteristics of the audio data. Collectively, the state change characteristics, device state characteristics, video characteristics, and audio characteristics are referred to as modal characteristics;

[0010] Step S3: Calculate the correlation between different modal features, calculate the corresponding attention weights using the Sigmoid function based on the correlation scores, and also introduce a gating mechanism to adjust the attention weights in combination with historical data. Multiply the adjusted attention weights by each modal data to obtain weighted modal data, and splice the weighted modal data to obtain fused modal data;

[0011] Step S4: Generate a training set and a test set based on the fused modal data, train a security warning model based on the training set, input the test set into the security warning model to output a warning result, calculate the prediction error based on the warning result, and update the parameters of the security warning model according to the prediction error.

[0012] Combined with the first aspect, in the first implementation manner of the first aspect of this application, calculating the correlation between different modal features includes:

[0013] Align the sensor data, Internet of Things device data, video data, and audio data in the time dimension based on timestamps, use a linear transformation to map different modal features to the same dimension, introduce a non-linear transformation to the modal features after the linear transformation using a non-linear activation function, normalize the modal features after the non-linear transformation, calculate the similarity scores between different modal features, normalize the calculated similarity scores, and use the normalized similarity scores as the correlation between the corresponding different modal data.

[0014] Combined with the first aspect, in the second implementation manner of the first aspect of this application, calculating the similarity scores between different modal features includes:

[0015] Use multiple different similarity algorithms to calculate the similarity between every two different modal features to obtain multiple different basic similarity scores. For every two different modal features, calculate the sum of the absolute values of the corresponding multiple basic similarity scores. Divide the absolute value of each basic similarity score by the sum of the absolute values to obtain the initial weight factor of each similarity algorithm. Introduce a variation parameter for each similarity algorithm, adjust the weight factors of different similarity algorithms based on the variation parameter, perform a weighted operation on the adjusted weight factors and the corresponding basic similarity scores to obtain a mixed similarity score, and use the mixed similarity score as the similarity score between different modal features.

[0016] Combined with the first aspect, in the third implementation manner of the first aspect of this application, adjusting the weight factors of different similarity algorithms based on the variation parameter includes:

[0017] Initialize the change parameters, calculate the adjusted weight factor using the softmax function based on the change parameters, calculate the corresponding similarity score based on the adjusted weight factor, adjust each change parameter, calculate the similarity score after adjusting the change parameters, and calculate the change rate between the two similarity scores before and after each adjustment of the change parameter. When the change rate is less than the preset second threshold, stop adjusting the change parameter, and use the finally calculated similarity score as the final similarity score.

[0018] Combined with the first aspect, in the fourth implementation manner of the first aspect of this application, generating a training set and a test set includes:

[0019] Identify all security events to be recognized, set an event label for each security event, determine whether a security event has occurred for each modal data. If so, set the event label of the corresponding security event for the modal data; if not, set the default label for the modal data. Generate an event label vector for each fused modal data, where each data in the event label vector corresponds to the meaning of the security event type. Determine whether a security event has occurred for each modal data in the fused modal data. If so, set the corresponding data in the event label vector to the corresponding event label; if not, set the corresponding data in the event label vector to the default label. Use all the fused modal data marked with event labels as the comprehensive data set, and divide the comprehensive data set into a training set and a test set.

[0020] Combined with the first aspect, in the fifth implementation manner of the first aspect of this application, training a security warning model based on the training set includes:

[0021] Individually train each modal data in the fused modal data to generate an individual warning model. During the process of generating the individual warning model, verify the accuracy of the individual warning model based on the test set every preset first time period. If the corresponding accuracy does not improve within several consecutive training cycles, end the training of this individual warning model. Train the fused modal data to generate a comprehensive warning model. The training of the comprehensive warning model is carried out in parallel with the training of each individual warning model. Fuse the generated multiple individual warning models to generate a fused warning model. Compare the accuracy of the warning results of the fused warning model and the comprehensive warning model, and use the model with the higher accuracy as the security warning model.

[0022] Combined with the first aspect, in the sixth implementation manner of the first aspect of this application, calculating the prediction error based on the warning result includes:

[0023] When the security prediction model is a fusion early warning model, obtain the prediction accuracy of each individual early warning model, use the prediction accuracy as the reliability score of each modality data, calculate the sum of all reliability scores, divide each reliability score by the sum of scores, and use the result as the error weight corresponding to each modality data. Compare the early warning result of each individual early warning model with the pre-set event label vector, calculate the prediction error of each individual early warning model, and use the result of weighted addition of the error weight and the corresponding prediction error as the prediction error of the security early warning model. When the security early warning model is a comprehensive early warning model, calculate the difference between the prediction result of the comprehensive model and the pre-set event label vector as the prediction error.

[0024] Combined with the first aspect, in the seventh implementation manner of the first aspect of this application, update the parameters of the security early warning model according to the prediction error, including:

[0025] Backpropagate the calculated prediction error to the security early warning model, calculate the gradient of the parameters of the security early warning model, use an optimization algorithm to update the parameters of the security early warning model, and repeat this step until the prediction error of the security early warning model is less than or equal to the first threshold.

[0026] In a second aspect, this application provides a security early warning system based on multi-modal data fusion. The system includes:

[0027] A data collection module, which is used to deploy data acquisition devices in the target environment to collect multi-modal data of the target environment. The data acquisition devices include various sensors, Internet of Things devices, shooting devices, and audio acquisition devices. The multi-modal data includes sensor data, Internet of Things device data, video data, and audio data. According to the characteristics of the collected data and the early warning requirements, set the data collection frequency for different data acquisition devices, and preprocess the collected multi-modal data;

[0028] A feature extraction module, which is used to extract the state change features of sensor data using a time series analysis model, extract the device state features of Internet of Things devices from Internet of Things device data using a machine learning algorithm, extract the video features of video frames using a pre-trained convolutional neural network, and extract the audio features of audio data using a deep learning model. The state change features, device state features, video features, and audio features are collectively referred to as modal features;

[0029] A data fusion module, which calculates the correlation between different modal features, calculates the corresponding attention weights using the Sigmoid function based on the correlation scores, also introduces a gating mechanism to adjust the attention weights in combination with historical data, multiplies the adjusted attention weights with each modality data to obtain weighted modality data, and splices the weighted modality data to obtain fused modality data;

[0030] The model training module generates a training set and a test set based on the fused modal data, trains a security early warning model based on the training set, inputs the test set into the security early warning model to output an early warning result, calculates a prediction error based on the early warning result, and updates the parameters of the security early warning model according to the prediction error.

[0031] The third aspect of this application provides a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it enables the computer to execute the above-mentioned security early warning method based on multi-modal data fusion.

[0032] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0033] In the technical solution provided by this application, by fusing various modal data such as sensor data, Internet of Things device data, video data, and audio data, the state of the target environment can be comprehensively perceived from different perspectives. Compared with a single data source, multi-modal data fusion can provide richer and more comprehensive information, which helps to more accurately identify and predict security events; extract features from each modal data respectively, such as using a time series analysis model to extract the state change features of sensor data, and using a pre-trained convolutional neural network to extract video features of video frames, etc., which can deeply mine the key information in each modal data. At the same time, calculate the correlation between different modal features, and calculate the attention weight based on the correlation score, which can highlight the modal data that is more important for security early warning, make the model pay more attention to the features closely related to security events, and further improve the accuracy of early warning; train each modal data in the fused modal data separately to generate a separate early warning model, and dynamically adjust the training cycle according to the accuracy during the training process to prevent overfitting or underfitting of some modalities. At the same time, train the fused modal data to generate a comprehensive early warning model, fuse the separate early warning model and the comprehensive early warning model, compare the accuracy of the early warning results of different models, and select the optimal model as the security early warning model. This training strategy can make full use of the advantages of each modal data, improve the efficiency and quality of model training, and thus achieve faster and more accurate early warning. Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic diagram of an embodiment of the security early warning method based on multi-modal data fusion in the embodiments of this application;

[0036] Figure 2It is a schematic diagram of an embodiment for calculating the correlation between the features of different modality data in an embodiment of the present application;

[0037] Figure 3 It is a schematic diagram of an embodiment for calculating the similarity score between different modality features in an embodiment of the present application;

[0038] Figure 4 It is a schematic diagram of an embodiment of a security warning system based on multi-modal data fusion in an embodiment of the present application. Detailed implementation manners

[0039] The embodiments of the present application provide a security warning method, system and storage medium based on multi-modal data fusion. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than those illustrated or described here. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0040] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 An embodiment of a security warning method based on multi-modal data fusion in an embodiment of the present application includes:

[0041] Step S1: Deploy data acquisition devices in the target environment to collect multi-modal data of the target environment. The data acquisition devices include various sensors, Internet of Things devices, shooting devices and audio acquisition devices. The multi-modal data includes sensor data, Internet of Things device data, video data and audio data. According to the characteristics of the collected data and the warning requirements, set the data acquisition frequency for different data acquisition devices and preprocess the collected multi-modal data.

[0042] Specifically, to improve the accuracy of safety warnings, a variety of data collection devices are deployed in the target environment to collect multi-modal data of the target environment. The data collection devices include various sensors such as temperature sensors, smoke sensors, and gas sensors, Internet of Things devices such as access control devices and alarm devices, imaging devices such as high-definition cameras, and audio collection devices such as high-quality microphones to collect audio data in the target environment. According to the characteristics of the collected data and the warning requirements, data collection frequencies are set for different data collection devices. For example, sensor data can be collected once per second, and video data and audio data can be continuously collected. Since multi-modal data includes data collected from different collection devices, such as video, audio, and sensor data, each of these data has unique information and characteristics. Subsequently, through effective fusion methods, the advantages of each modal data can be fully utilized to achieve accurate understanding and analysis of complex scenarios. To better fuse multi-modal data, it is first necessary to preprocess the multi-modal data. The preprocessing includes data cleaning, data alignment, and data normalization. Data alignment refers to aligning sensor data, Internet of Things device data, video data, and audio data in the time dimension based on timestamps to ensure that different modal data is comparable at the same time point or within the same time window for effective feature extraction. By preprocessing the original multi-modal data, the data quality and consistency are improved, preparing for subsequent feature extraction.

[0043] Step S2: Use a time series analysis model to extract the state change features of sensor data, use a machine learning algorithm to extract the device state features of Internet of Things devices from Internet of Things device data, use a pre-trained convolutional neural network to extract the video features of video frames, and use a deep learning model to extract the audio features of audio data. The state change features, device state features, video features, and audio features are collectively referred to as modal features.

[0044] Specifically, to better analyze and fuse multi-modal data subsequently, feature extraction is performed on each modal data after preprocessing. Use an event sequence analysis model (such as LSTM) to extract the state change features of sensor data, such as temperature changes and humidity trends. Based on Internet of Things device data, extract the device state features of Internet of Things device states, such as switch states. Use a pre-trained convolutional neural network (such as ResNet) to extract the deep features of videos, such as the categories, shapes, and movement trajectories of objects. Also use a deep learning model (such as CNN) to extract the audio features of audio, such as sound events (human voices, sirens), etc. The modal features of each modal data extracted provide basic data for the subsequent construction of the knowledge graph.

[0045] Step S4: Calculate the correlation between different modal features, calculate the attention weights of each modal data according to the correlation, introduce a gating mechanism to adjust the attention weights in combination with historical data, multiply the adjusted attention weights by each modal data to obtain weighted modal data, and splice the weighted modal data to obtain fused modal data.

[0046] Specifically, in order to fuse the features of different modalities, first calculate the correlation between different modal features. In multi-modal data, the contribution degrees of different modal data to security warning are different. By calculating the correlation, the association degree between different modal data can be quantified, and the attention weights can be allocated accordingly to highlight the modal data that is more important for security warning. The specific calculation method of the correlation will be explained in detail later. Since the modal features are extracted from the corresponding modal data, calculating the correlation of the modal features can represent the correlation between the modal data.

[0047] Calculate the attention weights of each modal data based on the correlation. Since the correlation score is the score for calculating the similarity or association degree between two modal data, usually in the attention mechanism, this score is obtained by calculating the dot product or scaled dot product between the query vector (the first modal data) and the key vector (the second modal data). The correlation score reflects the similarity between each key vector and the query vector, that is, the association degree between each element and the element at the current position.

[0048] The attention weights calculated through the correlation score are the weights for the data corresponding to the key vector (the second modal data). Specifically, after the correlation score between each key vector and the query vector is normalized (for example, using the Sigmoid function), the obtained weights represent the degree of attention that the model should give to the data corresponding to each key vector when calculating the final output. The attention weights are used to measure the contribution degree of different modal data to security warning. Introduce a gating mechanism to adjust the attention weights in combination with historical data so that the model can better capture the dynamic changes in the time series. Multiply the adjusted attention weights by each modal data to obtain weighted modal data, and splice the weighted modal data to obtain fused modal data. Through weighted fusion, the modal data that is more important for security warning can be highlighted. Through the splicing operation, the feature information of different modalities is integrated into a unified data, providing more comprehensive data support for the subsequent training of the security warning model.

[0049] Step S4: Generate a training set and a test set based on the fused modal data, train a security warning model based on the training set, input the test set into the security warning model to output a warning result, calculate the prediction error based on the warning result, and update the parameters of the security warning model according to the prediction error.

[0050] Specifically, to improve the accuracy of the security warning model, a training set and a test set are generated based on the fused modal data. The security warning model is trained based on the training set, and the test set is input into the security warning model to output a warning result. The warning result is the type of security event and the corresponding accuracy. For example, the fire event has an accuracy of 0.8, and the violent intrusion event has an accuracy of 0.9. The prediction error between the warning result and the correct result is calculated based on the attention weight. The correct result refers to the pre-set event label vector. The parameters of the security warning model are updated based on the prediction error. Through the above method, the security warning model can learn the complex relationship between different modal data and security events and achieve accurate prediction of security events.

[0051] In a specific embodiment, to calculate the correlation between different modal features, the following steps are further performed:

[0052] The different modal features are mapped to the same dimension using a linear transformation, a non-linear activation function is used to introduce a non-linear transformation to the modal features after the linear transformation, the modal features after the non-linear transformation are normalized, the similarity score between different modal features is calculated, the calculated similarity score is normalized, and the normalized similarity score is used as the correlation between the corresponding different modal data.

[0053] Specifically, as Figure 2 shown in the flowchart for calculating the correlation between different modal data features. First, the sensor data, Internet of Things device data, video data, and audio data are aligned in the time dimension based on the time stamp. Through time alignment, the time deviation can be eliminated, the data synchronization can be improved, and the feature mismatch caused by time asynchronization can be avoided. Also, since different modal data have different feature dimensions, directly calculating the similarity will result in dimension mismatch. Therefore, a linear transformation is used to unify different modal features to the same dimension for similarity calculation and fusion. Then, a non-linear activation function such as ReLU is used to introduce a non-linear transformation to the modal features after the linear transformation to facilitate capturing the complex relationship between different modal features. The modal features after the non-linear transformation are normalized to make different modal features have similar scales and avoid the excessive influence of some large modal features on the similarity calculation. The normalized similarity score is used as the correlation between the corresponding different modal data. The greater the similarity, the stronger the correlation between the two modal data.

[0054] In a specific embodiment, the steps for calculating the similarity score between different modal features specifically include the following:

[0055] Calculate the similarity between every two different modality features using multiple different similarity algorithms to obtain multiple different basic similarity scores. For every two different modality features, calculate the sum of the absolute values of the corresponding multiple basic similarity scores. Divide the absolute value of each basic similarity score by the sum of the absolute values to obtain the initial weight factor for each similarity algorithm. Introduce a variation parameter for each similarity algorithm and adjust the weight factors of different similarity algorithms based on the variation parameter. Perform a weighted operation on the adjusted weight factors and the corresponding basic similarity scores to obtain a mixed similarity score, and use the mixed similarity score as the similarity score between different modality features.

[0056] Specifically, as Figure 3 shown is the flowchart for calculating the similarity score between different modality features. There are multiple similarity calculation methods, such as pre-similarity, dot product similarity, and Euclidean distance similarity, etc. Calculate the similarity between two different modality features using multiple different similarity calculation methods to obtain multiple basic similarity scores. For example, use a to represent audio features and v to represent video features. Assume that three different basic similarity scores s1, s2, and s3 are calculated for the audio feature a and the video feature v. Calculate the sum of the absolute values of the three similarity scores sum = |s1| + |s2| + |s3|. Then divide the absolute value of each basic similarity score by sum to obtain the initial weight factor for each similarity algorithm. For example, the three weight factors are w1, w2, and w3 respectively. To make the weight factors more adaptable, introduce a variation parameter (b1, b2, b3) to adjust the weight factors. The specific adjustment method will be explained in detail later. Perform a weighted operation on the adjusted weight factors and the corresponding basic similarity scores to obtain a mixed similarity score, and use the mixed similarity score as the similarity score between different modality features. By using the above method, calculate the similarity score between video features and audio using multiple different similarity algorithms, calculate the weight factor for each similarity metric based on the similarity score, combine the basic similarity score with the adaptive weight to obtain the final similarity score. By combining multiple similarity metric methods and introducing an adaptive weight, it can more flexibly adapt to different modality data and calculate a more reasonable similarity score.

[0057] In a specific embodiment, adjusting the weight factors of different similarity algorithms based on the variation parameter includes the following steps:

[0058] Initialize the change parameters, calculate the adjusted weight factors using the softmax function based on the change parameters, calculate the corresponding similarity scores based on the adjusted weight factors, adjust each change parameter, calculate the similarity scores after adjusting the change parameters, and calculate the change rate between the two similarity scores before and after each adjustment of the change parameters. When the change rate is less than the preset second threshold, stop adjusting the change parameters, and use the finally calculated similarity score as the final similarity score.

[0059] Specifically, for example, the initial values of the change parameters are 1, 1, 1 respectively, and the change parameters are adjusted to 1.1, 0.9, 1.0. Use the change parameters to weight the corresponding similarity algorithms, and calculate the adjusted weight factors. For example b1, b2, b3 are change parameters used to adjust the weights of different similarity algorithms. Calculate the weighted hybrid similarity score SW = w1*s1 + w2*s2 + w3*s3. For each adjustment of the change parameters, calculate the change rate between the two similarity scores before and after the adjustment. When the change rate is less than the preset second threshold, stop adjusting the change parameters, and use the finally calculated similarity score as the final similarity score. The above method adjusts the weights of different similarity algorithms through change parameters, and can automatically optimize these weights according to the data, thereby improving the accuracy of similarity calculation.

[0060] In a specific embodiment, a training set and a test set are generated, which specifically include the following steps:

[0061] Identify all security events to be recognized, set event labels for each security event, and determine whether a security event has occurred for each modal data. If so, set the event label of the corresponding security event for the modal data; if not, set the default label for the modal data. Generate an event label vector for each fused modal data, where each data in the event label vector corresponds to the meaning of the security event type. Determine whether a security event has occurred for each modal data in the fused modal data. If so, set the corresponding data in the event label vector to the corresponding event label; if not, set the corresponding data in the event label vector to the default label. Use all the fused modal data marked with event labels as the comprehensive data set, and divide the comprehensive data set into a training set and a test set.

[0062] Specifically, clarify the types of security events to be warned about, and assign a unique label to each security event for subsequent classification and identification. For example, the types of security events include fire, intrusion, equipment failure, gas leakage, and no security event occurring. Use numbers to assign unique event labels to each security event. For example, fire is set to 1, intrusion is set to 2, equipment failure is set to 3, gas leakage is set to 4, and no security event is set to 0. For each modal data such as sensor data (temperature 35°C, smoke concentration 50 ppm), set the corresponding time label for the sensor data to 1. For Internet of Things device data (door lock status: closed, current: normal), it is judged that no security event has occurred, and the default label is set to 0. Generate an event label vector for the fused modal data. For example, if there are a total of four types of security events, the event label vector of the fused modal data is defaulted to [0, 0, 0, 0]. Each data in the time label vector corresponds one by one to each type of security event. For example, when the first digit is 0, it means that no fire has occurred, and if the first digit is 1, it means that a fire has occurred. Judge whether a security event has occurred for each modal data in the fused modal data. If so, set the corresponding event label to the corresponding time label. If not, set the corresponding data to 0. Use all the fused modal data marked with time labels as the comprehensive data set, and divide the comprehensive data set into a training set and a test set according to a preset ratio. For example, 80% of the data can be used as the training set, and 20% of the data can be used as the test set.

[0063] In a specific embodiment, train a security warning model based on the training set, which specifically includes the following steps:

[0064] Individually train each modal data in the fused modal data to generate an individual warning model. During the process of generating the individual warning model, at every preset first time period, verify the accuracy of the individual warning model based on the test set. If the corresponding accuracy does not increase within a continuous number of training cycles, end the training of this individual warning model. Train the fused modal data to generate a comprehensive warning model. The training of the comprehensive warning model and the training of each individual warning model are carried out in parallel. Fuse the generated multiple individual warning models to generate a fused warning model. Compare the accuracy of the warning results of the fused warning model and the comprehensive warning model, and use the model with the higher accuracy as the security warning model.

[0065] Specifically, for the training problem of multi-modal data, in order to prevent overfitting and improve the warning accuracy, individually train different modal data to generate individual warning models. During the process of generating the individual warning model, at every preset first time period, such as one training cycle, verify the accuracy of the individual warning model based on the test set. If within a continuous number (such as 5) of training cycles, the growth value of the corresponding accuracy is less than the preset second threshold, then end the training of this individual warning model.

[0066] Train a comprehensive security warning model by training the fused modal data. To improve the efficiency of model training, the training of the security warning model is carried out in parallel with the training of each individual warning model. During the process of training the security warning model, at a preset first time period, such as one training cycle, the accuracy of the security warning model is verified based on the test set. If the growth value of the corresponding accuracy is less than the preset second threshold within a number of consecutive (such as 5) training cycles, the training of the security warning model is terminated.

[0067] Fuse the generated multiple individual warning models to generate a fused warning model, enabling the comprehensive warning model to learn useful information from the individual warning models. The fusion method can adopt ways such as parameter sharing, feature fusion, etc. For example, use the parameters of the individual warning models as the initial parameters of the comprehensive warning model, or splice or perform weighted summation on the prediction results of the individual warning models to achieve model fusion. The fused warning model combines the advantages of the individual warning models, can better utilize the complementary information of different modal data, and improve the robustness and accuracy of the model.

[0068] Use the test set to evaluate the fused warning model and the comprehensive warning model, compare the accuracy of their warning results. The evaluation metrics can include accuracy, recall rate, F1 score, etc. Select the model with higher accuracy as the final security warning model. If the accuracies of the two models are not much different, factors such as the complexity and running time of the models can be further considered. By comparing the performances of different models, ensure to select the optimal model as the core of the security warning system to improve the overall performance and reliability of the system.

[0069] The above method controls the training process of each modality separately according to the learning state of each modality, preventing overfitting or underfitting of some modalities. By independently controlling the learning process of each modality, it avoids overtraining of the feature extraction model of some modalities while under-training of other modality models.

[0070] In a specific embodiment, calculate the prediction error based on the warning result, which specifically includes the following steps:

[0071] When the security prediction model is a fusion early warning model, obtain the prediction accuracy of each individual early warning model, use the prediction accuracy as the reliability score of each modality data, calculate the total score of all reliability scores, and use the result of dividing each reliability score by the total score as the error weight corresponding to each modality data. Compare the early warning result of each individual early warning model with the pre-set event label vector, calculate the prediction error of each individual early warning model, and use the result of weighted addition of the error weight and the corresponding prediction error as the prediction error of the security early warning model. When the security early warning model is a comprehensive early warning model, calculate the difference between the prediction result of the comprehensive model and the pre-set event label vector as the prediction error.

[0072] Specifically, the traditional method of calculating errors often performs unified calculations on all modality data without considering the reliability differences of different modality data for the early warning model, thereby affecting the accuracy of error calculation. When the security prediction model is a fusion early warning model, the fusion early warning model is obtained by fusing multiple individual early warning models and can effectively reflect the reliability of each modality data in actual applications. Therefore, use the prediction accuracy of each individual early warning model as the reliability score corresponding to each modality data, calculate the total score of all reliability scores, and use the result of dividing each reliability score by the total score as the error weight corresponding to each modality data. Through the error weight, the contribution of each modality data in the fusion early warning model can be dynamically adjusted. Perform weighted addition on the error weight and the prediction error corresponding to each individual early warning model, and use the obtained result as the prediction error of the fusion early warning model. Through weighted addition, ensure that the prediction error of each modality data is adjusted according to its reliability, so that the final prediction error can more accurately reflect the overall performance of the model.

[0073] The comprehensive early warning model is obtained by learning and fusing modality data. The prediction error of the comprehensive early warning model can be obtained by calculating the difference degree between the prediction result and the true security event label.

[0074] In a specific embodiment, update the parameters of the security early warning model according to the prediction error, which specifically includes the following steps:

[0075] Backpropagate the calculated prediction error to the security early warning model, calculate the gradient of the security early warning model parameters, and use an optimization algorithm to update the parameters of the security early warning model. Repeat this step until the prediction error of the security early warning model is less than or equal to the first threshold.

[0076] Specifically, the back-propagation algorithm can automatically calculate the gradient, provide precise direction and size for parameter updates, and ensure that the model can effectively learn the training data. The optimization algorithm (such as stochastic gradient descent, Adam, or RMSprop) can dynamically adjust the learning rate according to the gradient, accelerate convergence, and improve the training efficiency and stability of the model. By setting the threshold, the training accuracy of the model can be controlled to ensure that the model stops training after reaching the predetermined performance, avoiding over-training and waste of resources.

[0077] The above describes a security warning method based on multimodal data fusion in an embodiment of the present application. The following describes a security warning system based on multimodal data fusion in an embodiment of the present application. Figure 4 In an embodiment of the present application, a safety warning system based on multimodal data fusion includes:

[0078] The data collection module is used to deploy data collection equipment in the target environment to collect multimodal data of the target environment. The data collection equipment includes a variety of sensors, IoT devices, shooting equipment and audio collection equipment. The multimodal data includes sensor data, IoT device data, video data and audio data. According to the characteristics of the collected data and the early warning requirements, the data collection frequency is set for different data collection equipment, and the collected multimodal data is preprocessed;

[0079] A feature extraction module is used to extract state change features of sensor data using a time series analysis model, extract device state features of IoT devices from IoT device data using a machine learning algorithm, extract video features of video frames using a pre-trained convolutional neural network, and extract audio features of audio data using a deep learning model. State change features, device state features, video features, and audio features are collectively referred to as modal features.

[0080] The data fusion module calculates the correlation between different modal features, calculates the corresponding attention weights based on the correlation scores using the Sigmoid function, and introduces a gating mechanism to adjust the attention weights in combination with historical data. The adjusted attention weights are multiplied by each modal data to obtain weighted modal data, and the weighted modal data are concatenated to obtain fused modal data.

[0081] The model training module generates training sets and test sets based on the fused modal data, trains the safety warning model based on the training set, inputs the test set into the safety warning model to output the warning result, calculates the prediction error based on the warning result, and updates the parameters of the safety warning model according to the prediction error.

[0082] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the security warning method based on multimodal data fusion.

[0083] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0084] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., which can store program codes.

[0085] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A safety warning method based on multimodal data fusion, characterized in that: The method comprises: Step S1: deploying data acquisition equipment in the target environment to collect multimodal data of the target environment. The data acquisition equipment includes a variety of sensors, IoT devices, shooting equipment and audio acquisition equipment. The multimodal data includes sensor data, IoT device data, video data and audio data. According to the characteristics of the collected data and the early warning requirements, the data collection frequency is set for different data collection equipment, and the collected multimodal data is preprocessed; Step S2: using a time series analysis model to extract state change features of sensor data, using a machine learning algorithm to extract device state features of IoT devices from IoT device data, using a pre-trained convolutional neural network to extract video features of video frames, and using a deep learning model to extract audio features of audio data. State change features, device state features, video features, and audio features are collectively referred to as modal features. Step S3, calculate the correlation between different modal features, calculate the corresponding attention weights based on the correlation scores using the Sigmoid function, introduce a gating mechanism to adjust the attention weights in combination with historical data, multiply the adjusted attention weights with each modal data to obtain weighted modal data, and concatenate the weighted modal data to obtain fused modal data; Step S4: Generate a training set and a test set based on the fused modal data, train the safety warning model based on the training set, input the test set into the safety warning model to output the warning result, calculate the prediction error based on the warning result, and update the parameters of the safety warning model according to the prediction error.

2. The method according to claim 1, characterized in that Calculate the correlation between different modal features, including: Sensor data, IoT device data, video data, and audio data are aligned in the time dimension based on timestamps. Different modal features are mapped to the same dimension using linear transformation. Nonlinear activation functions are used to introduce nonlinear transformation into the modal features after linear transformation. The modal features after nonlinear transformation are normalized. The similarity scores between different modal features are calculated. The calculated similarity scores are normalized and the normalized similarity scores are used as the correlation between the corresponding different modal data.

3. The method according to claim 2, characterized in that Calculate similarity scores between features of different modalities, including: A plurality of different similarity algorithms are used to calculate the similarity between each two different modal features to obtain a plurality of different basic similarity scores. For each two different modal features, the sum of the absolute values ​​of the corresponding multiple basic similarity scores is calculated. The absolute value of each basic similarity score is divided by the sum of the absolute values ​​to obtain the initial weight factor of each similarity algorithm. A variable parameter is introduced for each similarity algorithm. The weight factors of different similarity algorithms are adjusted based on the variable parameters. The adjusted weight factors and the corresponding basic similarity scores are weighted to obtain a mixed similarity score. The mixed similarity score is used as the similarity score between different modal features.

4. The method according to claim 3, characterized in that Adjust the weight factors of different similarity algorithms based on the changing parameters, including: Initialize the change parameters, calculate the adjusted weight factors based on the change parameters using the softmax function, calculate the corresponding similarity scores based on the adjusted weight factors, adjust each change parameter, calculate the similarity scores after adjusting the change parameters, calculate the change rate between the similarity scores before and after the adjustment each time the change parameters are adjusted, stop adjusting the change parameters when the change rate is less than a preset second threshold, and use the last calculated similarity score as the final similarity score.

5. The method according to claim 1, characterized in that: Generate training and test sets, including: Confirm all security events to be identified, set an event label for each security event, and determine whether a security event has occurred for each modal data. If so, set the event label of the corresponding security event for the modal data; if not, set a default label for the modal data. Generate an event label vector for each fused modal data. Each data in the event label vector corresponds to the meaning of the security event type. Determine whether a security event has occurred for each modal data in the fused modal data. If so, set the corresponding data in the event label vector to the corresponding event label; if not, set the corresponding data in the event label vector to the default label. Use all fused modal data marked with event labels as a comprehensive data set, and divide the comprehensive data set into a training set and a test set.

6. The method according to claim 1, characterized in that The security early warning model is trained based on the training set, including: Each modal data in the fused modal data is trained separately to generate a separate warning model. In the process of generating the separate warning model, the accuracy of the separate warning model is verified based on the test set every preset first time period. If the corresponding accuracy does not improve within several consecutive training cycles, the training of the separate warning model is terminated, and the fused modal data is trained to generate a comprehensive warning model. The training of the comprehensive warning model and the training of each separate warning model are carried out in parallel, and the multiple generated separate warning models are fused to generate a fused warning model. The accuracy of the warning results of the fused warning model and the comprehensive warning model are compared, and the model corresponding to the greater accuracy is used as the safety warning model.

7. The method according to claim 1, characterized in that Calculate the prediction error based on the warning results, including: When the safety prediction model is a fusion warning model, the prediction accuracy of each individual warning model is obtained, and the prediction accuracy is used as the reliability score of each modal data. The total score of all reliability scores is calculated, and the result obtained by dividing each reliability score by the total score is used as the error weight corresponding to each modal data. The warning result of each individual warning model is compared with the preset event label vector, and the prediction error of each individual warning model is calculated. The result obtained by weighted addition of the error weight and the corresponding prediction error is used as the prediction error of the safety warning model. When the safety warning model is a comprehensive warning model, the difference between the prediction result of the comprehensive model and the preset event label vector is calculated as the prediction error.

8. The method according to claim 1, characterized in that: Update the parameters of the safety warning model according to the prediction error, including: The calculated prediction error is back-propagated to the safety warning model, the gradient of the safety warning model parameters is calculated, the parameters of the safety warning model are updated using the optimization algorithm, and this step is repeated until the prediction error of the safety warning model is less than or equal to the first threshold.

9. A safety warning system based on multimodal data fusion, used to implement the safety warning method based on multimodal data fusion as described in any one of claims 1 to 8, characterized in that: The system comprises: The data collection module is used to deploy data collection equipment in the target environment to collect multimodal data of the target environment. The data collection equipment includes a variety of sensors, IoT devices, shooting equipment and audio collection equipment. The multimodal data includes sensor data, IoT device data, video data and audio data. According to the characteristics of the collected data and the early warning requirements, the data collection frequency is set for different data collection equipment, and the collected multimodal data is preprocessed; A feature extraction module is used to extract state change features of sensor data using a time series analysis model, extract device state features of IoT devices from IoT device data using a machine learning algorithm, extract video features of video frames using a pre-trained convolutional neural network, and extract audio features of audio data using a deep learning model. State change features, device state features, video features, and audio features are collectively referred to as modal features. The data fusion module calculates the correlation between different modal features, calculates the corresponding attention weights based on the correlation scores using the Sigmoid function, and introduces a gating mechanism to adjust the attention weights in combination with historical data. The adjusted attention weights are multiplied by each modal data to obtain weighted modal data, and the weighted modal data are concatenated to obtain fused modal data. The model training module generates training sets and test sets based on the fused modal data, trains the safety warning model based on the training set, inputs the test set into the safety warning model to output the warning result, calculates the prediction error based on the warning result, and updates the parameters of the safety warning model according to the prediction error.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instruction is executed by the processor, the safety warning method based on multimodal data fusion as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Safety early warning system based on multi-modal data fusion

    CN116881850A

  • Security risk grading early warning system and method based on multi-modal data

    CN117852867A

Cited By

  • Pedestrian anomaly detection method and device and storage medium

    CN120726675A

  • Pet behavior prediction method and device, equipment and storage medium

    CN120726701A

  • Dynamic self-adaptive operation and maintenance data intelligent prediction system, method and server

    CN120806287A

  • Dynamic updating method and device of biological characteristics, equipment and storage medium

    CN120977024A

  • Well engineering-oriented multi-modal abstract generation method and device and storage medium

    CN121256057A