Adaptive disaster anomaly recognition method and early warning system based on multi-prime convolution kernel
By using the Omni-Scale convolutional neural network with multi-prime convolution kernel and the enhanced self-attention mechanism in disaster anomaly detection, combined with the mean teacher model of dynamic smoothing coefficient, the problem of difficulty in receptive field selection in the existing technology is solved, and the accuracy and generalization ability of disaster anomaly recognition are improved.
Patent Information
- Application Number
- CN202510331414.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-20
AI Technical Summary
When the existing disaster anomaly detection methods face complex dimensional relationship characteristics in the time series, they cannot effectively select the receptive field, resulting in large calculation volume and high parameter adjustment complexity, which cannot meet the performance requirements of disaster warning.
The Omni-Scale convolutional neural network block based on multi-prime convolution kernel and the enhanced self-attention mechanism are used, combined with the mean teacher model with dynamic smoothing coefficients, and semi-supervised learning is performed through no more than 10% of the label data, and the receptive field size is dynamically adjusted.
It improves the accuracy and generalization ability of disaster anomaly recognition, reduces the model training time and calculation complexity, and enhances the accuracy of abnormal monitoring and classification.
Smart Images

Figure CN119848747B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disaster warning technology, and in particular to an adaptive disaster anomaly recognition method and warning system based on multi-prime convolution kernels. Background Art
[0002] Faced with the huge amount of data in the massive, high-dimensional, and high-noise time series data of the disaster emergency Internet of Things, as well as the complex dependencies between multiple disaster emergency monitoring variables, the existing disaster anomaly detection methods have the following shortcomings: (1) As the most popular and common anomaly detection methods, they are all based on deep learning technology. However, in order to learn effective abstract features, traditional deep learning methods need to learn a large number of parameters in advance, which consumes a lot of time and computation, and the model is prone to overfitting; manually adding high-quality labels to the training data is usually expensive. (2) In order to effectively extract data features, traditional feature extraction methods need to select the appropriate receptive field size for different data sets, which greatly increases the system's computational workload and parameter adjustment complexity, and cannot meet the performance requirements of disaster warning.
[0003] Therefore, an adaptive disaster anomaly recognition method and early warning system based on multi-prime convolution kernels was developed to solve the above problems. Summary of the invention
[0004] The present invention proposes an adaptive disaster anomaly recognition method and early warning system based on multi-prime convolution kernels to solve the problem that the existing disaster anomaly detection method cannot effectively select the receptive field when facing the complex dimensional relationship characteristics in the time series.
[0005] The present invention achieves the above-mentioned purpose through the following technical solutions:
[0006] The present invention provides an adaptive disaster anomaly recognition method based on multi-prime convolution kernels, comprising:
[0007] Acquiring first information and second information, wherein the first information includes historical data of geological disasters, forest fires, floods, and earthquakes, and the second information includes real-time on-site monitoring data of geological disasters, forest fires, floods, and earthquakes;
[0008] After labeling part of the data in the first information, a training set is obtained, wherein the part of the data accounts for no more than 10% of the first information;
[0009] Constructing a model based on the mean teacher method, wherein the model based on the mean teacher method includes a student model and a teacher model, and the data entering the student model and the teacher model are sequentially processed by an Omni-Scale convolutional neural network block based on a multi-prime convolution kernel and an enhanced self-attention mechanism based on a convolution operation;
[0010] The model based on the mean teacher method is trained by the training set based on the dynamic smoothing coefficient to obtain a disaster anomaly recognition model;
[0011] The second information is input into the disaster anomaly recognition model to obtain an anomaly recognition result.
[0012] Specifically, obtaining the first information and the second information includes:
[0013] Obtain original historical data on geological disasters, forest fires, floods, and earthquakes;
[0014] The format of the original historical data is standardized to obtain standardized historical data of geological disasters, forest fires, floods, and earthquakes;
[0015] The standardized historical data of multiple dimensions of geological disasters, forest fires, floods, and earthquakes are integrated into multivariate time series data to obtain the first information;
[0016] Obtain real-time collection of original on-site monitoring data of geological disasters, forest fires, floods, and earthquakes;
[0017] The format of the original on-site monitoring data is standardized to obtain the standardized original on-site monitoring data of geological disasters, forest fires, floods, and earthquakes;
[0018] The second information is obtained by integrating the original on-site monitoring data of multiple dimensions of geological disasters, forest fires, floods, and earthquakes into multivariate time series data.
[0019] Specifically, the data entering the student model and the teacher model are processed in turn by an Omni-Scale convolutional neural network block based on a multi-prime convolution kernel and an enhanced self-attention mechanism based on a convolution operation, including:
[0020] Performing a convolution operation of an Omni-Scale neural network on the data to obtain a concatenation of multiple layers of feature maps;
[0021] Performing average pooling processing on the concatenation of the multiple layers of feature maps to obtain an output;
[0022] Calculate the query, key and value based on the output respectively based on the one-dimensional convolution;
[0023] Calculating an attention weight matrix based on the query and the key;
[0024] Calculate the influence of the attention mechanism based on the output, the attention weight matrix, the value and the preset feature map weight;
[0025] Perform layer normalization on the influence of the attention mechanism.
[0026] Furthermore, the data is subjected to a convolution operation of an Omni-Scale neural network to obtain a concatenation of multiple layers of feature maps, including:
[0027]
[0028] in, Is the convolution kernel size used The generated feature map, Represents a one-dimensional Omni-Scale convolution operation;
[0029]
[0030] Represents the concatenation of multiple layers of feature maps.
[0031] Further, the query, key and value are calculated based on the one-dimensional convolution according to the output, including:
[0032]
[0033] in, , , Represents the convolution kernel size of Query, Key and Value respectively, Represents a one-dimensional convolution operation.
[0034] Further, an attention weight matrix is calculated according to the query and the key, comprising:
[0035]
[0036] in, It is the dimension of Key.
[0037] Further, the influence of the attention mechanism is calculated according to the output, the attention weight matrix, the value and the preset feature mapping weight, including:
[0038]
[0039] Among them, γ represents the weight of the feature map after the self-attention mechanism is applied, which is calculated by the back propagation algorithm. The back propagation algorithm adopts the BP back propagation algorithm and directly calls the bp algorithm to generate γ.
[0040] Further, training the model based on the mean teacher method based on a dynamic smoothing coefficient according to the first information includes:
[0041] Inputting the annotated data in the first information into the student model to obtain a first output;
[0042] Calculate the cross entropy loss between the first output and the label value of the labeled data;
[0043] Inputting the unlabeled data in the first information into the student model and the teacher model at the same time to obtain a second output and a third output;
[0044] Calculating an average error loss of the second output and the third output;
[0045] The cross entropy loss and the average error loss are weighted and summed to obtain the final mixed loss. The weight coefficient of the weighted sum is used as part of the model parameters. The acquisition and update of the weight coefficient are the same as the model parameter update method: input to the Adam optimizer to obtain the gradient of the model parameters, and then update the parameter value through the back propagation algorithm;
[0046] Optimizing and updating the parameters of the student model according to the final loss;
[0047] Optimizing and updating the parameters of the teacher model according to the optimized and updated parameters of the student model and the dynamic smoothing coefficient;
[0048] The parameters of the optimized student model and teacher model are input into the Adam optimizer to obtain the gradient of the model parameters, and then the parameters of the entire model are updated through the back propagation algorithm;
[0049] When the number of training times reaches the preset number, the model training ends and the model parameters obtained from the last round of training are saved.
[0050] Furthermore, the formula for optimizing and updating the parameters of the teacher model according to the parameters optimized and updated by the student model and the dynamic smoothing coefficient is:
[0051]
[0052] represents the optimization parameters of the teacher model at time t, represents the updated parameters of the student model at time t, Represents the memory parameter of the teacher model at time t-1 , Represents the dynamic smoothing coefficient.
[0053] Furthermore, the calculation formula of the dynamic smoothing coefficient is:
[0054]
[0055] in epochThe batches of data for training.
[0056] The present invention also provides an early warning system for any of the above-mentioned adaptive disaster anomaly identification methods based on multiple prime number convolution kernels, comprising:
[0057] An acquisition module, the acquisition module is used to acquire first information and second information, the first information includes historical data of geological disasters, forest fires, floods, and earthquakes, and the second information includes real-time on-site monitoring data of geological disasters, forest fires, floods, and earthquakes;
[0058] a labeling module, wherein the labeling module is used to label part of the data in the first information to obtain a training set, wherein the part of the data accounts for no more than 10% of the first information;
[0059] A construction module, wherein the construction module is used to construct a model based on the mean teacher method, wherein the model based on the meanteacher method includes a student model and a teacher model, and the data entering the student model and the teacher model are sequentially processed by an Omni-Scale convolutional neural network block based on a multi-prime convolution kernel and an enhanced self-attention mechanism based on a convolution operation;
[0060] A training module, wherein the training module is used to train the model based on the mean teacher method based on the training set based on the dynamic smoothing coefficient to obtain a disaster anomaly recognition model;
[0061] An anomaly identification module is used to input the second information into the disaster anomaly identification model to obtain an anomaly identification result.
[0062] The beneficial effects of the present invention are:
[0063] The present invention proposes an adaptive disaster anomaly recognition method and early warning system based on multi-prime convolution kernels, which is based on a mean teacher model with a dynamic smoothing coefficient, a convolutional neural network with multi-prime convolution kernels, and an adaptive self-attention mechanism. The semi-supervised learning of the data is combined with no more than 10% of labels, which effectively improves the accuracy and generalization ability of the system in identifying anomalies, solves the problem of being unable to effectively select the receptive field when the dimensional relationship characteristics in the time series are complex, and improves the accuracy of anomaly monitoring and anomaly classification and enhances the generalization ability of the anomaly detection model through the strategy of dynamically adjusting the smoothing coefficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 This is a data processing flow chart of an adaptive disaster anomaly recognition method based on multi-prime convolution kernels in an embodiment of the present invention;
[0065] Figure 2Schematic diagram of the model training process based on the mean teacher method in an embodiment of the present invention;
[0066] Figure 3 A schematic diagram of the model architecture of the student model and the teacher model described in an embodiment of the present invention;
[0067] Figure 4 This is a schematic diagram of data collection in an embodiment of the present application. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0069] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0071] In the description of the present invention, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, or are directions or positional relationships in which the product of the invention is usually placed when in use, or are directions or positional relationships commonly understood by those skilled in the art. These directions or positional relationships are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.
[0072] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and should not be understood as indicating or implying relative importance.
[0073] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly stipulated and limited, terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal connection of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. The term "Omni-Scale Convolutional Neural Network" refers to a full-scale convolutional neural network.
[0074] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.
[0075] like Figure 1 As shown, the present invention provides an adaptive disaster anomaly recognition method based on a multi-prime convolution kernel, comprising:
[0076] Acquiring first information and second information, wherein the first information includes historical data of geological disasters, forest fires, floods, and earthquakes, and the second information includes real-time on-site monitoring data of geological disasters, forest fires, floods, and earthquakes;
[0077] After labeling part of the data in the first information, a training set is obtained, wherein the part of the data accounts for no more than 10% of the first information;
[0078] Constructing a model based on the mean teacher method, wherein the model based on the mean teacher method includes a student model and a teacher model, and the data entering the student model and the teacher model are sequentially processed by an Omni-Scale convolutional neural network block based on a multi-prime convolution kernel and an enhanced self-attention mechanism based on a convolution operation;
[0079] The model based on the mean teacher method is trained by the training set based on the dynamic smoothing coefficient to obtain a disaster anomaly recognition model;
[0080] The second information is input into the disaster anomaly recognition model to obtain an anomaly recognition result. The anomaly recognition result includes a disaster recognition category and an anomaly score.
[0081] In some embodiments, obtaining the first information and the second information includes:
[0082] Obtain original historical data on geological disasters, forest fires, floods, and earthquakes;
[0083] The format of the original historical data is standardized to obtain standardized historical data of geological disasters, forest fires, floods, and earthquakes;
[0084] The standardized historical data of multiple dimensions of geological disasters, forest fires, floods, and earthquakes are integrated into multivariate time series data to obtain the first information;
[0085] Obtain real-time collection of original on-site monitoring data of geological disasters, forest fires, floods, and earthquakes;
[0086] The format of the original on-site monitoring data is standardized to obtain the standardized original on-site monitoring data of geological disasters, forest fires, floods, and earthquakes;
[0087] The second information is obtained by integrating the original on-site monitoring data of multiple dimensions of geological disasters, forest fires, floods, and earthquakes into multivariate time series data.
[0088] In this embodiment, the original historical data is the time series data set UCR, which covers many fields and contains 128 types of industry data, among which the distribution, dimension and proportion of anomalies are different. The data is diverse, ensuring the generalization ability of the model and the wide application.
[0089] like Figure 4 As shown, in this embodiment, the original on-site monitoring data specifically includes:
[0090] Landslide monitoring data: displacement sensor data, video monitoring sensor data, rain gauge data, soil moisture meter data in the landslide monitoring area;
[0091] Earthquake monitoring data: seismic wave sensor data, acceleration sensor data, and pressure sensor data in the earthquake monitoring area;
[0092] Flood monitoring data: water level sensor data, video monitoring sensor data, water depth sensor data, and rainfall sensor data in the flood monitoring area;
[0093] Forest fire monitoring data: wind direction and speed sensor data, far-infrared sensor data, video monitoring sensor data, and air temperature and humidity sensor data in the forest fire monitoring area.
[0094] After the above original on-site monitoring data is transmitted to the base station, the base station transmits the original on-site monitoring data to the disaster anomaly recognition model in the present invention.
[0095] Combination Figure 1 and Figure 3 As shown, in some embodiments, the data entering the student model and the teacher model are sequentially passed through the Omni-Scale convolutional neural network block based on multiple prime convolution kernels (i.e. Figure 3 OS_CNN Block in ), enhanced self-attention mechanism based on convolution operation (i.e. Figure 3Conv_SelfAt) in, including:
[0096] Performing a convolution operation of an Omni-Scale neural network on the data to obtain a concatenation of multiple layers of feature maps;
[0097] Performing average pooling processing on the concatenation of the multiple layers of feature maps to obtain an output;
[0098] Calculate the query, key and value based on the output respectively based on the one-dimensional convolution;
[0099] Calculating an attention weight matrix based on the query and the key;
[0100] Calculate the influence of the attention mechanism based on the output, the attention weight matrix, the value and the preset feature map weight;
[0101] Perform layer normalization on the influence of the attention mechanism.
[0102] The Omni-Scale convolutional neural network in the present invention uses multiple prime numbers as the size of the convolution kernel to capture multi-scale features. For different data sets, time dependencies of different lengths can be covered by prime convolution kernels, thereby simplifying the difficulty of selecting the size of the receptive field, while achieving better performance and solving the problem of how to select the size of the receptive field in the classification task of CNN in time series data. An enhanced self-attention mechanism based on convolution operation is designed, which introduces convolution operation instead of linear transformation, and combines learnable parameter γ and layer normalization to improve the model's ability to capture local information of time series data.
[0103] like Figure 3 As shown, in some embodiments, the size of the convolution kernel in OS_CNN Block is selected from a prime number set , where each is a prime number (e.g., 1, 2, 3, 5, 7, etc., 1 is reserved). The data is convolved by an Omni-Scale neural network to obtain a concatenation of multiple layers of feature maps, including:
[0104]
[0105] in, Is the convolution kernel size used The generated feature map, Represents a one-dimensional Omni-Scale convolution operation;
[0106] The output of OS_CNN Block is the concatenation of multiple layers of feature maps.
[0107]
[0108] Represents the concatenation of multiple layers of feature maps.
[0109] Output Y after average pooling:
[0110]
[0111] In some embodiments, respectively calculating the query, the key, and the value based on the one-dimensional convolution according to the output includes:
[0112]
[0113] in, , , Represents the convolution kernel size of Query, Key and Value respectively, Represents a one-dimensional convolution operation.
[0114] In some embodiments, calculating an attention weight matrix based on the query and the key includes:
[0115]
[0116] in, It is the dimension of Key.
[0117] In some embodiments, calculating the influence of the attention mechanism according to the output, the attention weight matrix, the value and the preset feature map weight includes:
[0118]
[0119] Among them, γ represents the weight of the feature map after the self-attention mechanism is applied, which is calculated by the back propagation algorithm. The back propagation algorithm adopts the BP back propagation algorithm and directly calls the BP back propagation algorithm to generate .
[0120] To flexibly control the influence of the attention mechanism.
[0121] like Figure 3 As shown, The input multivariate time series monitoring data is input into OS_CNN block and Conv_SelfAt respectively. After the output, the sequence is concatenated and input into OS_CNN block again for full connection operation. Finally, the abnormal classification result is obtained as the output.
[0122] like Figure 2As shown, in some embodiments, training the model based on the mean teacher method based on the dynamic smoothing coefficient according to the first information includes:
[0123] Inputting the annotated data in the first information into the student model to obtain a first output;
[0124] Calculate the cross entropy loss between the first output and the label value of the labeled data;
[0125] Inputting the unlabeled data in the first information into the student model and the teacher model at the same time to obtain a second output and a third output;
[0126] Calculating an average error loss of the second output and the third output;
[0127] The cross entropy loss and the average error loss are weighted and summed to obtain the final loss of the mixture. Generally, the meanteacher model sets this weight value to 0.1 or 0.2. According to this application scenario, the weight value is set to 0.2 in this embodiment;
[0128] Optimizing and updating the parameters of the student model according to the final loss;
[0129] Optimizing and updating the parameters of the teacher model according to the optimized and updated parameters of the student model and the dynamic smoothing coefficient;
[0130] The parameters of the optimized student model and teacher model are input into the Adam optimizer to obtain the gradient of the model parameters, and then the parameters of the entire model are updated through the back propagation algorithm;
[0131] When the number of training times reaches the preset number, the model training ends and the model parameters obtained from the last round of training are saved.
[0132] like Figure 1 As shown, in some embodiments, the formula for optimizing and updating the parameters of the teacher model according to the parameters after optimization and update of the student model and the dynamic smoothing coefficient is:
[0133]
[0134] represents the optimization parameters of the teacher model at time t, represents the updated parameters of the student model at time t, Represents the memory parameter of the teacher model at time t-1 , Represents the dynamic smoothing coefficient.
[0135] epoch is the batch of data training.
[0136] (1.1)
[0137] (1.2)
[0138] (1.3)
[0139] (1.4)
[0140] in, To input time series data, It represents the result after the labeled data is input into the student model. They represent the results of unlabeled data input into the student model and the teacher model respectively. represents the cross entropy loss calculation, express The result of calculating the cross entropy loss with the label Lable of the labeled data is shown in formula (1.1). MSE stands for mean square error calculation. express The result of calculating the mean square error loss is shown in formula (1.2). yes and The final loss obtained by weighted summation is used to optimize the parameters of the student model, as shown in formula (1.3). represents the optimization parameters of the teacher model at time t, represents the updated parameters of the student model at time t, By the parameters of the student model and the memory parameter of the teacher model at time t-1 The weighted sum is obtained as shown in formula (1.4). The smoothing coefficient α is calculated by formula (1.5):
[0141] (1.5)
[0142] in epoch The traditional mean teacher method relies on a fixed smoothing coefficient α to control the degree of parameter update of the teacher model and the student model. However, the fixed α lacks flexibility in the training process and cannot fully utilize the collaborative learning of the student model and the teacher model. The strategy of dynamically adjusting the smoothing coefficient α proposed in the present invention enhances the collaborative update capability of the student model and the teacher model and the generalization capability of the system.
[0143] like Figure 2 As shown, the training process of the model based on the mean teacher method in the present invention is as follows:
[0144] Input the training data in the dataset into the model, and take no more than 10% of the data for labeling (whether it is abnormal, what kind of abnormality); set the number of model training times N and start training; for each training, optimize the parameters of the student model and the teacher model according to the "final loss value" of the training; when the number of training times reaches N, end the model training, and the parameters of the entire model are now optimal.
[0145] The input data enters in sequence: Omni-Scale convolutional neural network block OS_CNNblock based on multi-prime convolution kernels, and enhanced self-attention mechanism algorithm based on convolution operation.
[0146] If it is labeled data (with labels), it is input into the student model, and according to formula (1.1), the output result is calculated with the label value to obtain the cross entropy loss; if it is unlabeled data (without labels), it is input into the student model and the teacher model at the same time, and according to formula (1.2), the output of the student model and the teacher model is calculated to obtain the average error loss.
[0147] According to formula (1.3), the "final loss value" of this training is calculated and the parameters of the student model are optimized.
[0148] According to formula (1.4), optimize the parameters of the teacher model; input the optimized parameters of the student model and the teacher model into the Adam optimizer to obtain the gradient of the model parameters, and then update the parameters of the entire model through the back propagation algorithm.
[0149] When the number of training times reaches N, the model training is terminated, and the model parameters obtained from the last round of training are saved to obtain a disaster anomaly recognition model.
[0150] Then, the test data is input to test the performance of the disaster anomaly recognition model. First, the test data is input into the trained disaster anomaly recognition model, and then the output of the disaster anomaly recognition model is compared with the anomaly classification results of the test data to evaluate the model performance.
[0151] The performance comparison of the disaster anomaly recognition model uses several main performance indicators based on the confusion matrix: precision, recall, and F1-score. For these three metrics, higher values indicate better performance.
[0152] Precision (Pre) refers to the proportion of samples predicted correctly in the prediction results. There are two types of results predicted as positive examples: either positive examples TP or negative examples FP. The formula is:
[0153]
[0154] Recall (Rec) refers to the ratio of the predicted correct positive examples to the total positive examples among the actual positive examples, which are based on the actual samples. Among the actual positive examples, the samples are either predicted correctly in the prediction (TP) or predicted incorrectly in the prediction (FN). It can be expressed as:
[0155]
[0156] The F1 value is the harmonic mean of precision and recall, and the calculation formula is:
[0157]
[0158] Specifically, in the time series data set, the precision MacroPre, recall MacroRec, and MacroF1 represent the arithmetic mean of each statistical indicator value of all categories.
[0159]
[0160]
[0161]
[0162] The experimental results of the adaptive disaster anomaly recognition method based on multi-prime convolution kernels (hereinafter referred to as OSS_CNN_MT) proposed in this invention on the public dataset UCR are shown in Tables 1, 2 and 3 below:
[0163] Table 1
[0164]
[0165] Table 2
[0166]
[0167] Table 3
[0168]
[0169] In Tables 1, 2, and 3, the method is compared with 6 baseline multi-classification anomaly detection methods. The method obtained 19 best results on 34 industry datasets, which is 10% higher than the second best method.
[0170] The present invention is applied to a multi-disaster disaster monitoring and early warning system in the field of emergency disasters, such as Figure 4 As shown in the figure, the monitoring areas of four typical disasters (geological disasters, forest fires, floods, and earthquakes) are displayed. The implementation plan is as follows:
[0171] 1. Deploy an adaptive disaster anomaly warning system: deploy the disaster anomaly recognition model of the present invention in the data center, and ensure that the data center is connected to the sensors on the monitoring site;
[0172] 2. Preprocess historical data and use historical data to train model parameters:
[0173] (1) Standardize the format of historical data on geological disasters, forest fires, floods, and earthquakes;
[0174] (2) Merge historical data from multiple dimensions into multivariate time series data, of which no more than 10% of the data is labeled (whether it is abnormal and what kind of abnormality it is);
[0175] (3) Using this historical multivariate time series data as the input of the disaster anomaly recognition model, and training the parameters of the disaster anomaly recognition model;
[0176] 3. Real-time monitoring data preprocessing:
[0177] (1) Collect on-site monitoring data of geological disasters, forest fires, floods, and earthquakes in real time, send them to the data center, and standardize the format of the real-time data;
[0178] (2) Preprocess the real-time collected data and fuse the monitoring data of multiple dimensions into multivariate time series data;
[0179] (3) Using this real-time multivariate time series data as input to the trained model for calculation;
[0180] 4. Real-time monitoring of data abnormal event identification:
[0181] (1) Use end-to-end classification to identify disaster types (geological disasters, forest fires, floods, earthquakes);
[0182] (2) Use the anomaly score to calculate the anomaly level, generate and send warning information (ID, monitoring point, disaster type, time, risk level, etc.) and divide the anomaly level according to the anomaly score. The higher the anomaly score, the higher the disaster level.
[0183] 5. Repeat 3-4 to identify real-time disaster abnormal events and update model parameters online.
[0184] As the most popular and common anomaly detection methods, they are all based on deep learning technology. However, in order to learn effective abstract features, traditional deep learning methods need to learn a large number of parameters in advance, which consumes a lot of time and computation. At the same time, the model is prone to overfitting. Manually adding high-quality labels to training data is usually expensive. To address this problem, a deep learning method based on mean teacher is used, and a very small amount (no more than 10%) of labeled data is used for semi-supervised learning, which greatly reduces the model training time and improves the accuracy of anomaly detection. However, in the traditional mean teacher method, the parameter update between the teacher model and the student model depends on a fixed smoothing coefficient α. After research, it was found that the fixed α coefficient may not be able to optimally control the parameter update between the teacher model and the student model at different stages of training.
[0185] In order to effectively extract data features, traditional feature extraction methods need to select appropriate receptive field sizes for different data sets, which greatly increases the system's computational workload and parameter adjustment complexity, and cannot meet the performance requirements of disaster warning. In response to this technical problem, in order to reduce model training time and improve anomaly detection accuracy, this technology is based on the mean teacher method that dynamically adjusts the smoothing coefficient α, combined with no more than 10% of labeled data for semi-supervised training and learning. In the early stages of training, the system relies more on the update of the student model, and gradually turns to the update of the teacher model in the later stages, thus enhancing the generalization ability of the entire system; in order to effectively identify complex emergency disaster IoT abnormal patterns and features, multiple prime numbers are used as the convolution kernel size, and the Omni-Scale block is designed to adaptively select the receptive field size. These prime convolution kernel sizes can cover the best receptive fields of various data sets without increasing computational complexity.
[0186] In summary, the advantages of the present invention compared to the prior art are:
[0187] (1) It solves the challenge of anomaly detection in multivariate time series data of emergency disasters. (2) It overcomes the shortcomings of traditional deep learning methods and adopts a mean teacher model with dynamic smoothing coefficients. It uses a small number of labels to reduce the amount of data calculation, improve the detection speed, and enhance the generalization ability of the anomaly detection model. (3) It uses an Omni-Scale convolutional neural network with multiple prime convolution kernels to effectively avoid the overlap of the receptive fields of different convolution kernels, thereby reducing information redundancy and improving the model's ability to capture multi-scale information. (4) It adopts an enhanced self-attention mechanism based on convolution operations, which helps to reduce the problem of gradient disappearance or explosion and improve the convergence of model training.
[0188] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An adaptive disaster anomaly recognition method based on multi-prime convolution kernels, characterized in that: include: Acquire first information and second information, wherein the first information includes historical data of geological disasters, forest fires, floods, and earthquakes, and the second information includes real-time on-site monitoring data of geological disasters, forest fires, floods, and earthquakes, and both the first information and the second information include landslide monitoring data, earthquake monitoring data, flood monitoring data, and forest fire monitoring data, wherein the landslide monitoring data includes displacement sensor data, video monitoring sensor data, rain gauge data, and soil moisture meter data in the landslide monitoring area, the earthquake monitoring data includes seismic wave sensor data, acceleration sensor data, and pressure sensor data in the earthquake monitoring area, the flood monitoring data includes water level sensor data, video monitoring sensor data, water depth sensor data, and rainfall sensor data in the flood monitoring area, and the forest fire monitoring data includes wind direction and wind speed sensor data, far infrared sensor data, video monitoring sensor data, and air temperature and humidity sensor data in the forest fire monitoring area; After labeling part of the data in the first information, a training set is obtained, wherein the part of the data accounts for no more than 10% of the first information; Constructing a model based on the mean teacher method, wherein the model based on the mean teacher method includes a student model and a teacher model, and the data entering the student model and the teacher model are sequentially processed by an Omni-Scale convolutional neural network block based on a multi-prime convolution kernel and an enhanced self-attention mechanism based on a convolution operation; The model based on the mean teacher method is trained by the training set based on the dynamic smoothing coefficient to obtain a disaster anomaly recognition model; Inputting the second information into the disaster anomaly recognition model to obtain an anomaly recognition result; Training the model based on the mean teacher method based on a dynamic smoothing coefficient according to the first information includes: Inputting the annotated data in the first information into the student model to obtain a first output; Calculate the cross entropy loss between the first output and the label value of the labeled data; Inputting the unlabeled data in the first information into the student model and the teacher model at the same time to obtain a second output and a third output; Calculating an average error loss of the second output and the third output; The cross entropy loss and the average error loss are weighted summed to obtain the mixed final loss; Optimizing and updating the parameters of the student model according to the final loss; Optimizing and updating the parameters of the teacher model according to the optimized and updated parameters of the student model and the dynamic smoothing coefficient; The parameters of the optimized student model and teacher model are input into the Adam optimizer to obtain the gradient of the model parameters, and then the parameters of the entire model are updated through the back propagation algorithm; When the number of training times reaches the preset number, the model training ends and the model parameters obtained from the last round of training are saved; The formula for optimizing and updating the parameters of the teacher model according to the parameters optimized and updated by the student model and the dynamic smoothing coefficient is: represents the optimization parameters of the teacher model at time t, represents the updated parameters of the student model at time t, Represents the memory parameter of the teacher model at time t-1 , represents the dynamic smoothing coefficient; The calculation formula of the dynamic smoothing coefficient is: in epoch The batches of data for training.
2. According to claim 1, the method for adaptive disaster anomaly recognition based on multi-prime convolution kernels is characterized in that: The first information and the second information are obtained, including: Obtain original historical data on geological disasters, forest fires, floods, and earthquakes; The format of the original historical data is standardized to obtain standardized historical data of geological disasters, forest fires, floods, and earthquakes; The standardized historical data of multiple dimensions of geological disasters, forest fires, floods, and earthquakes are integrated into multivariate time series data to obtain the first information; Obtain real-time collection of original on-site monitoring data of geological disasters, forest fires, floods, and earthquakes; The format of the original on-site monitoring data is standardized to obtain the standardized original on-site monitoring data of geological disasters, forest fires, floods, and earthquakes; The second information is obtained by integrating the original on-site monitoring data of multiple dimensions of geological disasters, forest fires, floods, and earthquakes into multivariate time series data.
3. The method for adaptive disaster anomaly recognition based on multi-prime convolution kernels according to claim 1 is characterized in that: The data entering the student model and the teacher model are processed in turn by the Omni-Scale convolutional neural network block based on multiple prime convolution kernels and the enhanced self-attention mechanism based on convolution operations, including: Performing a convolution operation of an Omni-Scale neural network on the data to obtain a concatenation of multiple layers of feature maps; Performing average pooling processing on the concatenation of the multiple layers of feature maps to obtain an output; Calculate the query, key and value based on the output respectively based on the one-dimensional convolution; Calculating an attention weight matrix based on the query and the key; Calculate the influence of the attention mechanism based on the output, the attention weight matrix, the value and the preset feature map weight; Perform layer normalization on the influence of the attention mechanism.
4. The method for adaptive disaster anomaly recognition based on multi-prime convolution kernels according to claim 3 is characterized in that: The data is convolved by an Omni-Scale neural network to obtain a concatenation of multiple layers of feature maps, including: in, Is the convolution kernel size used The generated feature map, Represents a one-dimensional Omni-Scale convolution operation; Represents the concatenation of multiple layers of feature maps.
5. The method for adaptive disaster anomaly recognition based on multi-prime convolution kernels according to claim 4 is characterized in that: The query, key, and value are calculated based on the one-dimensional convolution according to the output, including: in, , , Represents the convolution kernel size of Query, Key and Value respectively, represents a one-dimensional convolution operation, and Y represents the output after average pooling.
6. The method for adaptive disaster anomaly recognition based on multi-prime convolution kernels according to claim 5 is characterized in that: The influence of the attention mechanism is calculated according to the output, the attention weight matrix, the value and the preset feature map weight, including: Among them, A represents the attention weight matrix, Represents the weight of the feature map after the self-attention mechanism is applied, which is calculated by the back-propagation algorithm.
7. An early warning system for an adaptive disaster anomaly recognition method based on multiple prime convolution kernels as described in any one of claims 1 to 6, characterized in that: include: An acquisition module, the acquisition module is used to acquire first information and second information, the first information includes historical data of geological disasters, forest fires, floods, and earthquakes, and the second information includes real-time on-site monitoring data of geological disasters, forest fires, floods, and earthquakes; a labeling module, wherein the labeling module is used to label part of the data in the first information to obtain a training set, wherein the part of the data accounts for no more than 10% of the first information; A construction module, wherein the construction module is used to construct a model based on the mean teacher method, wherein the model based on the meanteacher method includes a student model and a teacher model, and the data entering the student model and the teacher model are sequentially processed by an Omni-Scale convolutional neural network block based on a multi-prime convolution kernel and an enhanced self-attention mechanism based on a convolution operation; A training module, wherein the training module is used to train the model based on the meanteacher method based on the training set based on the dynamic smoothing coefficient to obtain a disaster anomaly recognition model; an abnormality identification module, the abnormality identification module is used to input the second information into the disaster abnormality identification model to obtain an abnormality identification result; Wherein, training the model based on the mean teacher method based on the dynamic smoothing coefficient according to the first information includes: Inputting the annotated data in the first information into the student model to obtain a first output; Calculate the cross entropy loss between the first output and the label value of the labeled data; Inputting the unlabeled data in the first information into the student model and the teacher model at the same time to obtain a second output and a third output; Calculating an average error loss of the second output and the third output; The cross entropy loss and the average error loss are weighted summed to obtain the mixed final loss; Optimizing and updating the parameters of the student model according to the final loss; Optimizing and updating the parameters of the teacher model according to the optimized and updated parameters of the student model and the dynamic smoothing coefficient; The parameters of the optimized student model and teacher model are input into the Adam optimizer to obtain the gradient of the model parameters, and then the parameters of the entire model are updated through the back propagation algorithm; When the number of training times reaches the preset number, the model training ends and the model parameters obtained from the last round of training are saved; The formula for optimizing and updating the parameters of the teacher model according to the parameters optimized and updated by the student model and the dynamic smoothing coefficient is: represents the optimization parameters of the teacher model at time t, represents the updated parameters of the student model at time t, Represents the memory parameter of the teacher model at time t-1 , represents the dynamic smoothing coefficient; The calculation formula of the dynamic smoothing coefficient is: in epoch The batches of data for training.