Sand dust storm moving path prediction method, device and equipment and storage medium

By combining aerosol optical thickness data and geographical background information, using Otsu segmentation and random forest feature importance to screen key features, a hybrid model of 1DCNN and BiLSTM was constructed, which solved the complex spatial and temporal dependence and long-term memory problems of sandstorm path prediction in the existing technology, and achieved efficient and accurate sandstorm movement path prediction.

CN120375210AActive Publication Date: 2025-07-25INNER MONGOLIA NORMAL UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510403004.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-25
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing dust storm path prediction technology is difficult to deal with complex space-time dependence, lacks long-term memory capabilities, and has high dependence on meteorological data, resulting in high short-term prediction accuracy but decreasing long-term prediction accuracy.

Method used

By obtaining the aerosol optical thickness data set and geographical background information, the Otsu threshold segmentation algorithm is used to generate a binary marked dust storm image sequence, and key features are screened in combination with the random forest feature importance method, and a hybrid model is constructed, and feature stitching and training is used to achieve the prediction of the dust storm movement path.

Benefits of technology

It improves the accuracy and reliability of sandstorm movement path prediction, can effectively deal with complex space-time dependencies, reduce model complexity and computational costs, and provides scientific support for disaster prevention and mitigation work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375210A_ABST
    Figure CN120375210A_ABST
Patent Text Reader

Abstract

The invention relates to a sand storm moving path prediction method and device, equipment and a storage medium, and the method comprises the steps: obtaining an aerosol optical thickness data set and geographical background information of a sand storm event; generating a binarized marked sand storm image sequence based on the aerosol optical thickness data set, and screening key features from the geographical background information through a random forest feature importance method; performing feature splicing on the sand storm image sequence and the key features to obtain an input data set; constructing an initial hybrid model for predicting a sand storm moving path, and training the initial hybrid model by using the input data set to obtain a target hybrid model; and predicting a future sand storm moving path by using the target hybrid model. By adopting the scheme, efficient and accurate prediction of the sand storm moving path is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural disaster prevention, and in particular, to a method, device, equipment and storage medium for predicting the moving path of sandstorms. Background Art

[0002] With the gradual change of the global climate, sandstorm events have become increasingly frequent in recent years, and their influence scope has expanded to thousands of kilometers. According to the definition of the World Meteorological Organization, a sandstorm refers to a strong wind that blows a large amount of sand and dust from bare and dry soil into the atmosphere and carries it to places hundreds to thousands of kilometers away. Sandstorms mainly occur in arid and semi-arid regions, and their destructive power is huge, which may cause direct disasters such as house collapses, human and livestock casualties, etc. In addition, sandstorms absorb and scatter solar radiation, change the microphysical characteristics of clouds, reduce precipitation, further exacerbate drought, reduce air quality, and trigger various diseases. The short-term and long-term impacts of sandstorms on the health of urban residents are significant, especially the increased risk of respiratory and cardiovascular diseases. In view of this, spatio-temporal monitoring, modeling, forecasting of the moving path of sandstorms and the development of early warning systems are of great significance for mitigating and preventing the impacts of sandstorms on the environment, health and socio-economic aspects of urban areas where people live.

[0003] The existing sandstorm path prediction methods mainly include the following three categories: The first category is to simulate the propagation process of sandstorms based on the atmospheric dynamics equation combined with meteorological data. Such models rely on the input of high-precision meteorological data, are sensitive to initial conditions, are suitable for long-term large-scale prediction, but are difficult to meet the short-term prediction requirements in complex dynamic environments, and have limited ability to model the spatial distribution characteristics of sandstorms. The second category is to use satellite remote sensing data (such as aerosol optical depth AOD), ground lidar and other sensors to monitor the dust concentration and its spatial distribution in the atmosphere, and predict the moving path of sandstorms by analyzing these data. Remote sensing technology can provide high temporal and spatial resolution data, which is convenient for real-time monitoring of sandstorm dynamics and can be combined with ground observations to improve prediction accuracy. However, this method is too dependent on the acquisition and processing speed of remote sensing data, is difficult to perform complex dynamic simulations alone, and usually needs to be combined with other models (such as numerical models or deep learning models). The third category is deep learning methods, which can integrate multi-dimensional environmental factors and improve prediction accuracy. In particular, models such as convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) perform excellently in spatio-temporal feature extraction. However, CNNs lack time memory ability, and single LSTMs are difficult to capture spatial features, which limits their application effects in sandstorm path prediction.

[0004] Through research, it is found that the existing sandstorm path prediction technology has the following main defects: First, it is difficult to handle complex spatio-temporal dependencies. Specifically, traditional models such as numerical models or remote sensing data methods are difficult to effectively handle the complex spatio-temporal dependencies in the sandstorm movement path. Numerical models are highly dependent on initial conditions and are difficult to adapt to short-term high spatio-temporal resolution predictions; remote sensing technology is overly dependent on data acquisition and processing speed and is difficult to perform complex dynamic simulations. Second, it lacks long-term memory ability. Specifically, traditional CNN models lack long-term memory ability and cannot use historical data information for prediction, resulting in relatively high short-term prediction accuracy but decreased long-term prediction accuracy. Third, it has a high dependence on meteorological data. Specifically, numerical models and some remote sensing methods have a high dependence on meteorological data, and the prediction accuracy is greatly affected by the quality of meteorological data. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a sandstorm movement path prediction method, device, equipment and storage medium to achieve efficient and accurate prediction of the sandstorm movement path.

[0006] In the first aspect, an embodiment of the present application provides a sandstorm movement path prediction method, and the method includes:

[0007] Obtain the aerosol optical depth data set and geographical background information of the sandstorm event;

[0008] Generate a binary-labeled sandstorm image sequence based on the aerosol optical depth data set, and screen out key features from the geographical background information through the random forest feature importance method;

[0009] Perform feature splicing on the sandstorm image sequence and the key features to obtain an input data set;

[0010] Construct an initial hybrid model for sandstorm movement path prediction, and use the input data set to train the initial hybrid model to obtain a target hybrid model;

[0011] Use the target hybrid model to predict the future sandstorm movement path.

[0012] Optionally, the aerosol optical depth data set contains several aerosol optical depth distribution maps, and the generating a binary-labeled sandstorm image sequence based on the aerosol optical depth data set includes:

[0013] Perform binary processing on each aerosol optical depth distribution map using the Otsu threshold segmentation algorithm to obtain several images to be labeled;

[0014] Mark the pixel points in each image to be marked with pixels higher than the preset threshold as the sand and dust area, and mark the pixel points with pixels lower than the preset threshold in the image to be marked as the non-sand and dust area, to obtain a number of marked images;

[0015] Construct the sandstorm image sequence based on each marked image.

[0016] Optionally, the screening of the key features from the geographical background information by the random forest feature importance method includes:

[0017] Use the random forest model to calculate the mean decrease in impurity (MDI) score of each feature in the geographical background information;

[0018] Determine the key features according to the ranking of the MDI scores of each feature.

[0019] Optionally, the initial hybrid model includes 3 convolutional layers, 1 bidirectional long short-term memory network layer, 2 max pooling layers, 2 fully connected layers, and a dropout layer.

[0020] Optionally, in the 3 convolutional layers, the first convolutional layer uses 16 3*1 filters, the second convolutional layer uses 16 3*16 filters, and the third convolutional layer uses 16 3*16 filters; the activation function of each convolutional layer is ReLU, and the model weights are initialized by the GlorotUniform initialization method.

[0021] Optionally, the bidirectional long short-term memory network layer contains 32 units and can capture the forward and backward information of the sequence simultaneously.

[0022] Optionally, after predicting the future sandstorm movement path using the target hybrid model, the method further includes:

[0023] Obtain a number of movement path prediction results, and count the number of true positive results, true negative results, false positive results, and false negative results in the movement path prediction results;

[0024] Determine the overall accuracy, F1 value, and Kappa coefficient based on the number of true positive results, true negative results, false positive results, and false negative results in the movement path prediction results;

[0025] Evaluate the performance of the target hybrid model based on the overall accuracy, F1 value, and Kappa coefficient.

[0026] In a second aspect, an embodiment of the present application provides a sandstorm movement path prediction device, and the device includes:

[0027] A data acquisition module for acquiring an aerosol optical depth dataset and geographical background information of a sandstorm event;

[0028] A data processing module for generating a sequence of binary-labeled sandstorm images based on the aerosol optical depth dataset, and screening out key features from the geographical background information by a random forest feature importance method;

[0029] A dataset construction module for performing feature splicing on the sandstorm image sequence and the key features to obtain an input dataset;

[0030] A model training module for constructing an initial hybrid model for predicting the moving path of a sandstorm, and training the initial hybrid model with the input dataset to obtain a target hybrid model;

[0031] A path prediction module for predicting the future moving path of a sandstorm using the target hybrid model.

[0032] Optionally, the aerosol optical depth dataset contains several aerosol optical depth distribution maps, and generating the sequence of binary-labeled sandstorm images based on the aerosol optical depth dataset includes:

[0033] Performing binary processing on each aerosol optical depth distribution map using an Otsu threshold segmentation algorithm to obtain several images to be labeled;

[0034] Marking the pixel points with pixels higher than a preset threshold in each image to be labeled as sand dust regions, and marking the pixel points with pixels lower than the preset threshold in the image to be labeled as non-sand dust regions to obtain several labeled images;

[0035] Constructing the sandstorm image sequence based on each labeled image.

[0036] Optionally, screening out the key features from the geographical background information by the random forest feature importance method includes:

[0037] Calculating the mean decrease in impurity (MDI) scores of each feature in the geographical background information using a random forest model;

[0038] Determining the key features according to the ranking of the MDI scores of each feature.

[0039] Optionally, the initial hybrid model includes 3 convolutional layers, 1 bidirectional long short-term memory network layer, 2 max pooling layers, 2 fully connected layers, and a dropout layer.

[0040] Optionally, in the three convolutional layers, the first convolutional layer uses 16 3*1 filters, the second convolutional layer uses 16 3*16 filters, and the third convolutional layer uses 16 3*16 filters; the activation function of each convolutional layer is ReLU, and the model weights are initialized by GlorotUniform initialization method.

[0041] Optionally, the bidirectional long short-term memory network layer contains 32 units and can capture the forward and backward information of the sequence simultaneously.

[0042] Optionally, after predicting the future sandstorm movement path using the target hybrid model, the method further includes:

[0043] Obtain several movement path prediction results, and count the number of true positive results, true negative results, false positive results, and false negative results in the movement path prediction results;

[0044] Determine the overall accuracy, F1 value, and Kappa coefficient based on the number of true positive results, true negative results, false positive results, and false negative results in the movement path prediction results;

[0045] Evaluate the performance of the target hybrid model based on the overall accuracy, F1 value, and Kappa coefficient.

[0046] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the sandstorm movement path prediction method in any optional implementation manner in the first aspect are executed.

[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the sandstorm movement path prediction method in any optional implementation manner in the first aspect are executed.

[0048] The technical solutions provided by the present application include but are not limited to the following beneficial effects:

[0049] This application first collects two key data related to sandstorms, namely the aerosol optical depth (AOD) dataset and geographical background information. The AOD dataset is obtained through satellite remote sensing technology and can reflect the concentration and spatial distribution of dust in the atmosphere. Geographical background information includes environmental characteristics such as relative humidity, surface air temperature, surface skin temperature, wind speed and wind direction at different heights. These data provide a rich source of information for the model and help to more comprehensively understand the formation and movement mechanisms of sandstorms.

[0050] Then, the Otsu threshold segmentation algorithm is used to binarize the AOD dataset, generating a sequence of binarized sandstorm images. At the same time, the random forest feature importance method is adopted to screen out the key features from the geographical background information that have the greatest impact on the sandstorm movement path. This step improves the data processing efficiency by simplifying the data and screening key features, reduces the complexity and computational cost of the model, and avoids overfitting, thereby improving the generalization ability of the model.

[0051] Next, the processed sandstorm image sequence is feature - stitched with the screened key geographical background features to form a complete input dataset. The integrated dataset can provide richer information for the model, helping the model to more accurately capture the spatio - temporal characteristics of sandstorms, thus improving the accuracy and reliability of predictions.

[0052] Then, an initial hybrid model is constructed, which combines several convolutional layers and bidirectional long short - term memory network layers. The initial hybrid model is trained using the stitched input dataset to obtain the target hybrid model. This innovative model structure that combines 1DCNN and BiLSTM can effectively handle the complex spatio - temporal dependencies in the sandstorm movement path and significantly improve the model's prediction ability.

[0053] Finally, the trained target hybrid model is used to predict the future movement path of sandstorms. After obtaining the prediction results, the performance of the model is evaluated, the number of true positives, true negatives, false positives, and false negatives in the prediction results is counted, and metrics such as overall accuracy, F1 - score, and Kappa coefficient are calculated. This step can not only provide scientific support for actual disaster prevention and mitigation work but also ensure the reliability of the model's prediction results through multi - dimensional performance evaluation, providing a basis for further optimization of the model.

[0054] Through the above steps, from data acquisition to model training, and then to final prediction and evaluation, each step has clear beneficial effects, jointly constructing a complete and efficient sandstorm path prediction method. This method can not only achieve efficient and accurate prediction of the sandstorm movement path but also provide scientific support for actual disaster prevention and mitigation work, having important application value.

[0055] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following presents preferred embodiments in conjunction with the accompanying drawings and provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0057] Figure 1 Shows a flowchart of a method for predicting the movement path of a sandstorm provided in the first embodiment of the present invention;

[0058] Figure 2 Shows a flowchart of a method for generating a sandstorm image sequence provided in the first embodiment of the present invention;

[0059] Figure 3 Shows a flowchart of a method for screening key features provided in the first embodiment of the present invention;

[0060] Figure 4 Shows a flowchart of a method for evaluating the performance of a model provided in the first embodiment of the present invention;

[0061] Figure 5 Shows a schematic structural diagram of a device for predicting the movement path of a sandstorm provided in the second embodiment of the present invention;

[0062] Figure 6 Shows a schematic structural diagram of a computer device provided in the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0064] Embodiment 1

[0065] For ease of understanding of this application, the following will combine Figure 1 the content described in the flowchart of a sandstorm movement path prediction method provided by Embodiment 1 of the present invention shown in

[0066] Refer to Figure 1 as shown Figure 1 which shows the flowchart of a sandstorm movement path prediction method provided by Embodiment 1 of the present invention. Among them, the method includes steps S101 to S105:

[0067] S101: Obtain the aerosol optical depth dataset and geographical background information of the sandstorm event.

[0068] Specifically, collect two key data related to sandstorms, namely the aerosol optical depth (AOD) dataset and geographical background information. The AOD dataset is obtained through satellite remote sensing technology and can reflect the concentration and spatial distribution of dust in the atmosphere. The NASA's MERRA-2 dataset is an important source for obtaining AOD data with high temporal resolution (hourly) and spatial resolution (0.5°×0.625°), which is crucial for accurately capturing the dynamic changes of sandstorms. The MERRA-2 dataset is generated by NASA's Global Modeling and Assimilation Office (GMAO) and is a satellite-based reanalysis model that combines the AOD datasets from the Advanced Very High Resolution Radiometer (AVHRR), Multi-angle Imaging SpectroRadiometer (MISR), MODIS AOD dataset, and the Aerosol Robotic Network (AERONET) ground observations. This study uses the "dustextinctionaerosolopticalthickness (AOT) 550nm" data in MERRA-2, covering 84 dust events and 2016 storm hours from 2000 to 2024.

[0069] The geographical background information includes environmental characteristics such as relative humidity, surface air temperature, surface skin temperature, wind speed and direction at different heights, etc. These data can usually be obtained from meteorological observation stations or reanalysis datasets, providing the environmental background for the occurrence and movement of sandstorms for the model. By combining this geographical background information with the AOD dataset, the model can more comprehensively understand and predict the movement path of sandstorms, providing strong support for disaster prevention, mitigation and environmental protection.

[0070] S102: Generate a sequence of binarized marked sandstorm images based on the aerosol optical depth dataset, and screen out key features from the geographical background information through the random forest feature importance method.

[0071] Specifically, the aerosol optical depth dataset contains remote sensing images of dust storm events, where the target variable is AOD. Each remote sensing image is processed using the Otsu threshold segmentation algorithm to divide the pixels in the image into "dust pixels" and "non-dust pixels", thereby obtaining a series of labeled images that can clearly show the scope and location of the dust storm. Secondly, to improve the efficiency and accuracy of the model, it is necessary to screen out the key features that have the greatest impact on the movement path of the dust storm from numerous geographical background information. This is achieved through the random forest feature importance method, which can evaluate the importance of each feature in the prediction task and sort the features according to the mean decrease in impurity (MDI) score, and finally determine the key features.

[0072] S103: Feature splice the dust storm image sequence and the key features to obtain an input dataset.

[0073] Specifically, the purpose of this step is to combine spatial features and temporal features to form a complete input dataset. Specifically, the processed dust storm image sequence (containing spatial information) is feature spliced with the selected key geographical background features (containing temporal information) to provide comprehensive input for the model, enabling the model to simultaneously consider the spatial distribution and temporal changes of the dust storm during the learning process, thereby better capturing the spatio-temporal features of the dust storm movement path.

[0074] S104: Construct an initial hybrid model for predicting the dust storm movement path, and use the input dataset to train the initial hybrid model to obtain a target hybrid model.

[0075] Specifically, when constructing the initial hybrid model, the model combines a convolutional layer (1DCNN) and a bidirectional long short-term memory network layer (BiLSTM). 1DCNN is mainly used to extract spatial features from the dust storm image sequence. It identifies basic features in the image and reduces the data dimension through operations of the convolutional layer and the pooling layer. BiLSTM is used to capture long-term dependencies in the time series and can consider context information in both the forward and backward directions, thereby improving the model's understanding and prediction ability for time series data. During the model construction process, the model weights are initialized by the GlorotUniform initialization method. Then, the spliced input dataset is used to train this initial hybrid model. By adjusting the model parameters, it can accurately learn the features and laws of the dust storm movement path, and finally obtain the target hybrid model.

[0076] When performing model training, the deep learning framework consists of three main parts: model input, prediction model, and model output. The data (1992 hours out of 2016 storm hours) is divided into 1592 training samples (= 80% of the data), 200 test samples (= 10% of the data), and 200 validation samples (= 10% of the data). The input of the deep learning model includes the original MERRA-2 AOD data (image) at time step t-1 and the geographical environment information at time step t-1, and the output includes the predicted sandstorm images for 24 hours (t to t+24). The labeled AOD data layer is used as the target layer in the training stage of the deep learning model. The movement path of the sandstorm is greatly affected by the geographical background, so the AOD dataset and the geographical background information (i.e., background information) are superimposed to construct the input layer. Due to a time series problem, the original AOD images corresponding to each hour of the storm with 71×81 pixels and the context geographical information are input. The neural network is used to extract spatial and temporal features from the input data to achieve the prediction of the model.

[0077] S105: Predict the future movement path of the sandstorm using the target hybrid model.

[0078] Specifically, use the trained target hybrid model to predict the future movement path of the sandstorm. After obtaining the prediction results, in order to evaluate the performance of the model, it is necessary to perform statistical analysis on the prediction results. Specifically, count the number of true positives, true negatives, false positives, and false negative results in the prediction results, and then calculate indicators such as overall accuracy, F1 value, and Kappa coefficient based on these statistical data to quantify the prediction accuracy and reliability of the model. This step is crucial for verifying the effectiveness and practicality of the model and can also provide a basis for further optimization of the model.

[0079] In an alternative embodiment, refer to Figure 2 as shown Figure 2 shows a flowchart of a method for generating a sandstorm image sequence provided in the first embodiment of the present invention. Among them, the aerosol optical thickness dataset contains several aerosol optical thickness distribution maps. The method for generating a binary-labeled sandstorm image sequence based on the aerosol optical thickness dataset includes steps S201 to S203:

[0080] S201: Use the Otsu threshold segmentation algorithm to perform binary processing on each aerosol optical thickness distribution map to obtain several images to be labeled.

[0081] Specifically, the Otsu threshold segmentation algorithm is applied to each acquired AOD distribution map. This algorithm can automatically determine the optimal threshold for image segmentation, maximizing the between-class variance between the foreground (dust area) and the background (non-dust area). Through this process, each AOD distribution map is converted into a binary image, where the part with pixel values higher than the threshold is marked as "dust pixels", and the part lower than the threshold is marked as "non-dust pixels". The key to this step is that the Otsu algorithm can adaptively determine the optimal threshold without manual intervention, thus improving the processing efficiency and accuracy.

[0082] S202: Mark the pixel points with pixel values higher than the preset threshold in each image to be marked as the dust area, and mark the pixel points with pixel values lower than the preset threshold in the image to be marked as the non-dust area, obtaining a number of marked images.

[0083] Specifically, mark the pixel points with pixel values higher than the preset threshold in each image to be marked as the dust area. These pixel points usually appear as high-brightness areas in the image, indicating a high concentration of dust; while mark the pixel points with pixel values lower than the threshold as the non-dust area. These pixel points appear as low-brightness areas in the image, indicating no or very little dust. After this marking process, a number of marked images with clearly distinguished dust and non-dust areas are obtained, and these images will serve as the basis for constructing the subsequent sandstorm image sequence.

[0084] S203: Construct the sandstorm image sequence based on each marked image.

[0085] Specifically, process the hourly pictures into sequence data, perform data fusion, and delete the data outside the research area to obtain all the pixel points at the t - 1 moment, where t is the current moment. This sequence contains the distribution of the sandstorm at different time points and can intuitively display the movement and change process of the sandstorm. This image sequence will be used as one of the inputs of the model to capture the spatial characteristics and temporal dynamics of the sandstorm, providing important data support for subsequent model training and prediction.

[0086] In an alternative embodiment, refer to Figure 3 as shown, Figure 3 shows a flowchart of a key feature screening method provided in Embodiment 1 of the present invention. Among them, the step of screening key features from the geographical background information by the random forest feature importance method includes steps S301 - S303:

[0087] S301: Calculate the mean decrease in impurity (MDI) scores of each feature in the geographical background information using a random forest model.

[0088] Specifically, the MERRA-2 dataset provides a variety of environmental feature information, including 18 features such as relative humidity, surface air temperature, surface skin temperature, surface wind speed, surface wind direction, 10-meter wind direction, 50-meter wind direction, 500-meter wind direction, and their temperature and humidity, pressure and their temperature and humidity. In order to select the most effective parameters from these features for predicting the movement path of sandstorms, this paper uses the Random Forest Feature Importance (RFFI) technique. The RFFI technique can identify the most important features or variables in the dataset for a specific prediction task. It trains a random forest model on the dataset and then evaluates the contribution of each feature to the model's prediction accuracy.

[0089] Random forest is an ensemble learning method that improves the accuracy and stability of the model by constructing multiple decision trees and aggregating their results. Each decision tree splits according to the importance of the features during construction, and the Mean Decrease in Impurity (MDI) score is an indicator to measure the importance of each feature in the model. Specifically, the MDI score determines its importance by quantifying the decrease in impurity (or increase in purity) when a certain feature is used for decision tree splitting. The higher the MDI score, the greater the impact of the feature on the model. RFFI finally obtains the final importance score of each feature by averaging the MDI scores of each feature in all decision trees in the model.

[0090] Through this method, the most valuable features for predicting the movement path of sandstorms can be screened out from numerous geographical background information, thereby reducing the complexity and computational cost of the model, while improving the efficiency and accuracy of the model.

[0091] S302: Determine the key features according to the ranking of the MDI scores of each feature.

[0092] Specifically, first rank each feature according to its importance score in the model. Features with higher scores play a more important role in the model and can better explain and predict the movement path of sandstorms. By setting a threshold or selecting the top N features, the key features that have the greatest impact on the movement path of sandstorms can be determined. The purpose of this step is to reduce the complexity and computational cost of the model, while improving the efficiency and accuracy of the model. By only retaining the key features, we can construct a more concise and effective model, avoid overfitting and improve the generalization ability.

[0093] In addition, data standardization is an important step to ensure model performance. The features are standardized using the Z-score scaler so that the mean of the data is 0 and the standard deviation is 1, transforming the original data into a standard normal distribution. This standardization method helps ensure that all features are on the same scale so that the model can learn them equally. At the same time, standardization also helps reduce the impact of outliers, which can have a significant impact on model performance. The formula below shows the normalization result X of the Z-score scaler scaled The calculation process:

[0094]

[0095] Among them, X represents the original data, μ is the mean of the original data, and σ is the standard deviation of the original data. Through this standardization process, the stability and predictive ability of the model can be further improved.

[0096] In an optional embodiment, the initial hybrid model includes 3 convolutional layers, 1 bidirectional long short-term memory network layer, 2 maximum pooling layers, 2 fully connected layers and a dropout layer.

[0097] Specifically, the convolutional layer performs well in processing sequence data, and can automatically extract spatial features in the data through convolution operations, thereby reducing the reliance on data preprocessing and feature engineering. It uses multiple filters to slide on the input data to capture local high-dimensional features, and further reduces the data dimension through pooling operations, thereby improving the computational efficiency and robustness of the model.

[0098] The bidirectional long short-term memory network layer focuses on the processing of time series data. Its unique bidirectional structure enables it to model sequence data from both the forward and reverse directions. This mechanism enables the bidirectional long short-term memory network layer to capture the long-term dependencies between time steps in the time series and effectively use historical information to improve the accuracy of predictions. For the task of predicting the movement path of sandstorms, long-term dependencies in time series are crucial because the movement of sandstorms is often affected by a combination of various meteorological conditions and geographical environmental factors in the early stages.

[0099] In the hybrid model, the convolutional layer first extracts spatial features from the input sandstorm image sequence to capture the spatial pattern and structure of the sand distribution in the image. Subsequently, the bidirectional long short-term memory network layer takes over the processing, models these spatial features in the time dimension, and analyzes the movement trend and change law of the sandstorm at different time points. Through this deep fusion of spatial and temporal features, the hybrid model can fully understand and predict the movement path of the sandstorm, overcoming the limitations of a single model in processing complex spatiotemporal data.

[0100] In an alternative embodiment, in the said three convolutional layers, the first convolutional layer uses 16 3×1 filters, the second convolutional layer uses 16 3×16 filters, and the third convolutional layer uses 16 3×16 filters; the activation function of each convolutional layer is ReLU, and the model weights are initialized by the GlorotUniform initialization method.

[0101] Specifically, an overview of the convolutional layer structure is as follows: In the hidden layer of a 1DCNN one-dimensional convolutional neural network, the network involves two special matrix operations: the convolutional layer and the pooling layer. The convolutional layer, as a local feature extractor, is used to scan the input data and extract local high-dimensional features through multiple different filters. By sliding the filter (ω i ) over the input data (x t ) and performing convolution with the input data, the neurons in each convolutional layer perform non-linear calculations, perform dot products, and are optionally connected to the neurons in the next layer to generate multiple features (c t,i ). Each neuron implements local connection and shared weights, thereby reducing the complexity of the model and accelerating the training efficiency. The convolution equation is shown as follows:

[0102]

[0103] where, * represents the convolution operation, ω i represents the filter vector, b i represents the bias vector, and c t,i represents the 1DCNN output data features. After one-dimensional convolution, there is still redundant information. To reduce redundancy and improve the robustness of feature extraction, a pooling layer (downsampling) is added to reduce the computational cost and improve the learning efficiency through local averaging or max pooling. Convolution and pooling can be repeated multiple times until the feature map is reduced to 1×1. Finally, the fully connected layer flattens the features, and the output layer calculates the optimal parameters using the optimal loss function. CNN often uses the stochastic gradient descent method to optimize the training and improve the model performance.

[0104] In an alternative embodiment, the bidirectional long short-term memory network layer contains 32 units and can capture the forward and backward information of the sequence simultaneously.

[0105] Specifically, compared with the traditional Recurrent Neural Network (RNN), the Bidirectional Long Short-Term Memory Network layer BiLSTM can capture both forward and backward context information simultaneously. It can not only effectively handle the long-term dependence problem but also reduce the possibility of gradient vanishing, thus better predicting the current information. The core concepts of BiLSTM include the cell state and the gate structure. The cell state can transmit important information, overcome the limitations of short-term memory, and ensure the effective transmission of long-term dependence information. BiLSTM includes three gate structures: the input gate, the forget gate, and the output gate, each with its specific role, which are used to selectively remember, forget, or output information. For forward propagation, set the input sequence X = [x1, x2,..., xn], where xn is the input at time step n. The calculation process of the forward LSTM is as follows:

[0106] 1. The calculation formula of the forget gate:

[0107] where f t is the current output of the forget gate; σ is the Sigmoid activation function; W f is the learnable parameter matrix, and its dimension depends on the input features and the number of hidden units; h t-1 and represent the hidden state at the previous moment and the input at the current moment respectively; b f is the bias vector of the forget gate, which is used to adjust the output of the forget gate and, together with W f ensures the flexibility of the model to learn the forgetting rule.

[0108] 2. The calculation formula of the input gate:

[0109]

[0110] where i t is the current output of the input gate; σ is the Sigmoid activation function; W I is the weight matrix of the input gate, which is used to control the contribution degree of the current input and the previous hidden state h t-1 to the new cell state; h t-1 and represent the hidden state at the previous moment and the input at the current moment respectively; b i is the bias vector of the input gate, which is used to adjust the output of the input gate; is the candidate cell state, and its value is restricted to the interval [-1, 1] by the tanh activation function; W C is the weight matrix of the candidate cell state, which is used to calculate the weight parameter of ; b cis the bias vector of the candidate cell state, used to adjust the candidate state output.

[0111] 3. Cell update:

[0112] Among them, c t is the cell state at the current moment; c t-1 is the cell state at the previous moment, which is the long-term memory passed in the BiLSTM network at time t-1 and contains all past time step information; f t is the output of the forget gate at the current moment, used to control the information forgotten in c t-1 ; is the candidate cell state; i t is the current output of the input gate.

[0113] 4. Output gate and hidden state update:

[0114] h t = o t ⊙tanh(c t );

[0115] Among them, o t is the current output of the output gate; σ is the Sigmoid activation function; W o is the output gate weight matrix, h t-1 and respectively represent the hidden state at the previous moment and the input at the current moment; b o is the bias vector of the output gate, used to adjust the output result of the output gate, and cooperate with W o to ensure the model learning ability; h t is the hidden state at the current moment; c t is the cell state at the current moment.

[0116] The hybrid model developed based on convolutional layers and bidirectional long short-term memory network layers in this application contains 13 different layers, including 2 convolutional layers, 1 bidirectional LSTM layer, 1 Flatten layer, 3 fully connected layers, and multiple batch normalization layers and Dropout layers. These layers are organized into several main modules: a feature extraction module and a prediction module. The feature extraction module contains two convolutional layers and one bidirectional LSTM layer: the first convolutional layer uses 16 3*1 filters, the second convolutional layer uses 16 3*16 filters, and the third convolutional layer uses 16 3*16 filters. The activation function of the convolutional layer is ReLU, and the model weights are initialized by GlorotUniform initialization method to improve the stability and convergence efficiency of model training. A bidirectional LSTM layer, containing 32 units, can capture the forward and backward information of the sequence simultaneously, enhancing the model's ability to handle time-dependent features. Between the convolutional layers, dimensionality reduction is performed through two max pooling layers with a pooling window size of 2, reducing the feature dimension, improving the computational efficiency, and reducing the risk of overfitting.

[0117] The prediction module consists of two fully connected layers: one of the fully connected layers contains 64 neurons and uses the ReLU activation function, and the other output layer uses the softmax activation function according to the number of classifications (the softmax activation function is mostly used in the output layer of classification, and it interprets the output as a probability distribution, that is, the predicted probability of each category). Between these two fully connected layers, a Dropout layer is added to prevent overfitting. A total of 23,346 trainable parameters are used during the training process.

[0118] In an alternative embodiment, refer to Figure 4 as shown Figure 4 shows a flowchart of a model performance evaluation method provided in the first embodiment of the present invention. After predicting the future movement path of a sandstorm using the target hybrid model, the method further includes steps S401 to S403:

[0119] S401: Obtain several movement path prediction results, and count the number of true positive results, true negative results, false positive results, and false negative results in the movement path prediction results.

[0120] Specifically, first use the trained target hybrid model to predict the future movement path of a sandstorm, obtaining a series of prediction results. These results are presented in the form of images, showing the possible positions and ranges of the sandstorm at different time points. Next, a detailed statistical analysis of these prediction results is required. Specifically, compare the prediction results with the actual movement path of the sandstorm and count the number of the following four results:

[0121] True Positive (TP): The number of results where the model correctly predicts a dust area.

[0122] True Negative (TN): The number of results where the model correctly predicts a non-dust area.

[0123] False Positive (FP): The number of results where the model incorrectly predicts a non-dust area as a dust area.

[0124] False Negative (FN): The number of results where the model incorrectly predicts a dust area as a non-dust area.

[0125] These statistics are the basis for evaluating the model's performance and can intuitively reflect the accuracy and reliability of the model during the prediction process.

[0126] S402: Determine the overall accuracy, F1 value, and Kappa coefficient based on the numbers of true positive results, true negative results, false positive results, and false negative results in the predicted mobile path results.

[0127] Specifically, these data are used to calculate three key performance metrics: overall accuracy, F1 value, and Kappa coefficient. The overall accuracy is obtained by dividing the number of correctly predicted results (TP and TN) by the sum of the total number of predictions, which reflects the overall proportion of correct predictions by the model. The F1 value is a useful quantitative metric for measuring the balance between precision and recall, and is obtained by calculating their harmonic mean, which is used to measure the balance between the precision and completeness of the model. The Kappa coefficient represents the degree of agreement between the predicted data and the reference data. A Kappa coefficient value of 1 indicates 100% agreement, and a value of 0 indicates no agreement. These three metrics comprehensively reflect the performance of the model from different perspectives.

[0128] Furthermore, the calculation formula for the overall accuracy Overall Accuracy is as follows:

[0129]

[0130] The calculation formula for the F1 value is as follows:

[0131]

[0132] The calculation formula for the Kappa coefficient k is as follows:

[0133]

[0134] N = TP + TN + FP + FN;

[0135] Among them, TP is the number of results where the model correctly predicts the dust storm area; TN is the number of results where the model correctly predicts the non-dust storm area; FP is the number of results where the model incorrectly predicts the non-dust storm area as the dust storm area; FN is the number of results where the model incorrectly predicts the dust storm area as the non-dust storm area; Precision is the accuracy rate; Recall is the recall rate; ρ o represents the observed agreement ratio, that is, the ratio of the actual classification result to the true label being consistent; ρ e represents the expected agreement ratio, that is, the expected agreement ratio assuming the classification is random, and N is the total number of samples.

[0136] S403: Evaluate the performance of the target hybrid model based on the overall accuracy rate, F1 value, and Kappa coefficient.

[0137] Specifically, comprehensively evaluate the performance of the target hybrid model according to the calculated overall accuracy rate, F1 value, and Kappa coefficient. By analyzing the values of these indicators, the accuracy and reliability of the model in predicting the movement path of sandstorms can be judged. For example, a higher overall accuracy rate indicates that the model can correctly predict most situations; a higher F1 value indicates that the model has achieved a good balance between precision and recall; and a higher Kappa coefficient indicates that the prediction result of the model has a high degree of coincidence with the actual situation, exceeding the expectation of accidental agreement. This multi-dimensional evaluation method ensures a comprehensive and objective understanding of the model performance, providing an important reference basis for the further optimization and practical application of the model.

[0138] Embodiment 2

[0139] Embodiment 2 of the present invention provides a device for predicting the movement path of sandstorms. Refer to Figure 5 as shown Figure 5 shows a schematic structural diagram of a device for predicting the movement path of sandstorms provided by Embodiment 2 of the present invention. Among them, the device includes:

[0140] A data acquisition module 501, configured to acquire an aerosol optical depth data set of a sandstorm event and geographical background information;

[0141] A data processing module 502, configured to generate a sequence of binarized marked sandstorm images based on the aerosol optical depth data set, and screen out key features from the geographical background information by the random forest feature importance method;

[0142] A data set construction module 503, configured to perform feature splicing on the sandstorm image sequence and the key features to obtain an input data set;

[0143] The model training module 504 is used to construct an initial hybrid model for predicting the moving path of sandstorms, and train the initial hybrid model using the input data set to obtain a target hybrid model;

[0144] The path prediction module 505 is used to predict the future moving path of sandstorms using the target hybrid model.

[0145] In an optional implementation, the aerosol optical depth data set contains several aerosol optical depth distribution maps. Generating the binary-labeled sandstorm image sequence based on the aerosol optical depth data set includes:

[0146] Performing binary processing on each aerosol optical depth distribution map using the Otsu threshold segmentation algorithm to obtain several images to be labeled;

[0147] Marking the pixel points with pixels higher than the preset threshold in each image to be labeled as sand dust areas, and marking the pixel points with pixels lower than the preset threshold in the image to be labeled as non-sand dust areas to obtain several labeled images;

[0148] Constructing the sandstorm image sequence based on each labeled image.

[0149] In an optional implementation, screening out key features from the geographical background information by the random forest feature importance method includes:

[0150] Using a random forest model to calculate the mean decrease in impurity (MDI) scores of each feature in the geographical background information;

[0151] Determining the key features according to the ranking of the MDI scores of each feature.

[0152] In an optional implementation, the initial hybrid model includes 3 convolutional layers, 1 bidirectional long short-term memory network layer, 2 max pooling layers, 2 fully connected layers, and a dropout layer.

[0153] In an optional implementation, in the 3 convolutional layers, the first convolutional layer uses 16 3*1 filters, the second convolutional layer uses 16 3*16 filters, and the third convolutional layer uses 16 3*16 filters; the activation function of each convolutional layer is ReLU, and the model weights are initialized by the GlorotUniform initialization method.

[0154] In an optional implementation, the bidirectional long short-term memory network layer contains 32 units and can capture the forward and backward information of the sequence simultaneously.

[0155] In an alternative embodiment, after predicting the future movement path of sandstorms using the target hybrid model, the method further includes:

[0156] Obtain a number of movement path prediction results, and count the numbers of true positive results, true negative results, false positive results, and false negative results in the movement path prediction results;

[0157] Determine the overall accuracy rate, F1 value, and Kappa coefficient based on the numbers of true positive results, true negative results, false positive results, and false negative results in the movement path prediction results;

[0158] Evaluate the performance of the target hybrid model based on the overall accuracy rate, F1 value, and Kappa coefficient.

[0159] Embodiment III

[0160] Based on the same inventive concept, refer to Figure 6 as shown Figure 6 which shows a schematic structural diagram of a computer device provided in Embodiment III of the present invention. Among them, as Figure 6 shown, a computer device 600 provided in Embodiment III of the present application includes:

[0161] A processor 601, a memory 602, and a bus 603. The memory 602 stores machine-readable instructions executable by the processor 601. When the computer device 600 runs, communication is carried out between the processor 601 and the memory 602 through the bus 603. When the machine-readable instructions are run by the processor 601, the steps of the sandstorm movement path prediction method shown in the above Embodiment I are executed.

[0162] Embodiment IV

[0163] Based on the same inventive concept, the present application embodiment also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the sandstorm movement path prediction method described in any one of the above embodiments are executed.

[0164] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0165] The computer program product for predicting the sandstorm movement path provided by the embodiments of the present invention includes a computer-readable storage medium storing program codes. The instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For specific implementation, refer to the method embodiments, and will not be elaborated herein.

[0166] The sandstorm movement path prediction device provided by the embodiments of the present invention can be specific hardware on a device, or software or firmware installed on the device, etc. The device provided by the embodiments of the present invention has the same implementation principle and the same technical effects as the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference can be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing described systems, devices, and units can all refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0167] In the embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0168] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] In addition, each functional unit in the embodiments provided by the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0170] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0171] It should be noted that like reference numerals and letters denote like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.

[0172] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for predicting the moving path of sandstorms, characterized in that, The method includes: Obtaining an aerosol optical depth dataset and geographical background information of a sandstorm event; Generating a sequence of binarized sandstorm images based on the aerosol optical depth dataset, and screening out key features from the geographical background information by means of the random forest feature importance method; Performing feature splicing on the sandstorm image sequence and the key features to obtain an input dataset; Constructing an initial hybrid model for predicting the moving path of a sandstorm, and training the initial hybrid model with the input dataset to obtain a target hybrid model; Predicting the future moving path of a sandstorm by using the target hybrid model.

2. The method according to claim 1, wherein The aerosol optical depth dataset contains several aerosol optical depth distribution maps. The generating of the sequence of binarized sandstorm images based on the aerosol optical depth dataset includes: Performing binarization processing on each aerosol optical depth distribution map by using the Otsu threshold segmentation algorithm to obtain several to-be-labeled images; Marking the pixel points with pixels higher than a preset threshold in each to-be-labeled image as sand dust regions, and marking the pixel points with pixels lower than the preset threshold in the to-be-labeled image as non-sand dust regions to obtain several labeled images; Constructing the sandstorm image sequence based on each labeled image.

3. The method according to claim 1, wherein The screening out of the key features from the geographical background information by means of the random forest feature importance method includes: Calculating the mean decrease in impurity (MDI) scores of each feature in the geographical background information by using a random forest model; Determining the key features according to the ranking of the MDI scores of each feature.

4. The method according to claim 1, characterized in that, The initial hybrid model includes 3 convolutional layers, 1 bidirectional long short-term memory network layer, 2 max pooling layers, 2 fully connected layers, and a dropout layer.

5. The method according to claim 4, wherein In the 3 convolutional layers, the first convolutional layer uses 16 3×1 filters, the second convolutional layer uses 16 3×16 filters, and the third convolutional layer uses 16 3×16 filters; the activation function of each convolutional layer is ReLU, and the model weights are initialized by means of the GlorotUniform initialization method.

6. The method according to claim 5, characterized in that, The bidirectional long short-term memory network layer contains 32 units and can capture the forward and backward information of the sequence simultaneously.

7. The method according to claim 1, characterized in that After predicting the future moving path of a sandstorm by using the target hybrid model, the method further includes: Obtaining several moving path prediction results, and counting the numbers of true positive results, true negative results, false positive results, and false negative results in the moving path prediction results; Determining the overall accuracy, F1 value, and Kappa coefficient based on the numbers of true positive results, true negative results, false positive results, and false negative results in the moving path prediction results; Evaluating the performance of the target hybrid model based on the overall accuracy, F1 value, and Kappa coefficient.

8. A sandstorm movement path prediction device, characterized in that, The device includes: A data acquisition module, configured to obtain an aerosol optical depth dataset and geographical background information of a sandstorm event; A data processing module, configured to generate a sequence of binarized sandstorm images based on the aerosol optical depth dataset, and screen out key features from the geographical background information by means of the random forest feature importance method; A dataset construction module, configured to perform feature splicing on the dust storm image sequence and the key features to obtain an input dataset; A model training module, configured to construct an initial hybrid model for predicting the moving path of a dust storm, and use the input dataset to train the initial hybrid model to obtain a target hybrid model; A path prediction module, configured to use the target hybrid model to predict the future moving path of a dust storm.

9. A computer device, characterized in that, It includes: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the dust storm moving path prediction method described in any one of claims 1 to 7 are executed.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the dust storm moving path prediction method described in any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Small convolutional nuclear cell counting method and system based on deep convolutional neural network

    CN110659718A

  • CNN and BILSTM-based ship trajectory prediction method

    CN114154619A

  • Feature extraction method based on automatic parking model and convolutional neural network

    CN114913335A

  • IGBT module state prediction method based on domain adversarial long short-term memory network

    CN118095346A

  • Systems and methods for digitizing electrocardiograms

    US20190298204A1