Surface mine slope monitoring and early warning system and method
By combining spatiotemporal perception multimodal fusion analysis and visual risk identification modules with digital twin technology, the problems of superficial multimodal data fusion and insufficient visual recognition accuracy in open-pit mine slope monitoring have been solved. This has enabled high-precision, dynamic early warning and closed-loop optimization, thereby improving the system's intelligence level.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH BEIJING
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-15
AI Technical Summary
Existing open-pit mine slope monitoring technologies suffer from problems such as superficial multimodal data fusion, insufficient visual recognition accuracy, open-loop systems, and static and rigid early warning thresholds, making it difficult to achieve dynamic intelligent early warning and closed-loop optimization.
By employing a spatiotemporal perception multimodal fusion analysis module, a visual risk identification module, and a dynamic early warning and decision support module, combined with digital twin technology, a closed-loop optimization system is formed, achieving deep fusion, accurate identification, and dynamic early warning of multimodal data.
It improves the accuracy and robustness of slope stability prediction, enhances the ability to identify small targets and weak edge features, and enables timely and intelligent early warning and system self-optimization.
Smart Images

Figure CN122045693A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of mine safety production, geological disaster prevention and control and artificial intelligence technology, specifically to an open-pit mine slope monitoring and early warning system and method that integrates artificial intelligence, Internet of Things and digital twin technology. Background Technology
[0002] The stability of open-pit mine slopes is directly related to the safety of life and property and the sustainability of production. Traditional assessment methods rely heavily on manual inspections and single-point sensors, which have inherent drawbacks such as long response delays, high false negative rates, and difficulty in dealing with complex geological and climatic conditions.
[0003] In recent years, intelligent monitoring technologies have made progress. For example, UAV remote sensing is used for rapid identification of landslide areas, but it does not combine geological context and continuous monitoring data, resulting in a single evaluation dimension; IoT-based monitoring systems have achieved multi-parameter acquisition, but lack deep data fusion and intelligent decision-making capabilities; some studies have attempted to use machine learning for multimodal data fusion, but these have mostly remained at the level of simple feature splicing and have failed to effectively model the complex spatiotemporal relationships between multi-source data.
[0004] Existing technologies suffer from the following prominent contradictions: First, the methods for fusing multimodal data (geological text, time-series sensing, and visual images) are rudimentary, lacking detailed modeling of the spatiotemporal correlation and dynamic weight allocation between data, resulting in insufficient information utilization. Second, visual recognition models for complex mining scenarios (light variations, dust interference, and terrain undulations) lack sufficient detection accuracy for key risk targets such as minute cracks on slope surfaces and small isolated loose rocks. Third, existing systems are mostly in an "open-loop" state, with the assessment, early warning, and decision-making processes being fragmented. Early warning thresholds are static and rigid, unable to be dynamically adjusted according to the environment, and the assessment results are difficult to directly guide engineering practice, failing to form an intelligent closed loop of "monitoring-assessment-early warning-decision-optimization".
[0005] Therefore, there is an urgent need for a systematic solution that can deeply integrate multimodal information, accurately identify various risks, and achieve dynamic intelligent early warning and closed-loop optimization. Summary of the Invention
[0006] This invention aims to overcome the shortcomings of existing open-pit mine slope monitoring and early warning technologies, and to provide an intelligent monitoring and early warning solution with high monitoring accuracy, timely dynamic early warning, and closed-loop optimization capabilities, so as to achieve early detection, early warning, and early handling of slope instability risks.
[0007] Specifically, the present invention provides the following technical solutions: On one hand, the present invention provides a monitoring and early warning system for open-pit mine slopes, the system comprising: The data preprocessing module preprocesses the multimodal data collected during the monitoring of open-pit mine slopes. The spatiotemporal sensing multimodal fusion analysis module extracts fusion feature vectors from preprocessed multimodal data to obtain stability quantification indicators, risk levels, and future deformation trends. This module is composed of a spatiotemporal sensing multimodal fusion network, which includes multiple specific encoders, spatiotemporal sensing weight allocation units, and progressive multiscale feature decoupling fusion units. The preprocessed multimodal data includes geological text data, sensor time-series data, and image data collected by UAVs. The visual risk identification module, implemented by a dual-stream visual recognition network, processes UAV remote sensing image data in the preprocessed multimodal data to obtain segmentation result data. The dynamic early warning and decision support module, based on the output results of the spatiotemporal perception multimodal fusion analysis module and the visual risk recognition module, combined with real-time environmental data and historical case data, calculates the early warning threshold, generates graded early warning information, and sends it to the terminal. The model closed-loop optimization and digital twin management module receives early warning and handling feedback and new monitoring data provided by the terminal, and optimizes the parameters of the spatiotemporal perception multimodal fusion analysis module and the dynamic early warning and decision support module.
[0008] Preferably, the plurality of modality-specific encoders include: a modality-specific encoder based on a pre-trained language model for feature encoding of geological text data to obtain text features; a modality-specific encoder based on a temporal convolutional network for feature encoding of sensor time-series data to obtain temporal features; and a modality-specific encoder based on a deep convolutional network for feature encoding of image data to obtain image features.
[0009] Preferably, the spatiotemporal awareness weight allocation unit is composed of a lightweight neural network, used to generate a set of weight vectors based on the temporal context information and spatial location encoding corresponding to text features, temporal features, and image features, so as to perform weighted summation of text features, temporal features, and image features to obtain weighted features: ,in, , , These respectively represent the features assigned to the text. Time series characteristics Image features The dynamic weighting coefficients.
[0010] Preferably, the progressive multi-scale feature decoupling and fusion unit receives text features, temporal features, and image features, and extracts feature sub-maps at different scales through multiple convolutional layers of different scales set in parallel for decoupling; then, the obtained multiple feature sub-maps are stitched together, fused through a convolutional layer, and then aggregated through a cross-scale connection method to obtain a fused feature vector.
[0011] Preferably, the dual-stream visual recognition network includes parallel edge enhancement channels and semantic parsing channels, followed by an attention-guided fusion layer; The edge enhancement channel extracts primary edge features from image data using multi-directional gradient operators; the semantic parsing channel extracts deep semantic features from image data based on an encoder-decoder architecture. The attention-guided fusion layer calculates the dependence weight of each spatial location in the deep semantic features on edge information using spatial attention. It then multiplies the primary edge features element-wise with the dependence weights and adds them to the deep semantic features to achieve the fusion of primary edge features and deep semantic features, resulting in a fused feature vector that serves as the basis for subsequent image segmentation. The spatial attention method is as follows: perform 1x1 convolution on deep semantic features to compress the number of channels to 1, and then generate normalized dependency weights through the Sigmoid activation function.
[0012] Preferably, a small target sensitive detection head is set at the end of the decoder of the semantic parsing channel; the small target sensitive detection head includes a channel attention layer and a convolutional segmentation head; the channel attention layer adopts a squeeze-excitation network structure to recalibrate the shallow high-resolution features from the semantic parsing channel; the convolutional segmentation head fuses the recalibrated features with the fused feature vector output by the attention-guided fusion layer, and outputs the small target segmentation result through the convolutional layer.
[0013] Preferably, the dynamic early warning and decision support module includes a dynamic feedback early warning unit based on digital twins. The dynamic feedback early warning unit constructs a digital twin of the slope and maps the output results of the spatiotemporal perception multimodal fusion analysis module and the visual risk identification module to the digital twin. The unit has a built-in threshold dynamic model, which is used to calculate the early warning threshold for different areas based on real-time environmental parameters. The threshold dynamic model adopts an encoder-decoder architecture. The encoder consists of LSTM layers, and the decoder decodes based on an attention mechanism and outputs a warning threshold curve.
[0014] Preferably, the decoupling and splicing method is as follows: the received text features, temporal features, image features, and weighted features are used as inputs to three convolutional layers of different scales, respectively, to obtain the following sub-maps in sequence: shallow feature sub-maps, mid-level feature sub-maps, and deep feature sub-maps corresponding to text features, shallow feature sub-maps, mid-level feature sub-maps, and deep feature sub-maps corresponding to temporal features, shallow feature sub-maps, mid-level feature sub-maps, and deep feature sub-maps corresponding to image features, and shallow feature sub-maps, mid-level feature sub-maps, and deep feature sub-maps corresponding to weighted features; During the stitching process: For each feature sub-map at each scale, feature sub-maps at the same scale from text features, temporal features, image features, and weighted features are stitched and fused across modally through a convolutional layer to obtain cross-modal fused features; finally, the cross-modal fused features at three different scales are aggregated through a cross-scale connection method to obtain a fused feature vector.
[0015] Preferably, the multi-directional gradient operator is a convolution kernel with a set of trainable parameters, and its calculation formula is as follows:
[0016] Where I represents the input image. For learnable convolutional kernel weights, Indicates the direction parameter. express Gradient operator for direction.
[0017] Preferably, at the end of the decoder of the semantic parsing channel, a conventional target detection head is connected. The conventional target detection head receives the fused feature vector output by the attention-guided fusion layer and outputs the segmentation result of a conventional-sized target.
[0018] Secondly, the present invention also provides a method for monitoring and early warning of open-pit mine slopes, which is applied to the system described above, and the method includes: S1. Collect text data, time series data and image data related to the monitoring of unfilled mine slopes, and perform preprocessing; S2. Construct a spatiotemporal awareness multimodal fusion network and a dual-stream visual recognition network, and train the network model. S3. Deploy the trained spatiotemporal awareness multimodal fusion network and dual-stream visual recognition network on the edge computing server in the mine; construct a digital twin of the slope; initiate dynamic early warning and monitor environmental parameters in the mine area in real time; issue alarm information when an early warning is issued; S4. Closed-loop optimization: After the warning is issued, the accuracy of the warning and the effectiveness of the handling are recorded. These are then combined with the corresponding real-time collected data to form a new training sample, which is used to optimize the network parameters.
[0019] Furthermore, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the method described above when executing the program.
[0020] Compared with the prior art, the present invention has the following advantages: (1) Deeper integration and higher accuracy: This scheme fully explores the complementary and related information between multimodal data through dynamic weight allocation with spatiotemporal awareness and progressive multi-scale fusion, which greatly improves the accuracy and robustness of stability prediction.
[0021] (2) More accurate identification and more comprehensive coverage: This solution effectively solves the problem of identifying small targets and weak edge features in complex environments by combining a dual-stream visual network with an attention mechanism, and improves the detection rate and segmentation accuracy of key risk points such as cracks and floating stones.
[0022] (3) Smarter early warning and more timely response: The early warning engine based on digital twin and dynamic threshold model can adapt to changes in the environment, with longer early warning lead time, lower false alarm rate, and more intuitive and operable decision support.
[0023] (4) The system is more closed-loop and its capabilities can evolve: This solution realizes a complete closed loop from data collection to model optimization. The system can continuously improve itself by using new data and feedback, and has the ability to continuously learn and evolve. It has high engineering practical value. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of the overall system architecture provided by the present invention.
[0026] Figure 2 This is a schematic diagram of the spatiotemporal awareness multimodal fusion network (ST-MMFN) structure provided by the present invention.
[0027] Figure 3 This is a schematic diagram of the attention-guided edge-semantic dual-stream visual recognition network (AGE-SNet) structure provided by the present invention.
[0028] Figure 4 The flowchart of the dynamic feedback early warning engine based on digital twin provided by this invention.
[0029] Figure 5 This is a schematic diagram of the timing data flow of the system modules working together as provided by the present invention. Detailed Implementation
[0030] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0031] Those skilled in the art should understand that the following specific embodiments or implementation methods are a series of optimized configurations listed to further explain the specific content of the invention. These configuration methods can be combined or used in conjunction with each other, unless the invention explicitly states that some or a specific embodiment or implementation method cannot be associated with or used in conjunction with other embodiments or implementation methods. Furthermore, the following specific embodiments or implementation methods are merely optimized configurations and are not intended to limit the scope of protection of the invention.
[0032] Combination Figure 1 The diagram illustrates the overall architecture of the system of this invention, describing multi-source data acquisition, four core processing modules (fusion analysis, visual recognition, dynamic early warning, and closed-loop optimization), and the data flow relationships between them. The following describes an open-pit mine slope monitoring and early warning system provided by this invention, which mainly includes the following key modules: The data preprocessing module preprocesses the multimodal data collected during the open-pit mine slope monitoring process, including: structured geological text reports, sensor time-series monitoring data, and UAV remote sensing image data. Data preprocessing primarily includes: text cleaning (removing irrelevant characters and standardizing formats); entity annotation (extracting geological parameters using named entity recognition models); time-series denoising (applying wavelet transform or Kalman filtering); outlier removal (using statistical methods or isolated forest algorithms to remove outliers); image data calibration (primarily radiometric and geometric correction); and image data registration (matching feature points to a unified coordinate system).
[0033] Spatiotemporal Aware Multimodal Fusion Analysis Module: This module is primarily implemented using a Spatiotemporal Aware Multimodal Fusion Network (ST-MMFN). Preprocessed multimodal data is input into the ST-MMFN, which extracts and fuses deep correlation features across modalities through a progressive multi-scale feature decoupling and fusion strategy. The module outputs three main indicators: slope stability quantification indicators, such as the safety factor Fs; risk levels, such as a four-level classification; and future deformation trends, such as the sidewave displacement prediction curve. These outputs are obtained by decoding the fused feature vectors using specific algorithms: stability quantification indicators are calculated through a fully connected regression layer; risk levels are determined through a Softmax classification layer; and future deformation trends are generated through a time-series prediction layer (such as LSTM or TCN).
[0034] In this embodiment, the Spatiotemporal Aware Multimodal Fusion Network (ST-MMFN) abandons simple feature splicing or decision-level voting, and innovatively introduces a spatiotemporal awareness weight allocation unit. This unit can dynamically adjust the contribution weights of text, temporal, and image features in the fusion process based on the time point (e.g., rainy season / dry season) and spatial source (e.g., different areas of a slope) of the monitoring data. Simultaneously, a progressive multi-scale feature decoupling and fusion strategy is adopted, aligning and fusing features of different modalities across multiple network layers. First, detailed features are fused, then semantic features are fused, achieving a deep understanding from local to global perspectives.
[0035] Visual Risk Recognition Module: This module is mainly implemented by a dual-stream visual recognition network. The UAV remote sensing image data is input into the dual-stream visual recognition network (AGE-SNet), which processes edge details and global semantic information in parallel to achieve high-precision identification and pixel-level segmentation of risk targets such as cracks, landslides, and loose rocks on slope surfaces. Finally, it outputs the identification and segmentation results of cracks, landslides, or loose rocks.
[0036] In this embodiment, the dual-stream visual recognition network (AGE-SNet) addresses the challenge of target recognition in complex scenes by employing a dual-parallel processing flow. One edge enhancement flow focuses on extracting pixel-level details from the edges of cracks and pumice stones; the other semantic parsing flow is responsible for understanding the global scene and object semantics of the image. Through an attention-guided fusion unit, the dual-stream visual recognition network can learn where to focus more on edge information (such as crack outlines) and where to rely more on semantic information (such as landslide areas), thereby significantly improving the recognition accuracy of blurry and small targets. Simultaneously, the network incorporates a small target sensitive detection head to enhance the pumice stone detection rate.
[0037] Dynamic Early Warning and Decision Support Module: Based on the output data of the spatiotemporal perception multimodal fusion analysis module and the visual risk recognition module, the module uses a digital twin-based dynamic feedback early warning unit to dynamically calculate the three-dimensional early warning threshold, generate graded early warning information and structured governance suggestions, and push them to the terminal output layer. This terminal may include, for example, terminal devices used by managers, monitoring center screens, or digital twin visualization platforms.
[0038] In this embodiment, a dynamic feedback early warning unit based on digital twins constructs a digital twin of the slope as a virtual mapping. The analysis results of ST-MMFN and AGE-SNet are mapped to the twin in real time, generating a three-dimensional dynamic risk heat map. The unit's built-in threshold dynamic model can automatically calculate and update the early warning thresholds for different areas based on environmental parameters such as real-time rainfall and displacement rate, achieving a transition from "static thresholds" to "dynamic thresholds." After an early warning is triggered, this unit automatically generates structured response suggestions. Based on feedback from the response effects and new data, the system continuously optimizes the fusion network and early warning model. Specifically, it updates the network weights and bias parameters of ST-MMFN, as well as the LSTM unit weights and attention mechanism parameters of the threshold dynamic model, through online learning or incremental learning algorithms (such as fine-tuning based on stochastic gradient descent), forming a complete intelligent closed loop of perception-analysis-decision-feedback.
[0039] Model closed-loop optimization and digital twin management module: Based on the early warning and handling feedback and new monitoring data provided by the terminal, the parameters of the fusion network of the spatiotemporal perception multimodal fusion analysis module and the dynamic feedback early warning unit in the dynamic early warning and decision support module are adaptively optimized to form an intelligent evaluation closed loop.
[0040] In addition, this system is equipped with multi-source data acquisition devices or layers to collect geological textual data, time-series sensor data, and UAV imagery data. Geological textual data includes slope technical literature, survey reports, and experimental data; sensor data includes GNSS, fracture gauges, and water level gauges; and UAV imagery includes visible light aerial photography data or infrared aerial photography data.
[0041] In a more preferred embodiment, combined with Figure 2 As shown, the internal structure of the Spatiotemporally Aware Multimodal Fusion Network (ST-MMFN) is presented, particularly the working principle of the spatiotemporally aware weight allocation unit and the progressive multi-scale feature decoupling and fusion unit. The Spatiotemporally Aware Multimodal Fusion Network includes: Multiple modality-specific encoders are implemented based on an improved domain-pre-trained language model, a temporal convolutional network, and a deep convolutional network, respectively. The modality-specific encoder based on the improved domain-pre-trained language model employs a BERT or similar Transformer model further pre-trained on a mine slope safety corpus. Its core processing functions are multi-head self-attention and a feed-forward network, used to encode geological slope text data into text features. The specific encoder based on the temporal convolutional network employs a TCN structure with causal dilated convolutions. Its core processing functions are dilated convolution and residual connections, used to encode sensor temporal data into temporal features. The specific encoder based on deep convolutional networks employs ResNet, VGG, or similar CNN architectures. Its core processing functions are convolution, pooling, and activation functions (such as ReLU), used to encode UAV image data into image features. .
[0042] Spatiotemporal Awareness Weight Allocation Unit: This layer receives feature-encoded data obtained from multiple specific encoders. Based on the temporal context (obtained from timestamps, seasonal, and weather cycle features) and spatial location encoding (obtained from the 3D coordinates of monitoring points or regional grid encoding) corresponding to the input time-series data segment, it processes this contextual information through a lightweight neural network (such as a multilayer perceptron) to dynamically generate a set of weight vectors for weighted fusion of feature data from different modalities, including the text features, temporal features, and image features obtained above. After obtaining the weight vectors, the features are weighted and summed. ,in, These respectively represent the features assigned to the text. Time series characteristics Image features The dynamic weighting coefficients.
[0043] Progressive multi-scale feature decoupling and fusion unit: This unit decouples the feature vectors output by each modality-specific encoder—that is, the feature vectors before weighted summation. and the weighted features output by the spatiotemporal perception weight allocation unit. These are all used as inputs and decoupled across multiple scales. Specifically, for each input feature vector, including... and Each feature map is extracted using three parallel 1x1 convolutional layers with different dilation rates (e.g., 1, 2, 4), representing three different scales of shallow details, mid-level structure, and deep semantics. These convolutional layers process different input features independently. Subsequently, at each scale (e.g., the shallow scale), the features from... and The feature sub-maps at each scale are cross-modal concatenated to form cross-modal fused scale features. Each scale's cross-modal fused features are concatenated and fused using a convolutional layer (e.g., a 3x3 convolution) to enhance feature consistency within the scale. Finally, the fused features from the shallow, mid-level, and deep scales are aggregated through cross-scale connections (e.g., skip connections or Feature Pyramid Network, FPN) to output a unified fused feature vector. This fused feature vector is ultimately used to calculate stability quantification indicators, risk levels, and future deformation trends.
[0044] In a more preferred embodiment, Figure 3 The dual-stream parallel architecture of the two-stream visual recognition network (AGE-SNet) is presented, along with how the attention-guided fusion module integrates edge and semantic information. The two-stream visual recognition network includes: Edge enhancement channel: A learnable multi-directional gradient operator is used to extract primary edge features from the image data. The learnable multi-directional gradient operator is a set of trainable convolutional kernels, simulating edge detection operators such as Sobel. Its calculation formula is as follows: Where I is the input image, For learnable convolutional kernel weights, Indicates the direction parameter. express The gradient operator in the direction is trained and optimized to obtain an edge filter suitable for mining scenarios.
[0045] Semantic parsing channel: Deep semantic features of the image are extracted based on an encoder-decoder architecture; the encoder-decoder architecture preferably adopts U-Net or its variants. The encoder consists of multiple downsampling stages, each containing convolutional layers, activation functions (such as ReLU), and pooling layers (such as max pooling), used to progressively extract hierarchical features of the image and reduce spatial resolution, while increasing the number of feature channels to encode rich semantic information. The decoder consists of multiple upsampling stages, each restoring spatial resolution through transposed convolution or upsampling operations, and combining with high-resolution and low-level features fused with the corresponding stages of the encoder through skip connections, so as to preserve deep semantic context while restoring details. The deep semantic feature extraction process is as follows: the original input image passes through multiple downsampling stages of the encoder in sequence, obtaining the most global and abstract deep feature representation at the network's "bottleneck" layer; this feature then passes through multiple upsampling and feature fusion stages of the decoder, finally outputting a deep semantic feature map corresponding to the spatial size of the input image. The feature vector at each position of this feature map fuses local appearance information and global scene semantics.
[0046] Attention-guided fusion layer: This layer receives the primary edge features and deep semantic features, and calculates the dependency weight of each spatial location (corresponding to each pixel region in the image) on edge information through a spatial attention mechanism. Specifically, the dependency weights... It is based on deep semantic feature maps The calculation results are as follows: ,in For 1x1 convolution, The Sigmoid activation function generates With deep semantic features The dimensions are the same. Then, the primary edge features are... With Dependency Weight Element-wise multiplication is performed to enhance regions in the deep semantic feature map that require attention to edge details, and then the result is combined with the original deep semantic features. Addition achieves dynamic fusion: ,in This indicates element-wise multiplication. The fused feature map highlights the outline and internal texture of the risky target.
[0047] The decoder of the semantic parsing channel is connected to a conventional target detection head, which directly receives the fused features output by the attention-guided fusion layer. Through a series of upsampling and convolutional layer processing, the segmentation results of targets of normal size (such as cracks and landslides) are output.
[0048] Small Target Sensitive Detection Head: This detection head is added at the end of the decoder (i.e., the decoder in the semantic parsing channel). Specifically, this detection head additionally receives high-resolution feature maps from the shallow layer of the encoder in the semantic parsing channel (denoted as...). First, a channel attention module (such as the Squeeze-and-Excitation module) is used to... Feature recalibration is performed to enhance feature channels sensitive to small targets. Then, the recalibrated features are combined with the fused feature map from the attention-guided fusion layer. (After upsampling to the same resolution) the data is stitched together. Finally, a lightweight convolutional segmentation head (consisting of several convolutional layers) processes the stitched features, outputting a segmentation mask specifically for small targets such as small pumice stones. This small target sensitive detection head, together with the regular target detection head, forms a dual-branch output structure. The outputs of the two branches (i.e., the regular target segmentation mask and the small target segmentation mask) are fused by taking the maximum value pixel by pixel or by passing through an additional convolutional layer to generate the final unified target segmentation map.
[0049] In a more preferred embodiment, combined with Figure 4 The diagram illustrates the workflow of the dynamic feedback early warning unit based on digital twins in this invention, primarily including data mapping, dynamic threshold calculation, early warning generation, and feedback optimization. Details are as follows: A digital twin of the slope is constructed to access and map the geometric and physical properties and monitoring data of the physical slope in real time; in the preferred embodiment, the digital twin integrates the geological model, monitoring points and three-dimensional terrain.
[0050] The stability index output by the spatiotemporal perception multimodal fusion analysis module and the risk spatial distribution identified by the visual risk identification module are mapped onto the digital twin to generate a three-dimensional risk heat map.
[0051] Based on real-time monitoring data, the red, orange, yellow, and blue warning thresholds for each zone at the current moment are calculated using a pre-trained threshold dynamic model. Real-time monitoring data includes, but is not limited to, GNSS displacement rate, crack opening and closing degree, groundwater level, and rainfall intensity.
[0052] When the monitored data exceeds the threshold, the engine automatically triggers an alert and generates a structured security briefing based on the case library and rule library through template filling and rule reasoning. The briefing is presented in the form of a graphic report or a structured data interface.
[0053] In a further optimized approach, the aforementioned dynamic threshold model employs a time-series prediction model based on a Long Short-Term Memory (LSTM) network and an attention mechanism. Its specific structure is an encoder-decoder architecture. The encoder, composed of LSTM layers, encodes historical monitoring sequences, while the decoder, guided by the attention mechanism, decodes and outputs the future warning threshold curve. The model learns online from the latest monitoring sequences and warning feedback data as new training samples, periodically incrementally learning and dynamically adjusting the internal network parameters of the dynamic threshold model (including the weights of the LSTM layers, attention mechanism parameters, and fully connected layer parameters) to achieve adaptive updates of the warning threshold.
[0054] The following section, using a practical application scenario, explains the operation method of the system proposed in this solution. Figure 5 This demonstrates the entire process from data input to early warning output in a real-world application scenario, as well as the collaborative relationships between various modules when processing data at different stages. Taking a specific iron ore mine as an example, the slope is approximately 350 meters high, with complex geological conditions, requiring 24 / 7 intelligent monitoring and early warning.
[0055] S1: Data Acquisition and Preprocessing S11: Text Data: Collected 800 geological survey reports and geotechnical test reports for the slope over the years. Natural language processing technology was used for cleaning and word segmentation, and key entities such as "lithology", "joint dip angle", and "groundwater" were labeled.
[0056] S12: Time-series data: 15 GNSS displacement monitoring points, 10 surface crack gauges, and 5 groundwater level gauges were deployed, with a data acquisition frequency of 1 time / hour. Wavelet denoising and outlier removal were performed on the data.
[0057] S13: Image Data: Two visible light and one thermal infrared aerial photographs are taken weekly using drones to generate orthophotos and 3D point cloud models. Radiometric correction and geometric registration are performed on the images.
[0058] S2: Model Building and Training S21: ST-MMFN Training: Building as follows Figure 2 The network is shown. The text encoder uses a pre-trained model finely tuned on a mining geology corpus; the temporal encoder uses a one-dimensional temporal convolutional network; and the image encoder uses ResNet. The spatiotemporal awareness weight allocation module is a small neural network that takes the index of the current temporal data segment and the corresponding sensor location code as input, and outputs the fused weights of the three modalities. This network is trained on three years of historical data, and the loss function is a weighted sum of the mean squared error of the stability coefficient (e.g., safety factor Fs) prediction and the cross-entropy loss of the risk level classification.
[0059] S22: AGE-SNet Training: Building as follows Figure 3The network is shown. The edge enhancement stream uses a set of trainable Sobel-like filters. The semantic parsing stream is based on the U-Net architecture. The attention-guided fusion module computes the attention weight map for each location of the semantic feature map to the edge features. Supervised training is performed using manually annotated crack, landslide, and pumice segmentation maps. The loss function combines Dice loss and Focal Loss to handle class imbalance.
[0060] S3: System Deployment and Operation S31: Deploy the trained ST-MMFN and AGE-SNet models on the edge computing server in the mine.
[0061] S32: Construct a digital twin of the slope, integrating geological models, monitoring points, and three-dimensional terrain.
[0062] S33: Dynamic warning engine (such as...) Figure 4 Startup and Operation: The engine receives preprocessed multimodal data in real time. ST-MMFN outputs the current overall slope stability score (e.g., Fs=1.15) and zonal risk level. AGE-SNet outputs the latest crack distribution map and pumice location. The engine maps these results to a digital twin to generate a three-dimensional risk heat map.
[0063] S34: At this time, the weather station reports heavy rainfall expected within the next 6 hours. The threshold dynamic model dynamically lowers the warning threshold for high-risk areas based on real-time displacement rate (e.g., 0.5 mm / d) and predicted rainfall. When the displacement rate of a GNSS monitoring point exceeds the dynamic threshold, the engine immediately triggers an orange warning.
[0064] S35: Warning information is pushed to the terminals used by relevant personnel via SMS, APP, etc. At the same time, the engine automatically generates a structured briefing, such as: "Warning location: Level 3 platform in Area A; Main risk: potential shallow landslides induced by rainfall; Recommended measures: immediately evacuate the equipment below, strengthen drainage in the area, and it is recommended to conduct a detailed drone survey after the rain."
[0065] S4: Closed-loop optimization: S41: After the warning is issued, on-site personnel take action and report the results. The system records the accuracy of the warning (whether it is a real emergency) and the effectiveness of the action.
[0066] S42: These feedback data, together with the new data collected subsequently, constitute new training samples. The threshold dynamic model of ST-MMFN and the early warning engine is incrementally learned and the parameters are fine-tuned regularly (e.g., monthly) so that the system evaluation and early warning capabilities can be continuously evolved.
[0067] The logic and / or steps represented in the flowchart or otherwise described herein may be specifically implemented in any readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0068] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A monitoring and early warning system for open-pit mine slopes, characterized in that, The system includes: The data preprocessing module preprocesses the multimodal data collected during the monitoring of open-pit mine slopes. The spatiotemporal sensing multimodal fusion analysis module extracts fusion feature vectors from preprocessed multimodal data to obtain stability quantification indicators, risk levels, and future deformation trends. This module is composed of a spatiotemporal sensing multimodal fusion network, which includes multiple specific encoders, spatiotemporal sensing weight allocation units, and progressive multiscale feature decoupling fusion units. The preprocessed multimodal data includes geological text data, sensor time-series data, and image data collected by UAVs. The visual risk identification module, implemented by a dual-stream visual recognition network, processes UAV remote sensing image data in the preprocessed multimodal data to obtain segmentation result data. The dynamic early warning and decision support module, based on the output results of the spatiotemporal perception multimodal fusion analysis module and the visual risk recognition module, combined with real-time environmental data and historical case data, calculates the early warning threshold, generates graded early warning information, and sends it to the terminal. The model closed-loop optimization and digital twin management module receives early warning and handling feedback and new monitoring data provided by the terminal, and optimizes the parameters of the spatiotemporal perception multimodal fusion analysis module and the dynamic early warning and decision support module.
2. The system according to claim 1, characterized in that, The multiple modality-specific encoders include: a modality-specific encoder based on a pre-trained language model, used to encode features from geological text data to obtain text features; a modality-specific encoder based on a temporal convolutional network, used to encode features from sensor time-series data to obtain time-series features; and a modality-specific encoder based on a deep convolutional network, used to encode features from image data to obtain image features.
3. The system according to claim 2, characterized in that, The spatiotemporal awareness weight allocation unit consists of a lightweight neural network. It generates a set of weight vectors based on the temporal context information and spatial location encoding corresponding to text features, temporal features, and image features. These weighted vectors are then summed to obtain a weighted feature: ,in, , , These respectively represent the features assigned to the text. Time series characteristics Image features The dynamic weighting coefficients.
4. The system according to claim 2, characterized in that, The progressive multi-scale feature decoupling and fusion unit receives text features, temporal features, and image features. It extracts feature sub-maps at different scales through multiple convolutional layers of different scales set in parallel to decouple them. Then, the obtained feature sub-maps are concatenated and fused through a convolutional layer. Finally, they are aggregated through a cross-scale connection method to obtain a fused feature vector.
5. The system according to claim 1, characterized in that, The dual-stream visual recognition network includes parallel edge enhancement channels and semantic parsing channels, followed by an attention-guided fusion layer. The edge enhancement channel extracts primary edge features from image data using multi-directional gradient operators; the semantic parsing channel extracts deep semantic features from image data based on an encoder-decoder architecture. The attention-guided fusion layer calculates the dependence weight of each spatial location in the deep semantic features on edge information using spatial attention. It then multiplies the primary edge features element-wise with the dependence weights and adds them to the deep semantic features to achieve the fusion of primary edge features and deep semantic features, resulting in a fused feature vector that serves as the basis for subsequent image segmentation. The spatial attention method is as follows: perform 1x1 convolution on deep semantic features to compress the number of channels to 1, and then generate normalized dependency weights through the Sigmoid activation function.
6. The system according to claim 5, characterized in that, At the end of the decoder of the semantic parsing channel, a small target sensitive detection head is set; the small target sensitive detection head includes a channel attention layer and a convolutional segmentation head; the channel attention layer adopts a squeeze-excitation network structure to recalibrate the shallow high-resolution features from the semantic parsing channel; the convolutional segmentation head fuses the recalibrated features with the fusion feature vector output by the attention-guided fusion layer, and outputs the small target segmentation result through the convolutional layer.
7. The system according to claim 1, characterized in that, The dynamic early warning and decision support module includes a dynamic feedback early warning unit based on digital twins. The dynamic feedback early warning unit constructs a digital twin of the slope and maps the output results of the spatiotemporal perception multimodal fusion analysis module and the visual risk recognition module to the digital twin. The unit has a built-in threshold dynamic model, which is used to calculate the early warning threshold for different areas based on real-time environmental parameters. The threshold dynamic model adopts an encoder-decoder architecture. The encoder consists of LSTM layers, and the decoder decodes based on an attention mechanism and outputs a warning threshold curve.
8. The system according to claim 4, characterized in that, The decoupling and concatenation method is as follows: the received text features, temporal features, image features, and weighted features are used as inputs to three convolutional layers of different scales, respectively, to obtain the following sub-maps in sequence: shallow feature maps, mid-level feature maps, and deep feature maps corresponding to text features; shallow feature maps, mid-level feature maps, and deep feature maps corresponding to temporal features; shallow feature maps, mid-level feature maps, and deep feature maps corresponding to image features; and shallow feature maps, mid-level feature maps, and deep feature maps corresponding to weighted features. During the stitching process: For each feature sub-map at each scale, feature sub-maps at the same scale from text features, temporal features, image features, and weighted features are stitched and fused across modally through a convolutional layer to obtain cross-modal fused features; finally, the cross-modal fused features at three different scales are aggregated through a cross-scale connection method to obtain a fused feature vector.
9. The system according to claim 5, characterized in that, The multi-directional gradient operator is a set of trainable convolution kernels, and its calculation formula is as follows: Where I represents the input image. For learnable convolutional kernel weights, Indicates the direction parameter. express Gradient operator for direction.
10. A method for monitoring and early warning of slopes in open-pit mines, characterized in that, This method is applied to the system as described in any one of claims 1-9, and the method includes: S1. Collect text data, time series data and image data related to the monitoring of unfilled mine slopes, and perform preprocessing; S2. Construct a spatiotemporal awareness multimodal fusion network and a dual-stream visual recognition network, and train the network model. S3. Deploy the trained spatiotemporal awareness multimodal fusion network and dual-stream visual recognition network on the edge computing server in the mine; construct a digital twin of the slope; initiate dynamic early warning and monitor environmental parameters in the mine area in real time; issue alarm information when an early warning is issued; S4. Closed-loop optimization: After the warning is issued, the accuracy of the warning and the effectiveness of the handling are recorded. These are then combined with the corresponding real-time collected data to form a new training sample, which is used to optimize the network parameters.