Landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning

By employing a global multi-module model combining self-supervised learning and multi-agent reinforcement learning, the problem of low accuracy in existing landslide identification and early warning methods is solved, enabling more efficient landslide environmental information analysis and early warning.

CN122135530APending Publication Date: 2026-06-02CHENGDU UNIVERSITY OF TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU UNIVERSITY OF TECHNOLOGY
Filing Date
2026-03-04
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing landslide identification and early warning methods have low accuracy and poor robustness.

Method used

A global multi-module model based on self-supervised learning and multi-agent reinforcement learning is adopted, including a landslide displacement prediction module, a landslide displacement adjustment module, a threshold decision module, and an early warning decision module. The global multi-module model is constructed through federated learning for landslide identification and early warning.

Benefits of technology

It improves the accuracy and robustness of landslide identification and early warning, and enables more efficient analysis of landslide environmental information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135530A_ABST
    Figure CN122135530A_ABST
Patent Text Reader

Abstract

This invention provides a landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning, belonging to the field of landslide identification and early warning technology. The method sequentially employs a landslide displacement prediction module, a landslide displacement adjustment module, a threshold decision module, and an early warning decision module from a global multi-module model constructed using self-supervised learning, multi-agent reinforcement learning, and federated learning. This allows for the analysis of landslide environmental information in the target area within the current time period, yielding the first predicted value, the second predicted value, the landslide early warning threshold, and the landslide early warning result for the target area. This achieves landslide identification and early warning for the target area. The global multi-module model used in the above process is more robust and accurate than a single landslide early warning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of landslide identification and early warning technology, and in particular to a landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning. Background Technology

[0002] Landslide deformation refers to the change in shape and position of a mountain or soil mass under the influence of gravity due to changes in internal stress. This deformation is usually the result of the combined effects of multiple factors, including geological structure, soil and rock properties, and hydrological conditions.

[0003] Existing methods often use a single landslide early warning model for landslide identification and early warning. However, due to the lack of robustness of a single landslide early warning model, the accuracy of existing methods for landslide identification and early warning is relatively low. Summary of the Invention

[0004] This invention proposes a landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning. The global multi-module model constructed by self-supervised learning, multi-agent reinforcement learning and federated learning is used to identify and warn of landslides in the target area within the current time period. Compared with a single landslide early warning model, it has higher robustness and higher accuracy in landslide identification and early warning.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, this invention provides a landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning, comprising: acquiring landslide environmental information of a target area within the current time period; the landslide environmental information includes radar images, geological data, topographic data, climate data, and groundwater data. Based on the landslide environmental information of the target area within the current time period and a landslide displacement prediction module, a first predicted value of the landslide displacement of the target area within a future time period is determined; the landslide displacement prediction module is constructed using a gPINN model. Based on the first predicted value, radar images of the target area within the current time period, and a landslide displacement adjustment module, a second predicted value of the landslide displacement of the target area within a future time period is determined; the landslide displacement adjustment module is constructed using a Time-Bridge model. Based on the second predicted value, the landslide environmental information of the target area within the current time period, landslide early warning thresholds of the target area within historical time periods, and a threshold decision module, a landslide early warning threshold for the target area within the current time period is determined; the threshold decision module is constructed using a first Time-MOE large model. Based on the second predicted value, the landslide warning threshold for the target area within the current time period, the landslide environmental information of the target area within the current time period, and the warning decision module, the landslide warning result for the target area within the current time period is determined. Among them, the warning decision module is constructed based on the second Time-MOE large model. The landslide displacement prediction module, landslide displacement adjustment module, threshold decision module, and warning decision module constitute a global multi-module model. The global multi-module model is constructed by the landslide displacement prediction module, landslide displacement adjustment module, threshold decision module, and warning decision module through self-supervised learning, multi-agent reinforcement learning, and federated learning.

[0006] In one implementation of the first aspect, the construction process of the global multi-module model is as follows: Step 1: Build multiple landslide simulation environments in multiple clients, including: For each of the multiple clients, landslide environment information of the area where the client is located within a historical time period is obtained, and a landslide simulation environment is constructed based on the landslide environment information of the area where the client is located within a historical time period. The landslide simulation environment simulates the topographic changes, weather changes, and groundwater changes in the area where the client is located within a historical time period.

[0007] Step 2: For each of the multiple clients, construct an initial landslide displacement prediction module, an initial landslide displacement adjustment module, an initial threshold adjustment module, and an initial early warning decision module within the client's landslide simulation environment. Perform multi-agent reinforcement learning on these modules within the landslide simulation environment. The initial landslide displacement prediction module is constructed based on a pre-trained gPINN model; the initial landslide displacement adjustment module is constructed based on a pre-trained Time-Bridge model; the initial threshold adjustment module is constructed based on a pre-trained first Time-MOE large model; and the initial early warning decision module is constructed based on a pre-trained second Time-MOE large model. The pre-training method is self-supervised learning.

[0008] Step 3: For each of the multiple clients, the client transmits the landslide displacement prediction module, landslide displacement adjustment module, threshold adjustment module, and early warning decision module obtained after the multi-agent reinforcement learning in the landslide simulation environment to the server.

[0009] Step 4: The server performs hierarchical aggregation of the landslide displacement prediction module, landslide displacement adjustment module, threshold adjustment module, and early warning decision module transmitted by each client from multiple clients to construct a global multi-module model.

[0010] In one implementation of the first aspect, determining the first predicted value of the landslide displacement of the target area within a future time period includes: inputting the landslide environment information of the target area within the current time period into the gPINN model; after analyzing the landslide environment information of the target area within the current time period, the gPINN model outputs the first predicted value of the landslide displacement of the target area within a future time period.

[0011] In one implementation of the first aspect, determining a second predicted value of landslide displacement in a target area within a future time period includes: inputting a first predicted value and a radar image of the target area within the current time period into a Time-Bridge model; adjusting the first predicted value using the radar image of the target area within the current time period; and outputting a second predicted value of landslide displacement in the target area within a future time period.

[0012] In one implementation of the first aspect, determining the landslide warning threshold for the target area within the current time period includes: inputting a second predicted value, landslide environmental information of the target area within the current time period, and landslide warning thresholds of the target area within historical time periods into a first Time-MOE large model; after adjusting the landslide warning thresholds of the target area within historical time periods based on the second predicted value and the landslide environmental information of the target area within the current time period, the first Time-MOE large model outputs the landslide warning threshold for the target area within the current time period.

[0013] In one implementation of the first aspect, determining the landslide early warning results for the target area within the current time period includes: The second predicted value, the landslide warning threshold for the target area within the current time period, and the landslide environment information for the target area within the current time period are input into the second Time-MOE large model. The second Time-MOE large model analyzes the second predicted value, the landslide warning threshold for the target area within the current time period, and the landslide environment information for the target area within the current time period, and outputs the landslide warning result for the target area within the current time period.

[0014] Compared with the prior art, the present invention has the following beneficial effects:

[0015] The landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided by this invention sequentially employs a landslide displacement prediction module, a landslide displacement adjustment module, a threshold decision module, and an early warning decision module from a global multi-module model constructed based on self-supervised learning, multi-agent reinforcement learning, and federated learning. This enables the analysis of landslide environmental information in the target area within the current time period, and sequentially obtains the first predicted value, the second predicted value, the landslide early warning threshold, and the landslide early warning result for the target area. This achieves landslide identification and early warning for the target area. The global multi-module model used in the above process is more robust and has higher accuracy in landslide identification and early warning compared to a single landslide early warning model. Attached Figure Description

[0016] Figure 1 This is one of the schematic diagrams of a landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided in the embodiments of this application; Figure 2 This is the second schematic diagram of the landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided in the embodiments of this application; Figure 3 This is the third schematic diagram of the landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided in the embodiments of this application; Figure 4 This is the fourth schematic diagram of the landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided in the embodiments of this application; Figure 5 This is the fifth schematic diagram of the landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided in the embodiments of this application; Figure 6 This is a schematic diagram of the federated learning process provided in an embodiment of this application; Figure 7 This is a schematic diagram of the modular multi-agent network architecture provided in the embodiments of this application; Figure 8 This is a schematic diagram of the Time-Bridge model provided in the embodiments of this application; Figure 9 This is a schematic diagram of the distribution of federated learning security policies provided in an embodiment of this application. Detailed Implementation

[0017] In the specification and claims of this invention, the terms "first" and "second," etc., are used to distinguish different objects, rather than to describe a specific order of objects.

[0018] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0019] The method and apparatus provided in this application relate to landslide identification and early warning. They can predict landslide environmental information of the target area within the current time period by constructing a global multi-module model through self-supervised learning, multi-agent reinforcement learning, and federated learning, thereby realizing landslide identification and early warning of the target area within the current time period.

[0020] To address the issue of low accuracy in landslide identification and early warning using existing methods in the background art, this application provides a landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning. The global multi-module model constructed through multi-agent reinforcement learning and federated learning identifies and warns of landslides in the target area within the current time period. Compared with a single landslide early warning model, this method is more robust and has higher accuracy in landslide identification and early warning.

[0021] like Figure 1 As shown, the landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided in this application includes S101-S105: S101. Obtain landslide environmental information for the target area within the current time period; The aforementioned landslide environmental information includes radar images, geological data, topographic data, climate data, and groundwater data; S102. Based on the landslide environment information of the target area and the landslide displacement prediction module in the current time period, determine the first predicted value of the landslide displacement of the target area in the future time period. The landslide displacement prediction module described above is constructed using the gPINN model. The input of the gPINN model is the landslide environment information of the target area within the current time period, and the output is the first predicted value of the landslide displacement of the target area within the future time period. The duration of the current time period and the future time period can be 15 days or 30 days, and the time difference between the current time period and the future time period can be 15 days or 30 days. This application embodiment does not limit the duration and time difference of the current time period and the future time period. Optionally, combined Figure 1 ,like Figure 2 As shown, S102 includes S1021; S1021. Input the landslide environment information of the target area within the current time period into the gPINN model. After analyzing the landslide environment information of the target area within the current time period, the gPINN model outputs the first predicted value of the landslide displacement of the target area within the future time period. S103. Based on the first predicted value, the radar image of the target area in the current time period, and the landslide displacement adjustment module, determine the second predicted value of the landslide displacement of the target area in the future time period. The landslide displacement adjustment module described above is constructed using the Time-Bridge model. The input of the Time-Bridge model consists of a first predicted value and a radar image of the target area within the current time period. The output is a second predicted value of the landslide displacement of the target area within a future time period. For example, in combination Figure 2 ,like Figure 3 As shown, S103 above includes S1031; S1031. Input the first predicted value and the radar image of the target area in the current time period into the Time-Bridge model. After adjusting the first predicted value with the radar image of the target area in the current time period, the Time-Bridge model outputs the second predicted value of the landslide displacement of the target area in the future time period. S104. Based on the second predicted value, the landslide environment information of the target area in the current time period, the landslide warning threshold of the target area in the historical time period, and the threshold decision module, determine the landslide warning threshold of the target area in the current time period. The threshold decision module described above is constructed using a first Time-MOE large model. The input to the Time-MOE large model consists of two predicted values, landslide environmental information of the target area within the current time period, and landslide warning thresholds of the target area within historical time periods. The output of the Time-MOE large model is the landslide warning threshold for the target area within the current time period. The length of the historical time period can be 15 days or 30 days, and the time difference between the current time period and the historical time period can be 15 days or 30 days. This embodiment of the application does not limit the duration and time difference between the current time period and the historical time period. In one application scenario, combined with Figure 3 ,like Figure 4 As shown, S104 above includes S1041; S1041. Input the second predicted value, the landslide environment information of the target area in the current time period, and the landslide warning threshold of the target area in the historical time period into the first Time-MOE large model. After adjusting the landslide warning threshold of the target area in the historical time period based on the second predicted value and the landslide environment information of the target area in the current time period, the first Time-MOE large model outputs the landslide warning threshold of the target area in the current time period. S105. Based on the second predicted value, the landslide warning threshold of the target area in the current time period, the landslide environmental information of the target area in the current time period, and the warning decision module, determine the landslide warning result of the target area in the current time period. The early warning decision module is constructed based on the second Time-MOE large model. The input of the Time-MOE large model includes the second predicted value, the landslide early warning threshold of the target area within the current time period, and the landslide environment information of the target area within the current time period. The output is the landslide early warning result of the target area within the current time period. For example, the landslide early warning result may indicate that a landslide has occurred or that a potential sliding surface exists. Furthermore, the first Time-MOE large model and the second Time-MOE large model can be any existing second Time-MOE large model. This application embodiment does not limit the type of the first Time-MOE large model and the second Time-MOE large model. Optionally, the input content of the second Time-MOE large model may also include decision prompt words, which are used to guide the second Time-MOE large model to analyze the second predicted value, the landslide early warning threshold of the target area in the current time period, and the landslide environmental information of the target area in the current time period. In some embodiments, combined with Figure 4 ,like Figure 5 As shown, the above S105 includes S1051; S1051. Input the second predicted value, the landslide warning threshold of the target area in the current time period, and the landslide environment information of the target area in the current time period into the second Time-MOE large model. The second Time-MOE large model analyzes the second predicted value, the landslide warning threshold of the target area in the current time period, and the landslide environment information of the target area in the current time period, and outputs the landslide warning result of the target area in the current time period. In this embodiment, the landslide displacement prediction module, the landslide displacement adjustment module, the threshold decision module, and the early warning decision module constitute a global multi-module model. This global multi-module model is constructed from the landslide displacement prediction module, landslide displacement adjustment module, threshold decision module, and early warning decision module through self-supervised learning, multi-agent reinforcement learning, and federated learning. The federated learning process is as follows: Figure 6 As shown.

[0022] In one implementation, the construction process of the global multi-module model is as follows: Step 1: Build multiple landslide simulation environments in multiple clients, including: For each of the multiple clients, landslide environment information of the client's location within a historical time period is obtained, and a landslide simulation environment is constructed based on the landslide environment information of the client's location within the historical time period; the landslide simulation environment simulates the topographic changes, weather changes, and groundwater changes in the client's location within the historical time period. Step 2: For each of the multiple clients, construct an initial landslide displacement prediction module, an initial landslide displacement adjustment module, an initial threshold adjustment module, and an initial early warning decision module within the client's landslide simulation environment. Perform multi-agent reinforcement learning on these modules within the landslide simulation environment. Specifically, the initial landslide displacement prediction module is constructed based on a pre-trained gPINN model; the initial landslide displacement adjustment module is constructed based on a pre-trained Time-Bridge model; the initial threshold adjustment module is constructed based on a pre-trained Time-MOE large model; and the initial early warning decision module is constructed based on a pre-trained Time-MOE large model. The pre-training method is self-supervised learning.

[0023] In one implementation, the specific process of the above self-supervised learning is as follows: Step 2.1: Service client self-monitored initialization; The self-supervised learning component of this application employs a large model based on an improved Time-MOE model structure as the backbone network to extract high-quality, representative feature representations in unlabeled environments. This model structure combines masking reconstruction, self-attention mechanisms, and asymmetric encoder design, making it suitable for learning tasks involving multi-source heterogeneous time-series data in a federated environment. 1. Input data preprocessing and masking strategy design Data input: The raw multidimensional time series data collected from data sources in various countries are first standardized and then normalized using the DYT (Dynamic Y-aware Transformation) method to improve the consistency and stability of cross-distribution data; Masking mechanism: Randomly mask a certain proportion of time steps or channel dimensions in the input sequence, and guide the model to learn the latent structure and temporal dynamics of the data during the pre-training stage through the loss function; 2. The improved structure is shown in Table 1 below: Table 1. Schematic diagram of the improved structure ; 3. Pre-training process In a federated scenario where labels are not required, a global self-supervised pre-training is performed using a large amount of unlabeled data provided by the data provider to extract a general representation. First, the service client utilizes local unlabeled data to perform self-supervised learning through an improved model structure called Time-MoE (Mixture of Experts for Temporal Representation Learning). This model integrates a temporal masking reconstruction mechanism, an expert dynamic selection (MoE) structure, and temporal consistency modeling, making it particularly suitable for processing multi-source heterogeneous time series data in a federated environment. Initialize M expert sub-models: each expert Different structures can be used (LSTM, 1DConv, Transformer block). Initialize the gating network Typically an MLP or a lightweight Transformer, the output is an expert-assigned probability. ; Initialize the aggregator: a linear fusion layer or an attention aggregator; During the pre-training phase, the model randomly masks parts of the input time series and then attempts to reconstruct these masked time slices through a multi-expert network in order to learn the potential dynamic features and spatiotemporal evolution of the data. Different expert modules are responsible for processing features of different types or time scales, thereby forming a multi-granular temporal representation. Through this unsupervised pre-training method, Time-MoE can automatically extract high-quality temporal representation features without manual annotation, providing a strong feature foundation for subsequent landslide identification, early warning and prediction tasks; when transferred to downstream tasks, the model only needs a small amount of labeled data for fine-tuning to quickly adapt to new regions or new scenarios, significantly improving the model's generalization performance and stability in low-sample and cross-domain scenarios. The server-side convergence and evaluation results are shown in Table 2 below: Table 2. Schematic diagram of server-side convergence and evaluation results. ; 4. Publish a secure communication and privacy policy (Security & Privacy Deployment) (1) Homomorphic Encryption (HE) Configuration: The server uses the CKKS approximate homomorphic encryption algorithm to generate an asymmetric key, as shown in the following formula: ; in For public key, This is the server-side private key; (2) Public key distribution: Distribute the public key Included in the initialization package and distributed to all clients, it is used by clients to encrypt gradient or model updates; (3) Private key retention: private key Strictly stored in protected memory on the server side, used to decrypt encrypted global updates during the aggregation phase; (4) Differential privacy: The server configures a DP protection strategy for the client's uploaded gradient: setting a noise budget. This determines the upper limit of the amount of information that can be leaked during model training; (5) Clipping threshold Set the gradient clipping threshold The client needs to restrict the gradient norm to within a certain range before encryption. within In conjunction with Gaussian noise injected on the server side, it prevents reconstruction attacks targeting specific samples; ; in For the client gradient vector, Gradient clipping threshold; when the gradient is too large ( If it is, then scale it to the maximum length. When the gradient is small ( ), remain unchanged; 5. Broadcast the initial model parameters to the client, using the following formula: ; in , Including A network of experts Gating network and aggregator ; Step 2.2: Self-supervised pre-training for each client In this embodiment, the client is mainly responsible for local data processing and model updates: first, it performs self-supervised pre-training using local unlabeled data to obtain a high-quality feature encoder; then, based on its own computing power, bandwidth, and other conditions, it dynamically determines the training rounds and communication frequency through multi-agent reinforcement learning; subsequently, it fine-tunes the model on a small amount of labeled data, and uploads the updated parameter increments to the server after homomorphic encryption, while protecting privacy through differential privacy and unintentional transmission mechanisms. 2.2.1 Client Initialization (1) The client loads the local model, receives the current global model parameters, and loads the global weights into the local model, as shown in the following formula: ; (2) Obtain the server's HE public key Encryption used for uploading digests / increments); 2.2.2 Client from local dataset Randomly sampled small batches of data

[0024] For each sample Data augmentation operations can be performed, such as Gaussian distribution and random pruning; that is, augmenting the same original sample locally. Two separate random data augmentations were applied; the client used a local time-series dataset. Randomly sample small batches of samples: ; And enhance the pipeline through data Generate two time series from different perspectives: ; in These enhancements may include: temporal masking, Gaussian noise, and random pruning; these enhancements are designed to simulate observational uncertainty, enabling the model to learn robust features of temporal patterns. 2.2.3 Mask Reconstruction and Expert Dynamic Selection The client employs a masking-time reconstruction mechanism during self-supervised training; temporal masking is performed on each window sequence. 1. Randomly select the masking ratio ; 2. Generate a set of masking fragments As shown in the following formula: ; 3. Record the masking code for monitoring signals during reconstruction; After the input sequence is randomly masked, the Time-MoE model performs the following calculations: ; in: Indicates the first A network of experts. It is the expert weight distribution output by the gating network. It refers to the number of experts; The model learns the latent structure by minimizing the reconstruction error, as shown in the following formula; ; 2.2.4 Comparison of Losses To ensure the model remains stable under temporal perturbations, a time-contrastive learning objective is employed: ; in For the temporal representation of the two augmented samples, For cosine similarity, For temperature parameters; 2.2.5 Expert Equilibrium Regularization and Sparse Activation ; Fine-tuning phase: The client performs supervised fine-tuning on a small amount of labeled data to optimize the loss. ; Hyperparameters and As weight; 2.2.6 Local Update Perform several local steps to update student parameters ; AdamW is used to update local parameters, and the update process is shown in the following formula: ; Local private gating can also be updated. ; 2.2.7 Differential Privacy Clipping and Noise Addition Gradient clipping, as shown in the following formula: ; Add Gaussian noise as shown in the following formula: ; 2.2.8 Client Upload Stage Using the HE public key After gradient encryption, the data is sent to the server, as shown in the following formula: ; At the same time, the client enters a waiting state; Step 2.3: Server-side global model aggregation The server receives encrypted updates from K clients, as shown in the following formula: ; 2.3.1 Receive and decrypt encryption gradients (client → server) The server collects uploads from K clients, as shown in the following formula: ; Perform HE decryption as shown in the following formula: ; 2.3.2. Similarity Calculation The server receives a collection of knowledge summaries from each client; ; Calculate the cosine similarity between the expert embedding for each client and the global expert; ; 2.3.3. Dynamic Weight Allocation Based on the similarity distribution, the aggregation weight is adaptively adjusted so that clients with high similarity account for a larger proportion of the expert's updates; ; in: These are temperature control parameters; when The time approaches the average; when Only retain knowledge from the most similar client; 2.3.4 Server Aggregation Strategy (Aggregate) During the aggregation phase, the server merges the expert update directions uploaded by each client, as shown in the following formula: ; in: For learning rate, Indicates client For experts The degree of contribution or the frequency of actual use of the expert by the client; The updated system provides experts with a globally shared knowledge structure, while retaining adaptability to regional differences. 2.3.5 Convergence Monitoring Mechanism ① The rate of reduction in reconstruction error is shown in the following formula: ; ② Whether the entropy of expert utilization rate tends to be uniform, as shown in the following formula: ; ③ The parameter update norm is shown in the following formula: ; If all conditions are met, the distillation phase is triggered; if convergence is not achieved, the next round of federated training begins. ; Step 2.4 Global Distillation Stage (Teacher–Student Deployment) After global Time-MoE convergence, the server performs distillation, the distillation process as follows: Figure 9 As shown; 2.4.1 Server-side distillation preparation 1. Freeze the Teacher model parameters, as shown in the following formula: ; 2. Constructing a sparse sub-model definition For each expert, the set of experts in the global model is defined. We use mapping functions Generate its corresponding low-dimensional distillation embedding This process utilizes the teacher-student paradigm to compress complex sequential logic into a lightweight encoder. From the parameter space or prototype vector; filter out One active expert constructs a sparse sub-model: ; in, It includes shared layer parameters (Input Layer includes PhysAdapter, AttentionLayers, RMSNorm, Prediction Head). Represented as the first The sparse sub-model is a set of active expert indexes selected by each client; compared to the full model, the number of parameters in this sparse sub-model is significantly reduced (approximately 5%-10% of the full model). The server packages the following parameters: Shared Layers: Input Layer (including PhysAdapter), Attention Layers, RMSNorm, Prediction Head; 3. Distribute distillation strategy templates Including: intermediate layer feature distillation, Logits distillation (soft targets), expert selection probability hint, and temporal consistency hint. 2.4.2 Client-side distillation training process After each client receives the Teacher (frozen) and Student (trainable) models: (1) Construct the Teacher feature and Logits as shown in the following formula: ; (2) Student forward calculation, as shown in the following formula: ; (3) Multi-objective distillation loss Logits, the distillation mixing loss (Hinton KD), are shown in the following formula: ; Characteristic distillation (L2 / CKA), as shown in the following formula: ; Expert Assignment Distillation, as shown in the following formula: ; The final distillation loss is shown in the following formula: ; (4) Local Student Training The client only updates the Student model, as shown in the following formula: ; Similarly, DP+HE protection is required to differentiate uploaded parameters; 2.4.3 Server-side Student Aggregation Server-to-server sparse aggregation; Shared layer: The network layer that is common to all clients. The weighted average of the performance standards is shown in the following formula: ; Expert layer: Selective aggregation for specific experts (like Only the set of clients that activated and used the expert in this round of training are aggregated. If a certain expert This round was not selected by any client. If the parameters remain unchanged, as shown in the following formula: ; Repeat the iteration until convergence; 2.4.4 Final Deployment Phase After aggregation is complete, the server performs the following on the global model: (1) Convergence detection (e.g., the decreasing trend of global loss); (2) Robustness detection (checking whether malicious clients upload infected parameters, and correcting or discarding abnormal updates through aggregation); (3) The server distributes the final lightweight Student model to all clients for use in downstream client tasks; 2.4.5 During the distillation phase, the client may need to obtain expert selection hints or soft labels from the teacher model, but it cannot expose its own request content, nor can it let the server know the selection details. Therefore, OT services are used to ensure secure interaction. OT startup process: A. Server preprocessing (after each model update) includes the following steps; Teacher division: generation (Subcontracted by expert type); Generate a symmetric key for each packet And calculate the encrypted packet ; Each Stored in a downloadable location on the server, and recording packet metadata. (Such as package type, size, permissions, and usage limits); key set The server input is added to OT (meaning OT will offer these keys as "options" for the client to choose from in subsequent sessions). The packet metadata table (plaintext) is published for clients to view, but does not contain the key or packet plaintext; B. Client Request (Private Option); The client determines which type of hint package is needed locally based on the task / data (e.g., if you want the expert 3 hint, then you need a set of corresponding package indexes). The client initiates a 1-of-N OT (or multiple 1-of-N OT) request, and the OT protocol ensures that the server is completely unaware of the selected index; OT returns the corresponding symmetric key to the client. ; Step 3: For each of the multiple clients, the client transmits the landslide displacement prediction module, landslide displacement adjustment module, threshold adjustment module, and early warning decision module obtained after the multi-agent reinforcement learning in the landslide simulation environment to the server. Step 4: The server performs hierarchical aggregation of the landslide displacement prediction module, landslide displacement adjustment module, threshold adjustment module, and early warning decision module transmitted by each client from multiple clients to construct a global multi-module model; In one application scenario of the above implementation, the multi-agent reinforcement learning and collaborative training mechanism is as follows; (1) Reuse of distillation features and state space construction Each client loads the global distillation model. And using it as a feature extractor (PartialObservations), high-level feature mapping is performed on local landslide environment data to construct the state space of multi-agent reinforcement learning: (Radar imagery, topographic factors, rainfall data, groundwater level); in, For the client at time Representation of environmental state; This is the optimized encoding function after distillation; The sample data varies over time and includes radar images, terrain factors, rainfall data, and groundwater levels. This feature space maintains consistency across different clients to ensure a uniform state distribution in subsequent federated reinforcement learning. To further enhance feature stability, momentum smoothing is introduced in the state space for each client; ; To ensure a smooth transition of states over time and reduce noise sensitivity, a geomechanical model based on physical constraints (such as the Mohr-Coulomb criterion and stability coefficients) is introduced to impose physical consistency constraints on the state transition process, thereby constructing a virtual environment and ensuring that its dynamic evolution is consistent with real geological processes. (2) Multi-agent partitioning, agent structure and corresponding process To achieve collaborative perception and adaptive threshold control among clients in different landslide areas, this system establishes a modular multi-agent network architecture, such as... Figure 7 As shown; refer to Figure 7 Each client corresponds to one intelligent agent. It interacts with neighboring agents in a distributed communication graph; each agent contains four core modules. The module includes a Belief Module (corresponding to the landslide displacement prediction module), a Threshold Decision Module (corresponding to the threshold decision module), a Communication Module, and an Observation & Prediction Module (corresponding to the landslide displacement adjustment module and the early warning decision module). The modules work together to predict landslide conditions, adaptively adjust thresholds, and optimize regional early warning systems. The information flow process (data flow) is shown below; Obtaining partial observations from the environment. →Send to the data processing module to extract features; →Feature input BrLSTM Unit updates belief; BrLSTM Unit output values ​​to: The results are fed into the Prediction Module to generate predictions. The message is encoded into a communication module. The data is then sent to the decision unit for processing. The Communication Module sends messages to other agents. Other agents also send their messages to Agent i; →Agents i integrates neighbor messages, its own observations, and predictions to update its beliefs; Ultimately, all agents obtain a unified Predicted Class (decision result) through consensus mechanism or distributed averaging. (3) The temporal evolution of agent beliefs; Each agent at time step Maintaining its own belief state vector and contextual information: ; Each agent At any moment It has the following variables: Hidden state represents the agent's temporal memory; Hidden state represents the agent's temporal memory; The input information vector at each time step consists of three parts: Input information It consists of three parts: Local features extracted from InSAR and multi-source meteorological observation data; : An aggregate vector of communication messages between neighboring intelligent agents; Location and geocoding vector; ; ; Wherein, belief state = the agent's "best guess" of the true state of the current environment based on historical observations, communication information and model knowledge; Updating the belief state means integrating the "newly received information" into the "old belief" to form a "more accurate new belief." In a multi-agent system, there are two sources of information: local new observations (sensors, images, radar, etc.) and neighbor messages (belief summaries sent by other agents). In this embodiment, the trainable nonlinear mapping function combines an LSTM unit with a Time-Bridge model. The Time-Bridge model structure is as follows: Figure 8 As shown; by combining historical state and semantic information, a higher level of spatiotemporal dependency modeling and decision interpretability can be achieved; The following describes the fusion mechanism between LSTM units and the Time-Bridge model; (1) Standard LSTM Belief Update Mechanism At time step The agent receives the input feature vector: ; in: Local images and sensor observations of CNN features Average neighbor communication messages Location coding and geological background parameters; LSTM cell execution state update: ; in: Indicates a hidden state (short-term trend characteristics); Indicates cell state (long-term geological trends and cumulative displacement characteristics). (2) Time-Bridge-LSTM fusion mechanism This system introduces a Time-Bridge mechanism on the basis of traditional LSTM, and uses a cross-time attention structure to dynamically aggregate the hidden states of key historical time slices to enhance the ability to recognize future trends. First, construct cross-time attention convergence, as shown in the following formula; ; The attention weights for the above-mentioned cross-temporal attention convergence are shown in the following formula: ; in, Indicates the length of the time window; Represents a learnable mapping matrix; A dynamic weighted aggregation representing key historical states; (3) Belief Enhancement and Integration The Time-Bridge output is concatenated and merged with the current LSTM output, as shown in the following formula: ; in, To reach the final enhanced belief state, It is a non-linear activation function (ReLU / GELU). This is a feature splicing operation; and After similar operations, we obtained ; The effects of belief states are as follows; Updated and Simultaneously participate in the following three types of tasks: Prediction Module: Provides input to the gPINN model for displacement prediction; Communication Module: Encoded as shared messages ; Decision Module: Inputs are fed into the large model decision unit for threshold-based decision-making; (4) Threshold decision and action strategy generation Setting up intelligent agents At any moment The input is: These inputs are fed into the threshold decision unit, which uses the Time-MOE large model as the core inference engine to generate action strategies and threshold adjustment schemes based on the current belief state. (4.1) Threshold and action joint strategy generation; Action decision-making consists of two parts: Threshold adjustment strategy: The Time-MOE large model generates a new landslide early warning threshold based on current observation features, landslide probability prediction, and neighbor status. ,Right now: ; in The inference function representing a large model can be determined based on the environmental context and the hidden state of beliefs. Adjust the optimal output threshold; Action strategy generation: The agent needs to perform physical or logical actions according to the environmental state in order to coordinate the global behavior of multiple agents; Position (pose) Updated according to known dynamics, where position dynamics refer to the agent's current position information. How to perform a certain action In the next moment, it will change to a new position. ; ; action From a finite set of actions Chinese strategy sampling: ; in This unit represents the hidden state output of the Time-MOE large model and shares the same input as the Belief BrLSTM. However, the parameters are independent to ensure the decoupling of threshold judgment and action selection; (4.2) Joint optimization and feedback adjustment After performing an action, the agent updates its belief state and threshold parameters based on environmental feedback (such as changes in landslide displacement or risk probability). This update process is optimized using a reinforcement learning strategy, enabling both threshold adjustment and action decision-making to simultaneously minimize prediction error and response latency. The formula is as follows: ; in, For the ideal threshold, For instant rewards, The learning rate is used to achieve adaptive collaborative optimization of the global landslide early warning system through this mechanism, enabling each agent to continuously learn the optimal threshold-action joint strategy in a dynamic environment. (5) Inter-agent communication and message aggregation Communication topology and basic symbols: Set of intelligent agents: ; Directed edge set: ,like Then the intelligent agent Able to Send a message; In-degree (number of neighbors): ; Encoder: Decoder: ; A communication graph is defined as a directed graph. Each intelligent agent at any time Conceal one's own beliefs Encode it into a low-dimensional message and broadcast / unicast it to your neighbors: ; Meanwhile, this application introduces a graph attention mechanism (GAM) to adaptively assign weights to neighbor information, thereby more effectively modeling spatial heterogeneity and geological zoning differences. To distinguish the importance of neighbor information, the graph attention mechanism is used to calculate and aggregate neighbor weights: (5.1) Calculate the neighbor correlation score: ; in, It is a linear transformation matrix. Indicates splicing, These are trainable vectors; Measurement from neighbors Information The relative importance; (5.2) Normalized attention weights: ; From the above equation, it can be seen that after softmax normalization... and To achieve a weighted average; (5.3) The decoded information from aggregated neighbors is shown in the following formula: ; (6) Model-related information for each model (6.1) Observation & Prediction Module Observation Composed of image patches of the landslide area and their corresponding geological attributes: ; in, : The cropped InSAR image (two-dimensional matrix); : Local rainfall or meteorological time series (vector); Groundwater level or seepage observation (scalar or vector); Fixed terrain / geological feature encoding (location-related constant vector); processed by a feature extraction network (Composed of VIM backbone trained from a large model) Output local feature vectors ; (6.2) Physical Model and PINN Representation In landslide displacement prediction tasks, the displacement field is assumed to be... Satisfying a certain type of physical constraint (simplified example): ; It can include residual expressions for equilibrium equations, particle acceleration-stress relationships, or seepage-mechanical coupling equations; gPINN uses a parameterized neural network. Represent the displacement field and learn it through physical residual constraints: For intelligent agents For local prediction, the loss function of gPINN is defined as a weighted sum of data terms and physical terms: ; in: These are observation points (partially labeled or pseudo-observations). For the number of data points; These are the sampling points used for calculating the physical residuals. This represents the number of residual sampling points; For hyperparameters; For gPINN model parameters; For the client gPINN is giving a vision of the future at this moment. The preliminary time prediction is shown below; ; (6.3) Time-Bridge temporal correction and multi-scale enhancement gPINN output Although physically consistent, it may neglect local short-term nonlinear disturbances (such as sudden rainfall, local soil fracturing, etc.); Time-Bridge is responsible for: using the most recent observation sequence and local features to... Perform time series correction; integrate multi-scale time memories (short-term fluctuations + medium-term trends); output a final prediction with confidence level. ; The Time-Bridge is represented by a conditional correction function: ; in: For Time-Bridge parameters; Indicates recent Feature windows for each time step; The agent's current belief (including neighborhood information); In implementation, Time-Bridge can employ a Transformer- or hybrid gating architecture (including time and position embedding). The specific calculation example is as follows; The temporal embedding and normalization of the input sequence are shown in the following equation; ; in, Encoding for time location, For splicing or addition; The sequence context obtained by multi-head self-attention encoding is shown in the following equation; ; The gPINN output is concatenated with the context and passed through a residual network to obtain the correction term as shown in the following equation; ; ; in, It can be a two-layer MLP or Transformer decoder; Sequence pooling (such as average / attention-weighted pooling); Time-Bridge training aims to balance the adjusted results between data rigor, temporal smoothness, and physical consistency. The training loss is defined as follows: ; in, The following formula represents the fidelity of the adjusted prediction on the physical residual (which can be the physical residual of PINN); ; The specific process is as follows: Phase 1 (Pre-training): First, train gPINN (parameters) Perform self-supervised / physically supervised training to minimize ; Phase Two (Time-Bridge Training): Fix or fine-tune gPINN, train the Time-Bridge (parameters) ) minimize ; Phase Three (Joint Fine-tuning): End-to-End Fine-tuning The common minimum of the comprehensive loss is shown in the following equation; ; in Weigh the importance of the two parts; (7) Construction of distributed prediction and consensus mechanism This system integrates local predictions from multiple clients (agents) into a consistent risk assessment result for the region or the entire network. It employs a decentralized consensus mechanism to achieve stable, interpretable, and robust prediction fusion while ensuring decentralization, limited communication bandwidth, and privacy protection. (7.1) In After rounds of communication and prediction, each agent generates an original prediction vector, which is then processed... After rounds of observation, belief updates, and communication, each agent... The final cell state is shown in the following formula: ; Using the prediction module The agent maps this state to The prediction vector (logits) of dimension is shown in the following formula; ; in, : Number of categories (risk level and displacement velocity); : indicates the first Unnormalized confidence score of the class; (7.2) Distributed Average Consensus Without relying on a server, each agent approximates the global average prediction through an iterative information exchange process, as shown in the following equation: ; in: Initial prediction by the agent; Intelligent agent The set of neighbors; Step size parameter; : Consensus iteration count; in a strongly connected communication graph where the step size satisfies ( When the maximum in-degree is reached, the above process converges to the global average. ; (7.3) Final prediction and decision-making stage After consensus is reached, each agent receives the same global prediction vector. After Softmax normalization, the category with the highest probability is taken as the final decision output: ; Softmax transformation transforms the logits vector Transform into a probability distribution; The final predicted category is (e.g., landslide risk level "high / medium / low"); all agents share this result to form a globally consistent early warning conclusion; (8) Reinforcement learning and differentiable reward optimization A differentiable reward reinforcement learning (DRL) strategy is employed to optimize landslide prediction accuracy and threshold response capability; reward function: ; in, For the prediction error term, This indicates the accuracy rate of local early warnings. Indicates the threshold oscillation stationarity; To achieve end-to-end joint optimization of the prediction module (PINN, Time-Bridge), threshold adjustment module, and early warning decision module, this invention introduces reinforcement learning (RL) and a differentiable reward mechanism into the multi-agent system. This mechanism allows the reward signal to directly backpropagate gradients to the neural network parameters (including CNN, LSTM, and communication module), thereby achieving adaptive learning across the entire "prediction-communication-decision" chain. (8.1) Optimization objective and definition of differentiable reward Each agent in the time series The goal is to maximize the cumulative expected reward: ; in This represents globally learnable parameters (including module parameters for each agent). Differentiable reward function definition: ; in: :go through The average logits obtained by consensus among all agents after the step; : The one-hot vector corresponding to the true class; the negative sign makes minimizing the prediction error equivalent to maximizing the reward; This reward is continuously differentiable and can influence the prediction result. Each dimension generates a smooth gradient signal, thereby supporting backpropagation; The optimization goal is The generalized policy gradient is shown in the following equation: ; The second term is added because the reward is differentiable, allowing the gradient to be directly propagated back to the communication parameters. Depends on ,and It is jointly generated by the LSTM, communication, and prediction modules of each agent, as shown in the following formula: ; This enables the gradient signal to be directly fed back to the parameters of the prediction and feature extraction modules, including: LSTM module: Optimizes the time-based belief update process; Communication module: Enhances information fusion and collaborative consistency; In actual training, to prevent differentiable reward terms from causing gradient oscillations or excessive variance, an unbiased surrogate objective is added, as shown in the following equation: ; in This represents the constant reward separated from the computation graph, used to construct an unbiased estimate; The final update method adopted is the following combination: ; The first term in the above formula For stable training, the second item It maintains end-to-end differentiability; this hybrid form simultaneously guarantees unbiasedness and convergence stability in practice. (8.2) Gradient Update and Joint Optimization Combining the aforementioned modules (PINN prediction and Time-Bridge adjustment module): PINN module (prediction constraints): Prediction output Participate in consensus, gradient from Backpropagation drives PINN parameter updates to match actual displacement evolution; Time-Bridge Module (Predictive Correction): As a timing adjuster, the output of the Time-Bridge directly affects... Therefore, it also receives from The gradient prompts the model to automatically correct for time bias; Communication and Decision Module: Communication weights and message encoding parameters are determined via... Update to achieve better collaborative prediction; Ultimately, the parameter update for the entire system can be written as: ; in, The learning rate; In some embodiments, within the reservoir bank landslide monitoring area, each monitoring point deploys an agent (Agent i), corresponding to a local monitoring unit; the core modules contained in each agent are shown in Table 3 below: Table 3 Core Module Table

[0025] In summary, the landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning provided in this application embodiment sequentially employs a landslide displacement prediction module, a landslide displacement adjustment module, a threshold decision module, and an early warning decision module from a global multi-module model constructed based on self-supervised learning, multi-agent reinforcement learning, and federated learning. This enables the analysis of landslide environmental information in the target area within the current time period, and sequentially obtains the first predicted value, the second predicted value, the landslide early warning threshold, and the landslide early warning result for the target area. This achieves landslide identification and early warning for the target area. The global multi-module model used in the above process is more robust and has higher accuracy in landslide identification and early warning compared to a single landslide early warning model.

[0026] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0027] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A landslide identification and early warning method based on self-supervised learning and multi-agent reinforcement learning, characterized in that, include: Obtain landslide environmental information for the target area within the current time period; The landslide environmental information includes radar images, geological data, topographic data, climate data, and groundwater data; Based on the landslide environment information of the target area within the current time period and the landslide displacement prediction module, a first predicted value of the landslide displacement of the target area within a future time period is determined; the landslide displacement prediction module is constructed using the gPINN model. Based on the first predicted value, radar images of the target area in the current time period, and the landslide displacement adjustment module, a second predicted value of the landslide displacement of the target area in the future time period is determined; the landslide displacement adjustment module is constructed using the Time-Bridge model. Based on the second predicted value, the landslide environment information of the target area in the current time period, the landslide warning threshold of the target area in the historical time period, and the threshold decision module, the landslide warning threshold of the target area in the current time period is determined; The threshold decision module is constructed using the first Time-MOE large model; Based on the second predicted value, the landslide warning threshold of the target area within the current time period, the landslide environmental information of the target area within the current time period, and the warning decision module, the landslide warning result of the target area within the current time period is determined; wherein, the warning decision module is constructed based on the second Time-MOE large model; the global multi-module model is constructed by the landslide displacement prediction module, the landslide displacement adjustment module, the threshold decision module, and the warning decision module after self-supervised learning, multi-agent reinforcement learning, and federated learning.

2. The method as described in claim 1, characterized in that, The construction process of the global multi-module model is as follows: Step 1: Build multiple landslide simulation environments in multiple clients, including: For each of the multiple clients, landslide environment information of the area where the client is located within a historical time period is obtained, and a landslide simulation environment is constructed using the landslide environment information of the area where the client is located within the historical time period; the landslide simulation environment simulates the topographic changes, weather changes, and groundwater changes in the area where the client is located within the historical time period. Step 2: For each of the multiple clients, construct an initial landslide displacement prediction module, an initial landslide displacement adjustment module, an initial threshold adjustment module, and an initial early warning decision module within the landslide simulation environment of the client. Perform multi-agent reinforcement learning on the initial landslide displacement prediction module, the initial landslide displacement adjustment module, the initial threshold adjustment module, and the initial early warning decision module within the landslide simulation environment. Specifically, the initial landslide displacement prediction module is constructed based on a pre-trained gPINN model; the initial landslide displacement adjustment module is constructed based on a pre-trained Time-Bridge model; the initial threshold adjustment module is constructed based on a pre-trained first Time-MOE large model; and the initial early warning decision module is constructed based on a pre-trained second Time-MOE large model. The pre-training method is self-supervised learning. Step 3: For each of the multiple clients, the client transmits the landslide displacement prediction module, landslide displacement adjustment module, threshold adjustment module, and early warning decision module obtained after the multi-agent reinforcement learning in the landslide simulation environment to the server. Step 4: The server performs hierarchical aggregation of the landslide displacement prediction module, landslide displacement adjustment module, threshold adjustment module, and early warning decision module transmitted by each of the multiple clients to construct the global multi-module model.

3. The method as described in claim 1, characterized in that, The determination of the first predicted value of landslide displacement in the target area within a future time period includes: The landslide environment information of the target area within the current time period is input into the gPINN model. After analyzing the landslide environment information of the target area within the current time period, the gPINN model outputs the first predicted value of the landslide displacement of the target area within the future time period.

4. The method as described in claim 1, characterized in that, The second predicted value for determining the landslide displacement of the target area within a future time period includes: The first predicted value and the radar image of the target area within the current time period are input into the Time-Bridge model. The Time-Bridge model adjusts the first predicted value based on the radar image of the target area within the current time period and outputs a second predicted value of the landslide displacement of the target area within the future time period.

5. The method as described in claim 1, characterized in that, Determining the landslide early warning threshold for the target area within the current time period includes: The second predicted value, the landslide environment information of the target area in the current time period, and the landslide warning threshold of the target area in the historical time period are input into the first Time-MOE large model. The first Time-MOE large model adjusts the landslide warning threshold of the target area in the historical time period based on the second predicted value and the landslide environment information of the target area in the current time period, and then outputs the landslide warning threshold of the target area in the current time period.

6. The method as described in claim 1, characterized in that, The determination of landslide early warning results for the target area within the current time period includes: The second predicted value, the landslide warning threshold of the target area within the current time period, and the landslide environment information of the target area within the current time period are input into the second Time-MOE large model. The second Time-MOE large model analyzes the second predicted value, the landslide warning threshold of the target area within the current time period, and the landslide environment information of the target area within the current time period, and outputs the landslide warning result of the target area within the current time period.