Multi-branch network and cross-attention multi-source data fusion slope displacement prediction method

CN121071345BActive Publication Date: 2026-09-22BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511110307.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-09-22
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

[0004]一方面,传统预测方法往往仅依赖于单一数据源,或者只是简单地将多个数据源进行拼接或加权融合,未能深入挖掘不同数据源之间的内在联系和互补信息

Benefits of technology

[0050]本发明提供的多分支网络和交叉注意力多源数据融合边坡位移预测方法通过构建多分支网络对不同模态数据进行针对性的特征提取,利用交叉注意力机制对多源数据进行深度融合,实现多源数据的不同模态特异性特征提取、跨模态关联建模及多尺度预测,能够充分挖掘多源数据中的潜在信息,有效提高了边坡位移预测的准确性和可靠性,为地质灾害预警和防治提供了更加有力的技术支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121071345B_ABST
    Figure CN121071345B_ABST
Patent Text Reader

Abstract

The present application provides a multi-branch network and cross-attention multi-source data fusion slope displacement prediction method, comprising: collecting meteorological data, GNSS node data, radar three-dimensional data, pre-processing and standardizing the multi-source data; constructing a multi-branch network to extract features of each type of data respectively; fusing the feature vectors of the multi-source data through a cross-attention mechanism; using a multi-scale memory network combined with LSTM to capture recent trends, a one-dimensional dilated convolutional neural network to extract periodic features, and self-adaptive fusion to obtain displacement prediction values; using a course learning mechanism to train the model and output the displacement prediction values. The present application constructs a multi-branch network to extract features of different modal data, uses a cross-attention mechanism to deeply fuse multi-source data, realizes multi-source data specific feature extraction of different modalities, cross-modal correlation modeling and multi-scale prediction, fully excavates the potential information in multi-source data, and improves the accuracy and reliability of slope displacement prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of slope displacement prediction technology, and in particular, to a slope displacement prediction technology that integrates multimodal data features and strengthens cross-source correlation modeling; specifically, it relates to a slope displacement prediction method that integrates multi-branch networks and cross-attention multi-source data. Background Technology

[0002] In the field of geological disaster prevention and control, slope displacement prediction is a crucial link in ensuring the safety of people's lives and property and the stability of infrastructure. With the rapid development of sensor technology, communication technology, and data processing technology, obtaining multi-source monitoring data has become relatively easy, such as meteorological data, GNSS (Global Navigation Satellite System) data, and radar data. These multi-source data reflect the state of the slope from different perspectives, providing a rich information foundation for accurately predicting slope displacement.

[0003] However, existing slope displacement prediction techniques have several problems when processing multi-source data:

[0004] On the one hand, traditional forecasting methods often rely on a single data source or simply splice or weighted merge multiple data sources, failing to delve into the intrinsic connections and complementary information between different data sources. For example, relying solely on GNSS data for displacement forecasting does not take into account the significant impact of meteorological factors (such as rainfall and earthquakes) on slope stability, resulting in forecasts that cannot accurately reflect the actual displacement changes of the slope.

[0005] On the other hand, while existing deep learning models (such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) have improved prediction accuracy to some extent, they still have significant shortcomings in feature extraction and fusion from multi-source data. When processing multimodal data, some models fail to effectively capture the complex spatiotemporal relationships between different modalities, resulting in incomplete and inaccurate features learned by the model, which affects the reliability and stability of predictions. Summary of the Invention

[0006] Therefore, the purpose of this invention is to develop and design a slope displacement prediction method that uses a multi-branch network and cross-attention multi-source data fusion. By constructing a multi-branch network to extract features from different modal data, and utilizing a cross-attention mechanism to deeply fuse multi-source data, this method achieves modality-specific feature extraction, cross-modal correlation modeling, and multi-scale prediction of multi-source data. This fully explores the potential information in multi-source data, improves the accuracy and reliability of slope displacement prediction, and provides stronger technical support for geological disaster early warning and prevention.

[0007] This invention provides a slope displacement prediction method using multi-branch networks and cross-attention multi-source data fusion, comprising the following steps:

[0008] S1. Collect meteorological data, GNSS node data, and radar 3D data of the slope area, and preprocess (clean, remove outliers and noise) and standardize the collected multi-source data to eliminate the differences in the dimensions and orders of magnitude of various data and unify the order of magnitude of the multi-source data.

[0009] S2. Construct a multi-branch network to extract features from various types of data, including: using a Transformer encoder to add learnable positional coding to encode meteorological data, and extracting meteorological feature vectors through a multi-head self-attention mechanism; using a temporal graph convolutional neural network (TGCN) to process GNSS data, performing adaptive adjacency matrix generation, convolution operations, and node-level attention pooling to obtain graph structure feature vectors; performing average pooling downsampling on radar data, using a CNN to extract spatiotemporal features, performing spatiotemporal positional coding, and calculating multi-head attention to obtain spatiotemporal feature vectors.

[0010] S3. The feature vectors extracted from the multi-source data are fused through a cross-attention mechanism, including: performing modality-specific gated projection; calculating cross-modality gated attention to obtain local correlation information; and using multimodal collaborative memory for global fusion, obtaining fused features through importance weighting.

[0011] S4. Construct a multi-scale prediction model, using a multi-scale memory network, combined with LSTM to capture recent trends, and a one-dimensional dilated convolutional neural network (1D DilatedCNN) to extract periodic features. The displacement prediction value is obtained through adaptive fusion. The multi-scale prediction model is trained using a course learning mechanism, and single-modal pre-training, fusion training and end-to-end joint optimization are performed in stages to output the displacement prediction value.

[0012] Furthermore, the method for encoding meteorological data in step S2 and extracting meteorological feature vectors through a multi-head self-attention mechanism includes:

[0013] use The long-range dependency problem of time-series features in meteorological data is solved by using a multi-head self-attention mechanism. The expression of the multi-head self-attention mechanism is as follows:

[0014] ;

[0015] in, For single-head attention, ; , , The matrices (features before linear transformation) correspond to the query, key, and value of meteorological data, respectively. , , The first Size Linear transformation weights; The linear transformation weights are used after concatenating the results from multiple sources. It is the key in single-head attention. Feature dimensions (due to multi-head attention) Split into Subspace, therefore ); It is a scaling factor used to mitigate the issue that the higher the dimension in dot product attention, the larger the dot product result. The vanishing gradient problem is addressed by making the attention distribution more even.

[0016] The expression for the learnable positional encoding is:

[0017]

[0018] ;

[0019] in, For the time step of meteorological data; Dimension index for location encoding ; The dimension for position encoding is the same as the dimension of the hidden layers in the Transformer model.

[0020] Specifically, Need and Dimensional matching is used to encode temporal information.

[0021] The learnable location encoding uses different frequencies of sine and cosine functions to encode the time sequence, enabling the Transformer model to capture the long-term temporal dependencies of meteorological data (the lower the frequency, the longer the encoding attention mechanism is for long-term dependencies; the higher the frequency, the shorter the encoding attention mechanism is for short-term dependencies), thus preserving the temporal information of meteorological data.

[0022] Furthermore, the expression for the spatiotemporal location encoding in step S2, which involves performing average pooling downsampling on the radar data, extracting spatiotemporal features using a CNN, and performing spatiotemporal location encoding, is as follows:

[0023] ;

[0024] Where t is the timestamp of the radar data (the time of the t-th radar scan); Spatial coordinates of the radar monitoring point (on the slope) Location); Encoding time location (as in question 2) (encoding the timing information of t); Encode the x-dimensional space (encode the spatial distribution of the x-coordinate, representing the left and right positions); Encode the spatial y-dimensional coordinates (encode the spatial distribution of the y-coordinates, representing the depth of the position).

[0025] The physical meaning of the spatiotemporal location encoding described in this invention is to fuse time (t) and space (x,y) information to assist the model in capturing the spatiotemporal dependence of radar data (representing the displacement evolution law at different locations of the slope) and to preserve the spatiotemporal information of radar data.

[0026] Furthermore, in step S2, the expression for the adaptive adjacency matrix generation in the process of using a time-map convolutional neural network to process GNSS data is as follows:

[0027] ;

[0028] in, , GNSS node characteristics For activation function, It is a multilayer perceptron.

[0029] Furthermore, the modality-specific gated projection in step S3 includes: transforming different modal features of the multi-source data to suppress redundant information. The expression for the modality-specific gated projection is:

[0030] , ;

[0031] in, Modal identifiers (multi-source data modes such as meteorological, GNSS, and radar data); These are the original features of mode m (such as meteorological temperature and humidity sequences, GNSS displacement sequences). The linear transformation weights and biases for mode m (dimensionality reduction and transformation of the original features); For activation function ( (Generate gating coefficients to suppress redundant features). The gating weights for mode m (adjusting the scale of the gated features); This is the modal feature after gated projection (filtering redundant information and enhancing effective features within the modality).

[0032] Furthermore, the method for calculating cross-modal gating attention in step S3 includes:

[0033] Local correlations between different modalities of multi-source data are calculated by evaluating gating weights and cross-modal interaction features; wherein, the formula for calculating the gating weights is:

[0034] ;

[0035] The formula for calculating the cross-modal interaction features is:

[0036] ;

[0037] in, For different modes (e.g.) Indicates weather, ); For modality Features (after gating projection) ); For feature concatenation (concatenating features from modality i and j to learn cross-modal associations); For gating weight matrix (learning cross-modal gating relationships); For activation function ( Generate gate weights Control mode right (Attention intensity) • To traverse all mode pairs (weather) GNSS, meteorology Radar, GNSS radar); For point-by-point multiplication (gating weights) Multiply the attention result and then weightedly fuse cross-modal information. For cross-attention (with modality i as the query) Modality For key Sum Capture right (related to) For cross-modal interaction features (fusing local correlations between different modalities, such as meteorological pairs) (Influence of displacement).

[0038] Furthermore, the expression for adaptive fusion in the displacement prediction value obtained through adaptive fusion in step S4 is:

[0039] ;

[0040] in, For adaptive fusion weights ( The complexity is dynamically calculated based on features, such as when there are many modalities and complex time sequences. Adjustment Branch weights are used to combine features at different scales for displacement prediction. These are the predicted values ​​for slope displacement. For fully connected layers (processing multimodal fusion features) (Capture global relationships); This is a multimodal fusion feature (a comprehensive feature after cross-modal attention). Temporal convolutional networks (for processing sequence features) (Capture long-term time dependencies) It is a time-series feature sequence (such as the time-series features of each modality, preserving the time-series evolution).

[0041] The physical meaning of adaptive fusion described in this invention is through... By balancing the features of "multimodal fusion (global)" and "temporal evolution (local)," the accuracy of displacement prediction is improved.

[0042] Furthermore, in step S4, the multi-scale prediction model is trained using a course learning mechanism, and the loss function for the end-to-end joint optimization in the phased single-modal pre-training, fusion training, and end-to-end joint optimization is:

[0043] ;

[0044] in, The mean absolute error, ; The actual displacement value of the slope monitoring point (e.g.) (Radar measured value); The predicted displacement value output by the model; To predict the total number of samples; For time-series smoothing loss, ,in, The number of time steps for predicting the time window; This represents the predicted displacement for adjacent time steps.

[0045] End-to-end joint optimization of the loss function can further improve the model training effect.

[0046] This invention extracts temporal, graph structure, and spatiotemporal features based on independent branches of a multi-branch network, avoiding modal interference and achieving decoupling of multimodal features, thus improving representation integrity. It employs cross-attention to dynamically capture cross-modal dependencies, achieving enhanced cross-source data association and overcoming the static weight defects of traditional fusion. It integrates short-term LSTM and long-term DilatedCNN, balancing high-frequency response and periodic learning, exhibiting better adaptability to unseen scale changes, good multi-scale adaptability, and suitability for complex scenarios. The course learning mechanism reduces the difficulty of multimodal training and improves training efficiency.

[0047] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method as described above.

[0048] The present invention also provides a computer device, the computer device including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the steps of the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method as described above.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0050] The multi-branch network and cross-attention multi-source data fusion slope displacement prediction method provided by this invention extracts features from different modal data by constructing a multi-branch network and deeply fuses multi-source data using a cross-attention mechanism. This enables the extraction of modal-specific features, cross-modal correlation modeling, and multi-scale prediction of multi-source data, which can fully explore the potential information in multi-source data and effectively improve the accuracy and reliability of slope displacement prediction, providing stronger technical support for geological disaster early warning and prevention. Attached Figure Description

[0051] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.

[0052] In the attached diagram:

[0053] Figure 1 This is a diagram of the multi-branch network feature extraction architecture according to an embodiment of the present invention;

[0054] Figure 2 This is a schematic diagram of the architecture of the multi-head self-attention mechanism according to an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of the cross-attention mechanism according to an embodiment of the present invention;

[0056] Figure 4 This is a flowchart illustrating the multi-scale prediction process according to an embodiment of the present invention.

[0057] Figure 5 This is a schematic diagram of the system architecture according to an embodiment of the present invention;

[0058] Figure 6 This is a flowchart of the slope displacement prediction method based on multi-branch network and cross-attention multi-source data fusion of the present invention;

[0059] Figure 7 This is a schematic diagram of the configuration of a computer device according to an embodiment of the present invention. Detailed Implementation

[0060] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and products consistent with some aspects of this disclosure as detailed in the appended claims.

[0061] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0062] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0063] The embodiments of the present invention will be described in further detail below.

[0064] This invention provides a slope displacement prediction method using multi-branch networks and cross-attention multi-source data fusion. (See also...) Figure 6 As shown, it includes the following steps:

[0065] S1. Collect meteorological data, GNSS node data, and radar 3D data of the slope area, and preprocess (clean, remove outliers and noise) and standardize the collected multi-source data to eliminate the differences in the dimensions and orders of magnitude of various data and unify the order of magnitude of the multi-source data.

[0066] Figure 1 The forecasting process is shown, displaying meteorological, GNSS, and radar point cloud data.

[0067] In this embodiment, meteorological data is collected at fixed time intervals, with meteorological factors (such as precipitation and wind speed) collected. The time window length is T, and the input dimension after standardization is [missing information]. ( For batch size, (For feature dimensions).

[0068] GNSS data acquisition and deployment At each monitoring point, three-dimensional coordinate changes are collected at a fixed frequency. A data structure is constructed based on the correlation between geographical distance and historical displacement. Each node A graph structure with striped edges.

[0069] The acquired radar data is processed through gridding to generate a spatial resolution of [missing information]. Deformation field data, with a time step of The input dimension is ( (Number of channels).

[0070] Detect outliers in the data using statistical methods. Outliers are removed using the criteria, and missing values ​​are repaired using sliding window interpolation.

[0071] To achieve radar data downsampling, by Average pooling reduces spatial resolution and computational complexity from Down to ;

[0072] in, The spatial height dimension of radar data; The spatial width dimension of radar data; This refers to the channel dimension of radar data. The downsampling factor for average pooling is determined by... Average pooling reduces spatial resolution, and the computational cost decreases quadratically with increasing spatial dimension (e.g., ...). At that time, the computational load was reduced to the original ).

[0073] S2. Construct a multi-branch network to extract features from various types of data, including: using a Transformer encoder to add learnable positional encoding to encode meteorological data, and employing a multi-head self-attention mechanism (such as...). Figure 2 (As shown) Extract meteorological feature vectors; use a temporal graph convolutional neural network (TGCN) to process GNSS data, perform adaptive adjacency matrix generation, convolution operation and node-level attention pooling to obtain graph structure feature vectors; perform average pooling downsampling on radar data, use CNN to extract spatiotemporal features, perform spatiotemporal position encoding and calculate multi-head attention to obtain spatiotemporal feature vectors;

[0074] Encoding meteorological data and extracting meteorological feature vectors using a multi-head self-attention mechanism includes:

[0075] use The long-range dependency problem of time-series features in meteorological data is solved by using a multi-head self-attention mechanism. The expression of the multi-head self-attention mechanism is as follows:

[0076] ;

[0077] in, For single-head attention, ; , , The matrices (features before linear transformation) correspond to the query, key, and value of meteorological data, respectively. , , The first Size , , Linear transformation weights; The linear transformation weights are used after concatenating the results from multiple sources. It is the key in single-head attention. Feature dimensions (due to multi-head attention) Split into Subspace, therefore ); It is a scaling factor used to mitigate the issue that the higher the dimension in dot product attention, the larger the dot product result. The vanishing gradient problem is addressed by making the attention distribution more even.

[0078] The expression for the learnable positional encoding is:

[0079]

[0080] ;

[0081] in, For the time step of meteorological data; Dimension index for location encoding ; The dimension for position encoding is the same as the dimension of the hidden layers in the Transformer model.

[0082] Need and Dimensional matching is used to encode temporal information.

[0083] The learnable location encoding uses different frequencies of sine and cosine functions to encode the time sequence, enabling the Transformer model to capture the long-term temporal dependencies of meteorological data (the lower the frequency, the longer the encoding attention mechanism is for long-term dependencies; the higher the frequency, the shorter the encoding attention mechanism is for short-term dependencies), thus preserving the temporal information of meteorological data.

[0084] The expression for spatiotemporal location encoding is:

[0085] ;

[0086] Where t is the timestamp of the radar data (the time of the t-th radar scan); Spatial coordinates of radar monitoring points [on the slope] Location】; Encoding time and location (similar to question 2) (encoding the timing information of t); Encode the x-dimensional space (encode the spatial distribution of the x-coordinate, representing the left and right positions); Encode the spatial y-dimensional coordinates (encode the spatial distribution of the y-coordinates, representing the depth of the position).

[0087] In this embodiment, the spatiotemporal location encoding integrates time (t) and spatial (x,y) information to assist the model in capturing the spatiotemporal dependence of radar data (representing the displacement evolution law at different locations of the slope) and preserving the spatiotemporal information of the radar data.

[0088] The expression for generating the adaptive adjacency matrix is:

[0089] ;

[0090] in, , GNSS node characteristics For activation function, It is a multilayer perceptron.

[0091] S3. The feature vectors extracted from the multi-source data are fused through a cross-attention mechanism, including: performing modality-specific gated projection; calculating cross-modality gated attention to obtain local correlation information; and using multimodal collaborative memory for global fusion, obtaining fused features through importance weighting.

[0092] Modality-specific gated projection includes: transforming different modal features of multi-source data to suppress redundant information. The expression for the modality-specific gated projection is:

[0093] , ;

[0094] in, Modal identifiers (multi-source data modes such as meteorological, GNSS, and radar data); These are the original features of mode m (such as meteorological temperature and humidity sequences, GNSS displacement sequences). The linear transformation weights and biases for mode m (dimensionality reduction and transformation of the original features); For activation function ( (Generate gating coefficients to suppress redundant features). The gating weights for mode m (adjusting the scale of the gated features); This is the modal feature after gated projection (filtering redundant information and enhancing effective features within the modality).

[0095] Methods for computing cross-modal gated attention include:

[0096] Local correlations between different modalities of multi-source data are calculated by evaluating gating weights and cross-modal interaction features; wherein, the formula for calculating the gating weights is:

[0097] ;

[0098] The formula for calculating the cross-modal interaction features is:

[0099] ;

[0100] in, For different modes (e.g.) =Meteorology, ); For modality Features (after gating projection) ); For feature concatenation (concatenating features from modality i and j to learn cross-modal associations); For gating weight matrix (learning cross-modal gating relationships); For activation function ( Generate gate weights · Control mode right (Attention intensity) To traverse all mode pairs (weather) GNSS, meteorology Radar, GNSS radar); For point-by-point multiplication (gating weights) Multiply the attention result and then weightedly fuse cross-modal information. For cross-attention (with modality i as the query) Modality For key Sum Capture right (related to) For cross-modal interaction features (fusing local correlations between different modalities, such as meteorological pairs) (Influence of displacement).

[0101] Figure 3 The processing logic of dynamic projection, cross-modal interaction, and global aggregation, which fuses feature vectors extracted from multi-source data through a cross-attention mechanism, is illustrated.

[0102] S4. Construct a multi-scale prediction model, using a multi-scale memory network, combined with LSTM to capture recent trends, and a one-dimensional dilated convolutional neural network (1D Dilated CNN) to extract periodic features. The displacement prediction value is obtained through adaptive fusion. The multi-scale prediction model is trained using a course learning mechanism, and single-modal pre-training, fusion training and end-to-end joint optimization are performed in stages to output the displacement prediction value.

[0103] The expression for adaptive fusion is:

[0104] ;

[0105] in, For adaptive fusion weights ( The complexity is dynamically calculated based on the "feature complexity," such as when there are many modalities and complex time sequences. Adjustment Branch weights are used to combine features at different scales for displacement prediction. These are the predicted values ​​for slope displacement. For fully connected layers (processing multimodal fusion features) (Capture global relationships); This is a multimodal fusion feature (a comprehensive feature after cross-modal attention). Temporal convolutional networks (for processing sequence features) (Capture long-term time dependencies) It is a time-series feature sequence (such as the time-series features of each modality, preserving the time-series evolution).

[0106] The adaptive fusion in this embodiment is achieved through... By balancing the features of "multimodal fusion (global)" and "temporal evolution (local)," the accuracy of displacement prediction is improved.

[0107] The loss function for end-to-end joint optimization is:

[0108] ;

[0109] in, The mean absolute error, ; The actual displacement value of the slope monitoring point (e.g.) (Radar measured value); The predicted displacement value output by the model; To predict the total number of samples; For time-series smoothing loss, ,in, The number of time steps for predicting the time window; This represents the predicted displacement for adjacent time steps.

[0110] In this embodiment, the single-modal pre-training phase includes: training Rounds round It is the process of "inputting the entire training dataset into the model and updating the parameters once", and validating the error metrics on the validation set. , This is the error threshold for the validation set in this stage. The mean squared error is used to ensure the basic feature extraction capability of each branch.

[0111] The fusion training phase includes: freezing the parameters in feature extraction, optimizing only the fusion, and training. indivual Through feature visualization (such as This improves the separability of verification classes.

[0112] The end-to-end joint optimization phase includes: activating all parameters and training based on the joint loss function. indivual Test set error metrics To verify the effectiveness of multimodal fusion. This is the error threshold for the test set in this phase. If... Then the performance of multimodal fusion should be better than that of single-modal fusion (otherwise fusion is worthless); if If so, it is necessary to go back to the single-modal pre-training stage and train again (possibly due to overfitting in single-modal pre-training, failure to learn effective correlations in fusion training, etc.).

[0113] The end-to-end joint optimization of the loss function further improves the model training effect.

[0114] Table 1 shows the key parameter configurations for each model in this embodiment;

[0115] Table 1

[0116]

[0117] Figure 4 The diagram illustrates the short-term trend capture, long-term pattern extraction, and adaptive fusion process presented in multi-scale prediction. Figure 5This embodiment illustrates the functional division and data flow of multi-branch network and cross-attention multi-source data fusion for slope displacement prediction.

[0118] This embodiment extracts temporal, graph structure, and spatiotemporal features based on independent branches of a multi-branch network, avoiding modal interference, achieving multimodal feature decoupling, and improving representation integrity. It adopts cross-attention to dynamically capture cross-modal dependencies, realize cross-source data association enhancement, and solve the static weight defects of traditional fusion. The course learning mechanism reduces the difficulty of multimodal training and improves training efficiency. It integrates short-term LSTM and long-term Dilated CNN, taking into account high-frequency response and periodic learning, showing better adaptability to unseen scale changes, with good multi-scale adaptability, and is suitable for complex scenarios.

[0119] The hardware environment in this embodiment is: GPU-accelerated computing, single-sample inference time. (batch size) ), The unit is milliseconds.

[0120] The software architecture is as follows: it is implemented using the deep learning framework PyTorch; graph data is processed through a graph neural network library; data loading supports asynchronous mechanisms; prediction results are output as displacement values; and it supports integration with external systems. Docking.

[0121] This embodiment improves the model's versatility through parametric design, avoids specific numerical limitations, and is suitable for flexible configuration and expansion in different monitoring scenarios.

[0122] This invention also provides a computer device. Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention; see the accompanying drawings. Figure 7 As shown, the computer device includes: an input system 23, an output system 24, a memory 22, and a processor 21; the memory 22 is used to store one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method provided in the above embodiments; wherein the input system 23, the output system 24, the memory 22, and the processor 21 can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.

[0123] The memory 22, as a read / write storage medium for a computing device, can be used to store software programs and computer-executable programs, such as the program instructions corresponding to the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method described in this embodiment of the invention. The memory 22 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the device. Furthermore, the memory 22 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 22 may further include memory remotely located relative to the processor 21, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0124] The input system 23 can be used to receive input digital or character information, and generate key signal inputs related to user settings and function control of the device; the output system 24 may include display devices such as a display screen.

[0125] The processor 21 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 22, thereby realizing the above-mentioned multi-branch network and cross-attention multi-source data fusion slope displacement prediction method.

[0126] The computer equipment provided above can be used to execute the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method provided in the above embodiments, and has corresponding functions and beneficial effects.

[0127] This invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method provided in the above embodiments. The storage medium can be any type of memory device or storage device, including: mounting media such as CD-ROM, floppy disk, or magnetic tape systems; computer system memory or random access memory such as DRAM, DDRRAM, SRAM, EDORAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (e.g., hard disk or optical storage); registers or other similar types of memory elements; the storage medium may also include other types of memory or combinations thereof; furthermore, the storage medium may reside in a first computer system in which the program is executed, or it may reside in a different second computer system connected to the first computer system via a network (such as the Internet); the second computer system can provide program instructions to the first computer for execution. The storage medium includes two or more storage media that can reside in different locations (e.g., in different computer systems connected via a network). The storage medium can store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.

[0128] Of course, the computer-executable instructions provided in the embodiments of the present invention are not limited to the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method described in the above embodiments, but can also perform related operations in the multi-branch network and cross-attention multi-source data fusion slope displacement prediction method provided in any embodiment of the present invention.

[0129] The technical solution of the present invention has been described in conjunction with preferred embodiments. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions resulting from these changes or substitutions will all fall within the scope of protection of the present invention.

[0130] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A slope displacement prediction method using multi-branch network and cross-attention multi-source data fusion, characterized in that, Includes the following steps: S1. Collect meteorological data, GNSS node data, and radar 3D data of the slope area, and preprocess and standardize the collected multi-source data to eliminate the differences in the dimensions and orders of magnitude of various data and unify the order of magnitude of the multi-source data. S2. Construct a multi-branch network to extract features from various types of data, including: using a Transformer encoder to add learnable position coding to encode meteorological data, and extracting meteorological feature vectors through a multi-head self-attention mechanism; using a temporal graph convolutional neural network to process GNSS data, performing adaptive adjacency matrix generation, convolution operations, and node-level attention pooling to obtain graph structure feature vectors; performing average pooling downsampling on radar data, using a CNN to extract spatiotemporal features, performing spatiotemporal position coding, and calculating multi-head attention to obtain spatiotemporal feature vectors. S3. The feature vectors extracted from the multi-source data are fused through a cross-attention mechanism, including: performing modality-specific gated projection; calculating cross-modality gated attention to obtain local correlation information; and using multimodal collaborative memory for global fusion, obtaining fused features through importance weighting. S4. Construct a multi-scale prediction model, using a multi-scale memory network, combined with LSTM to capture recent trends, and a one-dimensional dilated convolutional neural network to extract periodic features. Obtain displacement prediction values ​​through adaptive fusion. Train the multi-scale prediction model using a course learning mechanism, and perform single-modal pre-training, fusion training and end-to-end joint optimization in stages to output displacement prediction values. The expression for the adaptive fusion in the displacement prediction value obtained through adaptive fusion in step S4 is: ; in, For adaptive fusion weights; These are the predicted values ​​for slope displacement. It is a fully connected layer; This is a multimodal fusion feature; For temporal convolutional networks; It is a time-series characteristic sequence; The S4 step employs a course learning mechanism to train the multi-scale prediction model, performing single-modal pre-training, fusion training, and end-to-end joint optimization in stages. The loss function for the end-to-end joint optimization is as follows: ; in, The mean absolute error, ; This represents the actual displacement value of the slope monitoring point; The predicted displacement value output by the model; To predict the total number of samples; For time-series smoothing loss, ,in, The number of time steps for predicting the time window; This represents the predicted displacement for adjacent time steps.

2. The slope displacement prediction method based on multi-branch network and cross-attention multi-source data fusion according to claim 1, characterized in that, The method for encoding meteorological data in step S2 and extracting meteorological feature vectors through a multi-head self-attention mechanism includes: use The long-range dependency problem of time-series features in meteorological data is solved by using a multi-head self-attention mechanism. The expression of the multi-head self-attention mechanism is as follows: ; in, For single-head attention, ; , , A matrix corresponding to the query, key, and value of meteorological data, respectively; , , The first Size , , Linear transformation weights; The linear transformation weights are used after concatenating the results from multiple sources. It is the key in single-head attention. Feature dimensions; It is the scaling factor; The expression for the learnable positional encoding is: ; ; in, For the time step of meteorological data; Dimension index for location encoding ; The dimension for position encoding is the same as the dimension of the hidden layers in the Transformer model.

3. The slope displacement prediction method based on multi-branch network and cross-attention multi-source data fusion according to claim 1, characterized in that, The expression for the spatiotemporal location encoding in step S2, which involves average pooling downsampling of radar data, extracting spatiotemporal features using a CNN, and performing spatiotemporal location encoding, is as follows: ; Where t is the timestamp of the radar data; The spatial coordinates of the radar monitoring point; Encode the time location; Encode the x-dimensional space; Encode the y-dimensional space.

4. The slope displacement prediction method based on multi-branch network and cross-attention multi-source data fusion according to claim 1, characterized in that, The expression for the adaptive adjacency matrix generation in step S2, which uses a time-map convolutional neural network to process GNSS data, is as follows: ; in, , GNSS node characteristics For activation function, It is a multilayer perceptron.

5. The slope displacement prediction method based on multi-branch network and cross-attention multi-source data fusion according to claim 1, characterized in that, The modality-specific gated projection in step S3 includes: transforming different modal features of multi-source data to suppress redundant information. The expression for the modality-specific gated projection is: , ; in, Modal identifier; These are the original features of mode m; For the linear transformation weights and biases of mode m; For activation functions; Let be the gating weights for mode m; These are the modal features after gating and projection.

6. The slope displacement prediction method based on multi-branch network and cross-attention multi-source data fusion according to claim 1, characterized in that, The method for calculating cross-modal gating attention in step S3 includes: Local correlations between different modalities of multi-source data are calculated by evaluating gating weights and cross-modal interaction features; wherein, the formula for calculating the gating weights is: ; The formula for calculating the cross-modal interaction features is: ; in, For different modes; For modality Features; For feature splicing; This is the gate weight matrix; For activation functions; Gating weights for different modalities; To traverse all modal pairs; This is a point-by-point multiplication; For cross attention; This refers to cross-modal interaction features.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the slope displacement prediction method based on the multi-branch network and cross-attention multi-source data fusion method according to any one of claims 1-6.

8. A computer device, the computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the slope displacement prediction method based on the multi-branch network and cross-attention multi-source data fusion as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Video saliency object detection model and system based on cross attention mechanism

    CN112149459A

  • Target tracking method and system based on multi-scale aggregation attention feature extraction network

    CN118429389A