Tunnel face geological feature bimodal three-layer multi-scale multi-task fusion intelligent identification method and system

By integrating tunnel face images with drilling parameters into a three-layer, multi-scale, multi-task model, the limitations of a single data source in tunnel construction have been solved, enabling high-precision surrounding rock classification and construction optimization, and promoting the intelligent development of tunnel engineering.

CN120913053BActive Publication Date: 2026-05-08SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHWEST JIAOTONG UNIV
Filing Date
2025-06-09
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing tunnel construction, the analysis of a single data source has limitations in rock stratum identification and construction optimization, making it difficult to achieve high-precision stratum identification and construction quality assessment. In particular, it is highly dependent on traditional manual experience, and the identification accuracy is affected by lighting conditions and rock surface contamination.

Method used

A dual-modal, three-layer, multi-scale, and multi-task fusion intelligent identification method is adopted. By combining tunnel face images and drilling parameters, a deep fusion model is constructed. Through GeoDeformNet, GeoPyramidGAT, and GAV-Mamba networks, local, regional, and global multi-scale perception is achieved, thereby improving the accuracy of geological feature identification.

Benefits of technology

It significantly improves the accuracy and automation of surrounding rock classification and identification, reduces reliance on human experience, adapts to complex geological structures, and supports intelligent and safe tunnel construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913053B_ABST
    Figure CN120913053B_ABST
Patent Text Reader

Abstract

The application belongs to the field of tunnel engineering, and specifically discloses a double-mode three-layer multi-scale multi-task fusion intelligent identification method and system for geological features of a tunnel face, which comprises the following steps: collecting images of the tunnel face after on-site blasting is completed; collecting and sorting on-site construction data, including original drilling parameters and corresponding geological sketch information of the tunnel face; integrating and cleaning the original drilling parameters, combining the geological feature information extracted from the tunnel face images and the geological sketch to construct a multi-modal data set; and based on the multi-modal data set, constructing a multi-modal data fusion surrounding rock grade intelligent identification model and performing intelligent identification. The application realizes intelligent identification of the geological features of the tunnel face by fusing multi-dimensional information of the tunnel face images and the drilling parameters, significantly improves the accuracy and robustness of the geological feature identification, optimizes stratum judgment and blasting parameter adjustment in the construction process, and effectively promotes the intelligent development of tunnel construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tunnel engineering technology, specifically to a dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method and system for geological features of tunnel faces. Background Technology

[0002] Drill-and-blast method is one of the main methods for rock tunnel construction, and its construction quality and safety highly depend on the accurate assessment of geological conditions. During construction, the tunnel face is the area that most directly reflects lithological characteristics and blasting effects, traditionally relying mainly on manual observation and analysis. However, this method is greatly affected by subjective factors and is difficult to quantify. In recent years, with the development of computer vision technology, intelligent analysis of the tunnel face using image processing and deep learning methods has become possible. Through techniques such as semantic segmentation, object detection, and texture analysis, lithological information, joint and fracture characteristics, and blasting effects of the tunnel face can be extracted, providing data support for construction optimization. However, relying solely on image analysis has certain limitations; factors such as lighting conditions and rock surface contamination may affect recognition accuracy. Therefore, it is necessary to combine other data sources for multimodal information fusion to improve the accuracy and stability of geological identification.

[0003] Drilling parameters are another important type of formation characterization data during drill-and-blast drilling, mainly including drilling speed, torque, thrust, and impact frequency. Formations of different lithologies exhibit different parameter characteristics during drilling; for example, soft rock has a high drilling speed but low torque, while hard rock has a slow drilling speed but high torque. Based on these parameters, researchers have used statistical analysis, machine learning, and deep learning methods to classify formations, achieving some recognition results. However, drilling parameters only reflect the stress conditions of the drilling rig during drilling and cannot directly provide information on the spatial distribution or structural characteristics of the formation. Therefore, analysis based on a single data source has limitations. How to effectively integrate face images and drilling parameters to achieve more accurate formation identification and construction optimization has become a key research focus.

[0004] Multimodal learning offers a new technical approach to solving this problem. By fusing face images with drilling parameters, the complementary nature of visual information and drilling data can be fully utilized to improve the accuracy of formation identification. Current mainstream multimodal fusion methods include feature-level fusion, decision-level fusion, and deep fusion. Deep fusion typically employs multimodal neural networks, such as combining CNNs for image feature extraction, LSTMs or Transformers for processing time-series data, and using self-attention mechanisms to capture the correlation between the two. This fusion approach not only improves lithology identification accuracy but can also be further applied to blasting parameter optimization, construction quality assessment, and intelligent decision-making. In the future, combining 3D visual reconstruction, real-time monitoring, and automated control technologies will drive the development of drill-and-blast tunnel construction towards a more intelligent and efficient direction. Summary of the Invention

[0005] This invention aims to address the problems of single data utilization, insufficient information fusion, and low automation in existing tunnel surrounding rock identification technologies. This invention provides a dual-modal, three-layer, multi-scale, and multi-task fusion intelligent identification method and system for tunnel face geological features. This method comprehensively utilizes the spatial texture information of images and the physical feedback information of the drilling process to construct a deep fusion model with local, regional, and global multi-scale perception capabilities. This enables high-precision intelligent identification of tunnel face geological features, taking into account multiple application needs such as surrounding rock classification identification, blasting parameter optimization, and construction risk early warning, thus solving the problems mentioned in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces, comprising the following steps:

[0007] S1. Collect images of the tunnel face after the blasting is completed.

[0008] S2. Collect and organize on-site construction data, including raw drilling parameters and corresponding geological sketch information of the working face;

[0009] S3. Integrate and clean the original drilling parameters, and construct a multimodal dataset by combining the geological feature information extracted from the face image and geological sketch.

[0010] S4. Based on the multimodal dataset, construct a multimodal data fusion-based intelligent identification model for surrounding rock level, and intelligently identify the surrounding rock level at the working face.

[0011] Preferably, in step S1, when acquiring images of the tunnel face after the on-site blasting of the tunnel constructed by the drill-and-blast method, after ventilation for at least 30 minutes, a 10,000 Lux diffused light source is set up at a distance of 10 to 12 meters from the tunnel face and 1.5 meters on both sides of the tunnel centerline. The image acquisition device with a resolution of not less than 20 million pixels is used to take pictures, and multiple images are acquired from the same tunnel face. Finally, the image with the highest quality is selected and saved as the final image input.

[0012] Preferably, in step S2, the original drilling parameters include indicators of propulsion pressure, impact pressure, rotational torque, and feed rate, which are used to reflect the hardness and degree of fracturing of the surrounding rock; the geological sketch information includes information on lithology, joint and fracture development, and groundwater conditions, which serve as the source of subsequent identification labels.

[0013] Preferably, in step S3, the original drilling parameters automatically collected by the intelligent rock drilling rig are integrated, and the coordinates, feed rate, impact pressure, propulsion pressure and rotation pressure of each borehole are extracted, and the above parameters are standardized into fixed lengths.

[0014] Simultaneously, the mileage information and geological features corresponding to the working face are extracted from the geological sketch;

[0015] Using mileage and borehole number as index labels, a multimodal dataset consisting of face images, drilling parameters, and surrounding rock classification information is constructed.

[0016] Preferably, in step S4, the intelligent identification model is a multimodal fusion identification model with a three-layer cross-modal attention pyramid structure, specifically including:

[0017] 1) Bottom layer - GeoDeformNet network, used to focus on local features of regions with abrupt changes in lithological texture and drilling parameters in the image;

[0018] 2) The middle layer - GeoPyramidGAT network is used to extract regional features of joint distribution trends and local variation trends of drilling parameters in the image;

[0019] 3) Top-level - GeoAugmented Vision Mamba network, also known as GAV-Mamba network, is used to capture the overall stability characteristics of the image and the global distribution characteristics of the main variables in the drilling parameters.

[0020] Preferably, the GeoDeformNet network includes:

[0021] (1) Across attention layers, focus on the abrupt change regions of drilling parameters between adjacent boreholes;

[0022] First, select a borehole that is close to the current borehole by calculating the Euclidean geometric distance.

[0023] D ij =||X i -X j ||2

[0024] Where D ij X represents the Euclidean distance between the i-th and j-th boreholes. i ,X j ∈R 2 , indicating two different borehole coordinates;

[0025] Next, the statistical characteristics of the selected boreholes and the current borehole are calculated, specifically the difference in mean, standard deviation, and maximum absolute value. The calculation formulas are as follows:

[0026]

[0027] δ i =max|P i -P j |

[0028] Where μ i σ i δ i The mean difference, standard deviation, and maximum absolute difference of the i-th borehole are represented; k represents the number of boreholes involved in the calculation, which defaults to 2; P represents the original drilling parameter matrix.

[0029] The final output features are:

[0030] F drill,i =[P i ,μ i ,σ i ,δ i ]∈R 4d

[0031] Where F drill,i P represents the encoded borehole feature of the i-th borehole. i ,μ i ,σ i ,δ i The original parameter features, mean difference, standard deviation, and maximum absolute difference of the i-th borehole are represented sequentially, [,] denote vector concatenation, R represents the real number field, and d represents the dimension, i.e., P mentioned earlier. i ,μ i ,σ i ,δ i These four dimensions;

[0032] (2) Deformable convolutional layers are used to extract local features of abrupt changes in lithological texture in images;

[0033] The specific definition of deformable convolutional layers is as follows:

[0034] F1 = ReLU(Conv1(I))

[0035] Δ = OffsetConv(F1)

[0036] F2 = ReLU(DeformConv(F1,Δ))

[0037] F3 = MaxPool(F2)

[0038] F img =ReLU(Conv2(F3))

[0039] Where I represents the input face image, F1 represents a feature extracted from the original image convolutional layer and activated, Δ represents a learnable offset defined by F1 for constructing deformable layers, F2 represents a feature extracted from the input features by a deformable convolutional layer and activated, and F3 represents a feature obtained from the input features through max pooling. img The features are obtained from the input features via ReLU activation;

[0040] (3) Region attention mechanism, used to achieve a fine correspondence between image features and borehole drilling parameters based on pixel coordinates; firstly, the borehole coordinates need to be normalized, and then linearly mapped to project them onto a unified dimension for alignment:

[0041]

[0042] in, This represents the coordinates normalized to [-1, 1], W' and H' represent the dimensions after image processing, and GridSample() represents a sampling layer used to extract features from the image at the corresponding location based on the coordinates. The extracted drilling parameters and image features are aligned to facilitate subsequent fusion. The specific alignment steps are as follows:

[0043]

[0044] in P represents the input borehole coordinate matrix. b This represents the original, input drilling parameters. This represents the aligned characteristics of the drilling parameters. This represents the drilling parameter features processed by the cross-attention layer. `DrillParamEncoder()` is a custom cross-attention layer used to extract differential features between adjacent boreholes during drilling. p Let b represent a trainable projective weight matrix. p Represent a bias matrix;

[0045] Finally, the aligned drilling features and image features are fused using a multilayer perceptron.

[0046]

[0047] in The fused feature matrix, These represent the image features before fusion and the features during drilling, respectively.

[0048] Preferably, the GeoPyramidGAT network performs preliminary structural surface division on the facet image based on the superpixel segmentation method, and records the set of pixel coordinates for each region:

[0049] S = SLIC(I,n)

[0050] Ω k ={(x,y)|S(x,y)=k}

[0051]

[0052] Where SLIC() represents dividing the original image I into n blocks using a superpixel segmentation algorithm, Ω k Used to record the edge pixel coordinates of each image block, l i Used to map the recorded edge coordinates to the borehole position, thereby dividing the area described by the borehole; Represents the x and y coordinates of each point on the superpixel block;

[0053] Next, we calculate the statistical characteristics of the borehole drilling parameters in this area, including the mean, standard deviation, minimum, and maximum values:

[0054]

[0055] min kd =minf id

[0056] max kd =maxf id

[0057] Where μ kd σ kd min kd ,max kd This represents the mean, standard deviation, minimum, and maximum values ​​of the drilling parameters in the corresponding region; f id This represents the drilling parameters for the i-th borehole within this region;

[0058] Next, feature extraction needs to be performed on the image patches. The specific steps are as follows:

[0059]

[0060] PyramidPooling() represents the processing of image block I. k Pyramid pooling is used to transform it into a three-layer image feature pyramid;

[0061] Finally, a graph structure is constructed, where each node represents an image region. The node features include image region features and drilling parameter statistics. The specific steps are as follows:

[0062]

[0063] E = {(v k ,v j )|v j ∈KNN(C k ,K),j≠k}

[0064] in This represents the statistical characteristics of the current region during drilling, obtained through the above calculations. h represents the features of the current region image patch obtained through the above calculations. k C represents the stitched statistical features and image characteristics after drilling. k V represents the pixel coordinates of the graph nodes obtained through calculation, E represents the set of edges between the constructed nodes, and v k ,v j Representing two distinct graph nodes, v j ∈KNN(C k ,K) represents the number of edges that the current node should be adjacent to according to the K-nearest neighbor algorithm, where K is a hyperparameter of the K-nearest neighbor algorithm;

[0065] Finally, the constructed graph structure is fed into the graph neural network as input to learn region-level fusion features.

[0066] Preferably, the GAV-Mamba network maps the pixel coordinates of the image to the corresponding drilling parameters, so that the drilling parameters form a two-dimensional structure with the same size as the image, specifically including the following:

[0067] Define image network coordinates:

[0068] G={(x,y)∣x∈{0,1,...,W-1},y∈{0,1,...,H-1}}

[0069] Where W represents the width of the face image and H represents the height of the face image;

[0070] For each channel containing the drilling parameters, the sparse parameter values ​​are interpolated in the two-dimensional image domain:

[0071]

[0072] Interp method () indicates the two-dimensional interpolation method used;

[0073] Finally, the normalized image and the interpolation parameter map are stitched together along the channel dimension.

[0074]

[0075] I represents the input facet image. This represents the feature matrix of drilling parameters obtained after the above steps; the tensor F is used to extract global semantic features and train the model.

[0076] Similarly, this invention can be used to intelligently identify geological features of the working face, such as the surrounding rock grade, rock hardness, and rock mass integrity.

[0077] On the other hand, to achieve the above objectives, the present invention also provides the following technical solution: a dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification system for geological features of tunnel faces, comprising the following modules:

[0078] The image acquisition module acquires images of the tunnel face after blasting is completed.

[0079] The data collection module collects and organizes on-site construction data, including raw drilling parameters and corresponding geological sketch information of the drilling face;

[0080] The multimodal dataset construction module integrates and cleans the original drilling parameters, and combines geological feature information extracted from the face image and geological sketch to construct a multimodal dataset.

[0081] The intelligent identification module constructs a multimodal data fusion-based intelligent identification model for the surrounding rock level and performs intelligent identification of the surrounding rock level at the working face.

[0082] The beneficial effects of this invention are as follows: Compared with traditional methods that rely on a single modality (such as images or parameters) or on human experience for surrounding rock identification, this invention has the following significant advantages: 1) It introduces a dual-modal collaborative modeling mechanism of images and drilling parameters to fully explore the texture structure in the face image and the physical response features in the drilling parameters; 2) It innovatively designs a three-layer multi-scale cross-modal fusion architecture, corresponding to local mutation identification, regional structure modeling, and global semantic extraction, respectively, significantly improving the model's adaptability to complex geological structures and identification accuracy; 3) Through an end-to-end deep learning model, it improves the automation and intelligence level of surrounding rock classification judgment and reduces reliance on human experience; 4) The method of this invention has significant engineering application value and promotion prospects, and can be widely applied to geological identification, blasting optimization, and risk monitoring in drill-and-blast tunnel construction, providing key technical support for intelligent and safe tunnel construction. Attached Figure Description

[0083] Figure 1 This is a schematic diagram of the dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features at the tunnel face in Example 1;

[0084] Figure 2 This is a schematic diagram of the tunnel face image specifications in Example 1, where (a) shows image rotation and (b) shows cropping.

[0085] Figure 3 This is a schematic diagram of the overall construction of the intelligent identification model for dual-modal fusion of geological features in Example 1;

[0086] Figure 4 The diagram below illustrates the key mechanisms of the GeoDeformNet network in Example 1. (a) shows the learning of drilling fluctuations across attention layers, (b) shows the learning of local image features by deformable convolutional layers, and (c) shows the refined mapping of pixel coordinate space.

[0087] Figure 5 This is a schematic diagram of superpixel segmentation using the GeoPyramidGAT network in Example 1;

[0088] Figure 6 This is a schematic diagram of the GeoPyramidGAT network structure in Example 1;

[0089] Figure 7 This is a detailed structural diagram of the GAV-Mamba network in Example 1;

[0090] Figure 8 This is a comparison chart of the accuracy of surrounding rock classification between the present invention and the comparative model in Example 1;

[0091] Figure 9This is a schematic diagram of a dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification system module for geological features at the tunnel face, as described in an embodiment of the present invention.

[0092] In the diagram, 110 is the image acquisition module; 120 is the data collection module; 130 is the multimodal dataset construction module; and 140 is the intelligent identification module. Detailed Implementation

[0093] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0094] Example 1

[0095] Please see Figure 1 This invention provides a technical solution: a dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces, comprising the following steps:

[0096] S1. Collect images of the tunnel face after the blasting is completed.

[0097] When acquiring images of the tunnel face after blasting in drill-and-blast construction, after ventilation for at least 30 minutes, place a 10,000 Lux diffused light source 10-12 meters away from the tunnel face and 1.5 meters on each side of the tunnel centerline to provide stable and uniform illumination. Use an image acquisition device with at least 20 megapixels to capture images, and take multiple images of the same tunnel face. Finally, select the image with the highest quality (high clarity, even illumination, and no distortion or occlusion) and save it as the final image input. This provides a high-quality data foundation for subsequent image processing and model recognition.

[0098] S2. Collect and organize on-site data, including raw drilling parameters and corresponding geological sketches of the tunnel face. Based on the image acquisition results, simultaneously collect raw drilling parameters and geological sketches for the corresponding construction mileage sections of the tunnel face to ensure data consistency in time and space. Raw drilling parameters include indicators such as feed pressure, impact pressure, slewing torque, and feed rate, used to reflect the hardness and fracturing degree of the surrounding rock. The geological sketch information includes information on lithology, joint and fracture development, weathering degree, and groundwater distribution, serving as a source of subsequent identification labels and providing geological background support for the construction of multimodal features.

[0099] S3. Integrate and clean the original drilling parameters, and construct a multimodal dataset by combining the geological feature information extracted from the face image and geological sketch.

[0100] The raw drilling parameters automatically collected by the intelligent rock drilling rig are integrated, and the coordinates, feed rate, impact pressure, propulsion pressure and rotation pressure of each borehole are extracted. The above parameters are then standardized into fixed lengths.

[0101] The drilling parameter data is cleaned, denoised, and standardized to remove missing values ​​and outliers. Core features of each borehole (such as spatial coordinates, feed rate, impact pressure, propulsion pressure, and rotational torque) are extracted, and the parameter set length is standardized to meet model input requirements. Figure 2 As shown, this is a specification for tunnel face images, including rotation, cropping, etc.

[0102] Simultaneously, the mileage information and geological features corresponding to the working face are extracted from the geological sketch;

[0103] Using mileage and borehole number as index labels, a multimodal dataset consisting of face images, drilling parameters, and surrounding rock classification information is constructed.

[0104] Furthermore, textual structure extraction was performed on the geological sketches to obtain the corresponding rock mass level labels for the images. Regarding the images, the acquired images were cropped and scaled to ensure the complete representation of the tunnel face area (arch, sidewalls, and floor) and to achieve spatial alignment between the images and borehole coordinates. Finally, using the construction mileage and borehole number as indexes, the images, parameters, and labels were uniformly encoded to construct a structured and accurately labeled multimodal dataset.

[0105] S4. Based on a multimodal dataset, construct an intelligent identification model for surrounding rock grade using multimodal data fusion, and intelligently identify the surrounding rock grade at the working face. The intelligent identification model for surrounding rock grade using multimodal data fusion is constructed as follows: Figure 3 As shown, image features and drilling parameters are extracted and fused at three scales: local, regional, and global.

[0106] The intelligent identification model is a multimodal fusion identification model with a three-layer cross-modal attention pyramid structure, specifically including:

[0107] 1) The underlying layer - GeoDeformNet network, such as Figure 4 As shown, this network focuses on local features in regions of abrupt changes in lithological texture and drilling parameters within an image. It fuses image pixels with corresponding borehole drilling parameters via a cross-modal attention mechanism, emphasizing regions of abrupt changes in lithological texture and points of dramatic parameter shifts. Deformable convolution is used to extract irregular texture features, and pixel-to-borehole coordinate mapping is employed to achieve fine-grained alignment, improving local recognition accuracy.

[0108] 2) The middle layer – GeoPyramidGAT network – is used to extract regional features of joint distribution trends and local variation trends of drilling parameters in the image; such as... Figure 5 As shown, a superpixel segmentation algorithm is used to divide the image into multiple geological structure regions. Texture features of each region and the drilling parameter statistics (mean, variance, etc.) of the covered boreholes are extracted to construct a region-level map structure, such as... Figure 6 As shown, this structure is input to a Graph Attention Network (GAT) to learn the structural and semantic relationships between regions and capture fusion features at a mid-scale.

[0109] 3) Top-level - GeoAugmented Vision Mamba network, such as Figure 7 As shown, the GAV-Mamba network is used to capture the overall stability features of the image and the global distribution characteristics of the main variables in the drilling parameters. The structure is improved based on the Vision Mamba model. The drilling parameters are mapped to a two-dimensional matrix with the same size as the image and fused with the image feature map to construct a global fusion input. This model focuses on the overall stability of the surrounding rock structure and the macroscopic distribution pattern of the drilling parameters.

[0110] The three layers of features are dynamically weighted and fused through a gating mechanism, adjusting the weights of information at each scale according to different geological conditions to enhance the model's adaptability and robustness. Finally, the geological feature results are output through a fully connected layer, enabling intelligent and automated identification of geological features at the working face.

[0111] Furthermore, the GeoDeformNet network includes:

[0112] (1) Across attention layers, focus on the abrupt change regions of drilling parameters between adjacent boreholes;

[0113] First, select a borehole that is close to the current borehole by calculating the Euclidean geometric distance.

[0114] D ij =||X i -X j ||2

[0115] Where D ij X represents the Euclidean distance between the i-th and j-th boreholes. i ,X j ∈R 2 , indicating two different borehole coordinates;

[0116] Next, the statistical characteristics of the selected boreholes and the current borehole are calculated, specifically the difference in mean, standard deviation, and maximum absolute value. The calculation formulas are as follows:

[0117]

[0118] δ i =max|P i -P j |

[0119] Where μ i σ i δ i The mean difference, standard deviation, and maximum absolute difference of the i-th borehole are represented; k represents the number of boreholes involved in the calculation, which defaults to 2; P represents the original drilling parameter matrix.

[0120] The final output features are:

[0121] F drill,i =[P i ,μ i ,σ i ,δ i ]∈R 4d

[0122] Where F drill,i P represents the encoded borehole feature of the i-th borehole. i ,μ i ,σ i ,δ i The original parameters of the i-th borehole are represented by their respective differences: mean difference, standard deviation, and maximum absolute difference. [,] denotes vector concatenation, R represents the real number field, and d represents the dimension, i.e., P mentioned earlier. i ,μ i ,σ i ,δ i These four dimensions;

[0123] (2) Deformable convolutional layers are used to extract local features of abrupt changes in lithological texture in images;

[0124] The specific definition of deformable convolutional layers is as follows:

[0125] F1 = ReLU(Conv1(I))

[0126] Δ = OffsetConv(F1)

[0127] F2 = ReLU(DeformConv(F1,Δ))

[0128] F3 = MaxPool(F2)

[0129] F img =ReLU(Conv2(F3))

[0130] Where I represents the input face image, F1 represents a feature extracted from the original image convolutional layer and activated, Δ represents a learnable offset defined by F1 for constructing deformable layers, F2 represents a feature extracted from the input features by a deformable convolutional layer and activated, and F3 represents a feature obtained from the input features through max pooling. img The features are obtained from the input features via ReLU activation;

[0131] (3) A region attention mechanism is used to achieve a fine correspondence between image features and borehole drilling parameters based on pixel coordinates. To facilitate subsequent coordinate correspondence, the borehole coordinates need to be normalized first, and then linearly mapped to project them onto a unified dimension for alignment.

[0132]

[0133] in, This represents the coordinates normalized to [-1, 1], W' and H' represent the dimensions after image processing, and GridSample() represents a sampling layer used to extract features from the image at the corresponding location based on the coordinates. The extracted drilling parameters and image features are aligned to facilitate subsequent fusion. The specific alignment steps are as follows:

[0134]

[0135]

[0136] in P represents the input borehole coordinate matrix. b This represents the original, input drilling parameters. This represents the aligned characteristics of the drilling parameters. This represents the drilling parameter features processed by the cross-attention layer. `DrillParamEncoder()` is a custom cross-attention layer used to extract differential features between adjacent boreholes during drilling. p Let b represent a trainable projective weight matrix. p Represent a bias matrix;

[0137] Finally, the aligned drilling features and image features are fused using a multilayer perceptron.

[0138]

[0139] in The fused feature matrix, These represent the image features before fusion and the features during drilling, respectively.

[0140] Furthermore, the GeoPyramidGAT network performs preliminary structural surface division on the facet image based on the superpixel segmentation method, and records the set of pixel coordinates for each region:

[0141] S = SLIC(I,n)

[0142] Ω k ={(x,y)|S(x,y)=k}

[0143]

[0144] Where SLIC() represents dividing the original image I into n blocks using a superpixel segmentation algorithm, Ω k Used to record the edge pixel coordinates of each image block, l i Used to map the recorded edge coordinates to the borehole position, thereby dividing the area described by the borehole; Represents the x and y coordinates of each point on the superpixel block;

[0145] Next, we calculate the statistical characteristics of the borehole drilling parameters in this area, including the mean, standard deviation, minimum, and maximum values:

[0146]

[0147] min kd =minf id

[0148] maxkd=maxf id

[0149] Where μ kd σ kd min kd ,max kd This represents the mean, standard deviation, minimum, and maximum values ​​of the drilling parameters in the corresponding region; f id This represents the drilling parameters for the i-th borehole within this region;

[0150] Next, feature extraction needs to be performed on the image patches. The specific steps are as follows:

[0151]

[0152] PyramidPooling() represents the processing of image block I. k Pyramid pooling is used to transform it into a three-layer image feature pyramid;

[0153] Finally, a graph structure is constructed, where each node represents an image region. The node features include image region features and drilling parameter statistics. The specific steps are as follows:

[0154]

[0155] E = {(v k ,v j )|v j ∈KNN(C k ,K),j≠k}

[0156] in This represents the statistical characteristics of the current region during drilling, obtained through the above calculations. h represents the features of the current region image patch obtained through the above calculations. k C represents the stitched statistical features and image characteristics after drilling. k V represents the pixel coordinates of the graph nodes obtained through calculation, E represents the set of edges between the constructed nodes, and v k ,v j Representing two distinct graph nodes, v j ∈KNN(C k ,K) represents the number of edges that the current node should be adjacent to according to the K-nearest neighbor algorithm, where K is a hyperparameter of the K-nearest neighbor algorithm; specifically, the number of edges of a node is determined by its position in pixel coordinates, and the closer it is to the center of the image, the more edges it is connected to.

[0157] Finally, the constructed graph structure is fed into the graph neural network as input to learn region-level fusion features. The core approach consists of a multi-head attention layer and a single-head attention layer. The specific implementation of multi-head attention is as follows:

[0158]

[0159] in It is the weight matrix of the h-th attention head;

[0160] The attention coefficient is calculated as follows:

[0161]

[0162] Finally, the h attention heads are concatenated and activated using the ELU activation function.

[0163]

[0164] The approach to single-head attention is as follows:

[0165]

[0166] The definition of the attention coefficient is similar to that of multi-head attention, as follows:

[0167]

[0168] Furthermore, the GAV-Mamba network maps the pixel coordinates of the image to the corresponding drilling parameters, creating a two-dimensional structure for the drilling parameters that matches the image size. Specifically, this includes the following:

[0169] Define image network coordinates:

[0170] G={(x,y)∣x∈{0,1,...,W-1},y∈{0,1,...,H-1}}

[0171] Where W represents the width of the face image and H represents the height of the face image;

[0172] For each channel containing the drilling parameters, the sparse parameter values ​​are interpolated in the two-dimensional image domain:

[0173]

[0174] Interp method () indicates the two-dimensional interpolation method used;

[0175] Finally, the normalized image and the interpolation parameter map are stitched together along the channel dimension.

[0176]

[0177] I represents the input facet image. This represents the feature matrix of drilling parameters obtained after the above steps; the tensor F is used to extract global semantic features and train the model.

[0178] The accuracy of the model used in this embodiment for identifying the surrounding rock grade is compared with that of other existing models. Figure 8 As shown in the figure, the multimodal fusion identification model with a three-layer cross-modal attention pyramid structure used in this invention has higher identification accuracy than other comparative models.

[0179] This implementation method achieves deep integration of visual information and drilling parameters through a three-layer fusion structure, effectively improving the accuracy and stability of surrounding rock classification. It is suitable for intelligent construction assistance systems for tunnels under complex geological conditions, providing efficient and reliable technical support for geological identification and construction decision-making.

[0180] Based on the same inventive concept as the above-described method embodiments, this application also provides a dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification system for geological features of tunnel faces. This system can achieve the functions provided by the above-described method embodiments, such as... Figure 9 As shown, the system includes the following modules:

[0181] Image acquisition module 110: Acquires images of the tunnel face after blasting is completed at the tunnel site;

[0182] Data collection module 120: Collects and organizes on-site construction data, including raw drilling parameters and corresponding geological sketch information of the drilling face;

[0183] The multimodal dataset construction module 130 integrates and cleans the original drilling parameters, and constructs a multimodal dataset by combining the geological feature information extracted from the face image and geological sketch.

[0184] The intelligent identification module 140 constructs an intelligent identification model for the surrounding rock level based on multimodal datasets and performs intelligent identification of the surrounding rock level at the working face.

[0185] This invention constructs an intelligent analysis model for tunnel face photographs and drilling parameters based on multimodal data fusion theory, enabling intelligent identification of geological features at the tunnel face. By fusing multidimensional information from tunnel face images and drilling parameters, this model significantly improves the accuracy and robustness of geological feature identification, optimizes stratum assessment and blasting parameter adjustment during construction, provides more precise decision support for subsequent construction planning, and effectively promotes the intelligent development of tunnel construction.

[0186] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0187] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0188] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0189] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0190] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0191] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0192] The terms "first" and "second" used in the embodiments are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" can be interchanged in a specific order or sequence where permitted. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0193] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces, characterized in that, Includes the following steps: S1. Collect images of the tunnel face after the blasting is completed. S2. Collect and organize on-site construction data, including raw drilling parameters and corresponding geological sketch information of the working face; S3. Integrate and clean the original drilling parameters, and construct a multimodal dataset by combining the geological feature information extracted from the face image and geological sketch. S4. Based on the multimodal dataset, construct a multimodal data fusion-based intelligent identification model for surrounding rock level, and intelligently identify the surrounding rock level at the working face. The intelligent identification model for surrounding rock level is a multimodal fusion identification model with a three-layer cross-modal attention pyramid structure, specifically including: 1) Bottom layer - GeoDeformNet network, used to focus on local features in areas of abrupt changes in lithological texture and drilling parameters in the image; This network fuses image pixels with the corresponding drilling parameters of the borehole through a cross-modal attention mechanism, focuses on areas of abrupt changes in lithological texture and points of dramatic parameter changes, extracts irregular texture features by combining deformable convolution, and achieves fine-grained alignment by mapping pixels to borehole coordinates. 2) The mid-layer GeoPyramidGAT network is used to extract regional features of joint distribution trends and local variation trends of drilling parameters in the image; the superpixel segmentation algorithm is used to divide the image into multiple geological structure surface regions, extract the texture features of each region and the drilling parameter statistics of the boreholes it covers, and construct a regional-level graph structure. This structure is input to the graph attention network GAT to learn the structural and semantic relationships between regions and capture the fusion features at the mid-layer scale. 3) Top-level - GeoAugmented Vision Mamba network, namely GAV-Mamba network, is used to capture the overall stability features of the image and the global distribution characteristics of the main variables in the drilling parameters; the structure is improved based on the vision mamba model, the drilling parameters are mapped to a two-dimensional matrix with the same size as the image, and fused with the image feature map to construct a global fusion input; The three layers of features are dynamically weighted and fused through a gating mechanism. The weights of information at each scale are dynamically adjusted according to different geological conditions to enhance the adaptability and robustness of the model. Finally, the geological feature results are output through a fully connected layer.

2. The dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces according to claim 1, characterized in that: In step S1, when acquiring images of the tunnel face after blasting in the drill-and-blast method, after ventilation for at least 30 minutes, a 10,000 Lux diffused light source is placed 10-12 meters away from the tunnel face and 1.5 meters on both sides of the tunnel centerline. Images are captured using an image acquisition device with a resolution of at least 20 million pixels. Multiple images are acquired from the same tunnel face, and the image with the highest quality is selected and saved as the final image input.

3. The dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces according to claim 1, characterized in that: In step S2, the original drilling parameters include indicators such as propulsion pressure, impact pressure, rotational torque, and feed rate, which are used to reflect the hardness and degree of fracturing of the surrounding rock; the geological sketch information includes information on lithology, joint and fracture development, and groundwater conditions, which serve as the source of subsequent identification labels.

4. The dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces according to claim 1, characterized in that: In step S3, the raw drilling parameters automatically collected by the intelligent rock drilling rig are integrated, and the coordinates, feed rate, impact pressure, propulsion pressure and rotation pressure of each borehole are extracted, and the above parameters are standardized into fixed lengths. Simultaneously, the mileage information and geological features corresponding to the working face are extracted from the geological sketch; Using mileage and borehole number as index labels, a multimodal dataset consisting of face images, drilling parameters, and surrounding rock classification information is constructed.

5. The dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces according to claim 1, characterized in that: The GeoDeformNet network includes: (1) Across attention layers, focus on the abrupt change regions of drilling parameters between adjacent boreholes; First, select a borehole that is close to the current borehole by calculating the Euclidean geometric distance. ; in Indicates the first The first borehole and the first The Euclidean distance of a drill hole , indicating two different borehole coordinates; Next, the statistical characteristics of the selected boreholes and the current borehole are calculated, specifically the difference in mean, standard deviation, and maximum absolute value. The calculation formulas are as follows: ; ; ; in Indicates the first The mean difference, standard deviation, and maximum absolute difference of each borehole. k Indicates the number of boreholes included in the calculation; Represents the original drilling parameter matrix; The final output features are: ; in Indicates the first The encoded borehole features The numbers represent the order of the numbers. The original parameter characteristics of each borehole, including mean difference, standard deviation, and maximum absolute value difference. This represents vector concatenation. R Represents the real number field. d Indicates dimension; (2) Deformable convolutional layers are used to extract local features of abrupt changes in lithological texture in images; The specific definition of deformable convolutional layers is as follows: ; ; ; ; ; in This represents the input image of the face of the machine. This represents a feature extracted from a convolutional layer of the original image and obtained after activation. Indicate a basis A learnable offset is defined for constructing deformable layers. This represents a feature extracted from a deformable convolutional layer of input features and then activated. This represents a feature obtained from the input features through max pooling. The features are obtained from the input features via ReLU activation; (3) Region attention mechanism, used to achieve a fine correspondence between image features and borehole drilling parameters based on pixel coordinates; firstly, the borehole coordinates need to be normalized, and then linearly mapped to project them onto a unified dimension for alignment: ; ; in, This represents coordinates normalized to [-1, 1]. Indicates the size after image processing. This represents a sampling layer used to extract features from the image at corresponding locations based on coordinates. The extracted drilling parameters and image features are aligned to facilitate subsequent fusion. The specific alignment steps are as follows: ; ; in This represents the input borehole coordinate matrix. This represents the original, input drilling parameters. This represents the aligned characteristics of the drilling parameters. This represents the characteristics of the drilling parameters after processing across the attention layer. It is a custom cross-attention layer used to extract differential features of adjacent boreholes during drilling. This represents a trainable projected weight matrix. Represent a bias matrix; Finally, the aligned drilling features and image features are fused using a multilayer perceptron. ; in The fused feature matrix, These represent the image features before fusion and the features during drilling, respectively.

6. The dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces according to claim 1, characterized in that: The GeoPyramidGAT network uses a superpixel segmentation method to initially divide the structural surfaces of the facet image, recording the set of pixel coordinates for each region: ; ; ; in This indicates that the original image is segmented using a superpixel segmentation algorithm. Divided into piece, Used to record the edge pixel coordinates of each image block. This is used to map the recorded edge coordinates to the borehole positions, thereby delineating the area described by the borehole. Represents the x and y coordinates of each point on the superpixel block; Next, we calculate the statistical characteristics of the borehole drilling parameters in this area, including the mean, standard deviation, minimum, and maximum values: ; ; ; ; in This represents the mean, standard deviation, minimum, and maximum values ​​of the drilling parameters in the corresponding region. Indicates the first in this region i Drilling parameters for each borehole; Next, feature extraction needs to be performed on the image patches. The specific steps are as follows: ; in Indicates a block of images Pyramid pooling is used to transform it into a three-layer image feature pyramid; Finally, a graph structure is constructed, where each node represents an image region. The node features include image region features and drilling parameter statistics. The specific steps are as follows: ; ; ; ; in This represents the statistical characteristics of the current region during drilling, obtained through the above calculations. This represents the features of the image patch in the current region obtained through the above calculations. This represents the stitched statistical features and image characteristics obtained during drilling. This represents the pixel coordinates of the graph node obtained through calculation. This represents the set of edges between the constructed nodes. This represents two different graph nodes. This represents the number of edges that the current node should be adjacent to according to the K-nearest neighbor algorithm, where K is a hyperparameter of the K-nearest neighbor algorithm; Finally, the constructed graph structure is fed into the graph neural network as input to learn region-level fusion features.

7. The dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces according to claim 1, characterized in that: The GAV-Mamba network maps the pixel coordinates of the image to the corresponding drilling parameters, creating a two-dimensional structure for the drilling parameters that matches the image size. Specifically, this includes the following: Define image network coordinates: ; in This indicates the width of the facet image. Indicates the height of the working face image; For each channel containing the drilling parameters, the sparse parameter values ​​are interpolated in the two-dimensional image domain: ; in This indicates the two-dimensional interpolation method used; Finally, the normalized image and the interpolation parameter map are stitched together along the channel dimension. ; This represents the input image of the face of the machine. The obtained drilling parameter characteristic matrix is ​​represented; the tensor is obtained. F Used to extract global semantic features and train the model.

8. A system for a dual-modal, three-layer, multi-scale, multi-task fusion intelligent identification method for geological features of tunnel faces according to any one of claims 1-7, characterized in that: Includes the following modules: Image acquisition module (110) acquires images of the tunnel face after blasting is completed; Data collection module (120) collects and organizes on-site construction data, including raw drilling parameters and geological sketch information of the corresponding face; The multimodal dataset construction module (130) integrates and cleans the original drilling parameters, and constructs a multimodal dataset by combining the geological feature information extracted from the face image and geological sketch. The intelligent identification module (140) constructs a multimodal data fusion-based intelligent identification model for the surrounding rock level and performs intelligent identification of the surrounding rock level at the working face.