Tunnel face image descriptive model construction method, system, device and medium
By extracting geological descriptions from tunnel face images using deep learning algorithms, a Bert tunnel face image description model was constructed. This solved the problems of low efficiency and poor accuracy in recording geological information during tunnel construction, and enabled the direct generation of standardized geological descriptions from images. It is highly adaptable and reduces deployment costs.
Patent Information
- Application Number
- CN202511499954.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-10-21
AI Technical Summary
In current tunnel construction, the acquisition of geological information at the tunnel face relies on manual recording, which is inefficient, lacks standardization, and lacks models that can automatically generate standardized geological descriptions. In particular, the accuracy is difficult to guarantee under complex conditions such as insufficient light and dust.
Geological descriptions are extracted from tunnel face images using deep learning algorithms. By combining image preprocessing and feature extraction, a Bert tunnel face image description model is constructed to directly generate standardized geological descriptions from images.
It improves the efficiency and accuracy of geological information recording during tunnel construction, simplifies the multi-model construction process, reduces deployment costs, and is highly adaptable to different lighting and dust conditions.
Smart Images

Figure CN120976772B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of tunnel engineering and computer vision technology, and more specifically, to a method, system, device, and medium for constructing descriptive models of tunnel face images. Background Technology
[0002] During tunnel construction, geological information at the tunnel face is crucial for assessing surrounding rock stability, determining support parameters, and optimizing construction techniques. Traditional geological logging at the tunnel face relies primarily on manual on-site observation and recording, or is completed separately using multiple independent models such as lithology identification models, structural plane identification models, and groundwater identification models. However, these methods suffer from low efficiency, inconsistent standards, complex model chains, and high misclassification rates. Especially in environments with insufficient lighting, fractured tunnel face surfaces, and indistinct geological features, the accuracy and stability of manual identification and multi-model joint identification are difficult to guarantee. In recent years, with the development of deep learning and computer vision, automated geological information extraction from tunnel face images has become possible. However, existing research largely remains at the level of single tasks (such as fracture detection and lithology identification), lacking an end-to-end technical approach for directly generating complete and standardized geological descriptions from tunnel face images. Summary of the Invention
[0003] The purpose of this invention is to provide a method, system, device, and medium for constructing image descriptive models of tunnel faces, so as to realize the automated extraction, standardized processing, and high-precision description generation of geological information of tunnel faces, thereby improving the efficiency and accuracy of recording geological information during tunnel construction. This invention addresses the shortcomings of existing methods for acquiring geological information of tunnel faces, which rely on manual recording, have low information standardization, and whose image quality is severely affected by factors such as on-site lighting and dust, and lack a model construction method that can automatically generate standardized geological descriptions.
[0004] The above-mentioned technical objective of the present invention is achieved through the following technical solution:
[0005] Firstly, this application provides a method for constructing a descriptive model of a tunnel face image, including the following specific steps:
[0006] Obtain geological sketches of the tunnel face at each mileage during the tunnel construction process, as well as images of the tunnel face at each corresponding mileage.
[0007] Geological descriptions were identified and extracted from geological sketches of various working faces using deep learning algorithms, and the geological descriptions were then standardized.
[0008] The images of each tunnel face are preprocessed, and high-dimensional features are extracted from the preprocessed tunnel face images using a deep learning algorithm.
[0009] Training samples are constructed by using high-dimensional features of each image and corresponding geological descriptions that have been standardized.
[0010] The initial model is trained using training samples until it reaches the preset training termination condition. The initial model that reaches the preset training termination condition is then identified as the Bert tunnel face image description model.
[0011] Based on the above technical solution, the present invention can be further improved as follows.
[0012] Furthermore, the above geological description was obtained through the following methods:
[0013] The trained YOLOv5 object detection model was used to detect the text description region of the geological sketch at the working face, and the detected region was obtained.
[0014] The Convolutional Recurrent Neural Network algorithm is used to extract text from the detected area to obtain a geological description.
[0015] Furthermore, the images of each tunnel face are preprocessed, including image quality enhancement and brightness / color restoration. Image quality enhancement is performed using the Real-ESRGAN model, and brightness / color restoration is performed using the Retinex-Net algorithm.
[0016] Furthermore, the training samples mentioned above include bimodal datasets for each mileage. The bimodal datasets are formed by pairing high-dimensional image features of each mileage with standardized geological descriptions of the corresponding mileage.
[0017] Furthermore, the high-dimensional features of the aforementioned images are extracted using a Compact Convolutional Transformer model, combined with convolution and Transformer.
[0018] Secondly, this application provides a system for constructing a descriptive model of a tunnel face image, applicable to the method for constructing a descriptive model of a tunnel face image according to any one of the first aspects, including:
[0019] The construction data acquisition module is used to acquire geological sketches of the tunnel face at each mileage during the tunnel construction process, as well as images of each tunnel face at the corresponding mileage.
[0020] The geological sketching processing module is used to identify and extract geological descriptions from geological sketches of various working faces using deep learning algorithms, and to standardize the geological descriptions.
[0021] The high-dimensional feature extraction module is used to preprocess the images of each tunnel face and extract high-dimensional features from the preprocessed tunnel face images using a deep learning algorithm.
[0022] The training sample construction module is used to construct training samples by using the high-dimensional features of each image and the corresponding mileage and standardized geological descriptions.
[0023] The description model determination module is used to train a preset initial model using training samples until the initial model reaches a preset training termination condition. The initial model that reaches the preset training termination condition is then determined as the Bert tunnel face image description model.
[0024] Furthermore, in the aforementioned construction data acquisition module, the geological description is obtained through the following methods:
[0025] The trained YOLOv5 object detection model was used to detect the text description region of the geological sketch at the working face, and the detected region was obtained.
[0026] The Convolutional Recurrent Neural Network algorithm is used to extract text from the detected area to obtain a geological description.
[0027] Furthermore, the high-dimensional feature extraction module described above preprocesses the images of each tunnel face, including image quality enhancement and brightness / color restoration.
[0028] Image quality enhancement is performed using the Real-ESRGAN model;
[0029] Brightness and color restoration processing was performed using the Retinex-Net algorithm.
[0030] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for constructing a descriptive model of a tunnel face image according to any one of the first aspects.
[0031] Fourthly, this application provides a non-transitory computer-readable storage medium that stores computer instructions that cause a computer to execute the tunnel face image descriptive model construction method of any one of the first aspects.
[0032] Compared with the prior art, the present invention has at least the following beneficial effects:
[0033] 1. This invention realizes the technical route of directly generating geological sketch text from tunnel face images, avoiding the cumbersome process of constructing multiple independent models such as lithology identification, structural surface identification, and groundwater identification in traditional methods, and forming an integrated processing path of image input and text output.
[0034] 2. By eliminating the training and integration of multiple models, the overall system structure is simpler, and the deployment and operation and maintenance costs are significantly reduced, which is conducive to rapid deployment and long-term stable operation at the construction site.
[0035] 3. By generating standardized text descriptions through a unified model, discrepancies in results between different recorders or different models are avoided, ensuring the uniformity and standardization of geological information records and improving the usability of data in subsequent analysis and archiving.
[0036] 4. The one-step conversion mode reduces intermediate data processing steps, making geological information generation more real-time; at the same time, the method can still maintain high robustness under complex construction conditions such as different lighting and dust, and has stronger adaptability. Attached Figure Description
[0037] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings:
[0038] Figure 1 This is a flowchart of the construction method in an embodiment of the present invention;
[0039] Figure 2 This is a schematic diagram illustrating the standardization of geological descriptions in an embodiment of the present invention;
[0040] Figure 3 This is a schematic diagram of the enhanced resolution image of the tunnel face in an embodiment of the present invention. Figure 3 In the image, (a) is a low-quality image in one scenario, (c) is a low-quality image in another scenario, and (b) and (d) correspond to (a) and (c) respectively and are super-resolution enhanced images.
[0041] Figure 4 This is a schematic diagram of the tunnel face image after illumination enhancement in an embodiment of the present invention. Figure 4 In the image, (a) is a low-light image in one scene, (c) is a low-light image in another scene, and (b) and (d) correspond to (a) and (c) respectively and are light-enhanced images;
[0042] Figure 5 This is a flowchart illustrating the training process of the Bert tunnel face image description model in an embodiment of the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0044] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0045] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0046] In the description of the embodiments of the present invention, "multiple" means at least two.
[0047] Example 1: Since there is currently a lack of a technical method to directly generate standardized geological sketch text from tunnel face images, this example provides a method for constructing descriptive models of tunnel face images. This method should feature automated data acquisition, intelligent geological description extraction, refined image enhancement and feature extraction, and standardized text generation to improve the efficiency and accuracy of geological logging, reduce the influence of subjective human factors, and promote the informatization and intelligent development of tunnel construction. Figure 1 As shown, the method includes the following specific steps:
[0048] S1, acquire geological sketches of the tunnel face at each mileage during the tunnel construction process, as well as images of the tunnel face at each corresponding mileage.
[0049] During tunnel excavation, construction workers will draw geological sketches of the tunnel face after each construction mileage. These geological sketches include: geological personnel or supervision units drawing geological sketches of the tunnel face according to specifications during tunnel excavation, and making detailed annotations on rock strata distribution, joints and fissures, faults, water content, etc.; after the tunnel construction is completed, the geological sketches of the tunnel face of each corresponding mileage section of the entire line will be collected and organized to ensure that the sketches are clear and identifiable and the annotations are complete, so as to provide a reliable data basis for subsequent extraction and analysis of textual descriptions of the tunnel face. The geological sketches of the tunnel face are shown in Table 1.
[0050] Table 1
[0051]
[0052] The geological conditions are described as follows: the rock at the tunnel face is sandstone, purplish-red, weakly weathered (W), thickly layered, with the bedding planes intersecting the tunnel axis at a large angle, dipping to the left of the tunnel at the end of the long mileage section, with a dip angle of about 39 degrees. There is right-side bedding-parallel bias at the tunnel face, joints and fissures are relatively well-developed, the degree of bonding is generally moderate, the rock mass is relatively broken to relatively intact, the surrounding rock has an inlaid fragmented structure, the tunnel face is damp, and the exposed arch is prone to falling off. It is recommended that the surrounding rock be classified as Class III.
[0053] Furthermore, in order to construct data corresponding to the images and text, real-world images of the tunnel face corresponding to the geological sketch mileage are simultaneously collected during the tunnel excavation process. Image acquisition can be done using high-definition industrial cameras, SLR cameras, or existing monitoring equipment at the construction site to ensure that the features of the tunnel face are fully presented in the images and that they completely correspond to the mileage number of the aforementioned sketch data.
[0054] S2 uses deep learning algorithms to identify and extract geological descriptions from geological sketches of each face, and then standardizes these descriptions.
[0055] Optionally, the geological descriptions mentioned above can be automatically extracted from the collected geological sketches of the tunnel face based on deep learning object detection and text recognition algorithms. Specifically, this can be achieved through the following methods:
[0056] S21. The trained YOLOv5 object detection model is used to detect the text description area of the geological sketch of the working face, and the detection area is obtained.
[0057] The detection and localization algorithm includes the following steps:
[0058] YOLOv5 performs convolution, pooling, and SiLU activation as shown in formulas (1)-(3), and finally obtains the four coordinate positions of the detection box through the Sigmoid function as shown in formula (4).
[0059] (1)
[0060] In the formula, This represents the pixel value at depth k at position (i, j) in the input feature map of the convolutional layer. This represents the pixel value at depth p at position (i+m, j+n) in the input data. This represents the weight value at position (m, n) in the convolution kernel with depth p and at depth k in the output feature map. represents the bias term of depth k on the output feature map, and f represents the kernel size.
[0061] further, (2)
[0062] In the formula, Represents the pooling function (such as max pooling and average pooling). Indicates the input tensor All values within the pooling window.
[0063] Furthermore, in the above, (3)
[0064] In the formula, This represents the output of the activation function. Indicates input, e It is a mathematical constant, approximately equal to 2.71828.
[0065] Furthermore, in the above: (4)
[0066] in, It is the input value. These are the pixel values output by the model, and their values are between 0 and 1.
[0067] S22, the Convolutional Recurrent Neural Network (CRNN) algorithm is used to extract text from the detected area to obtain a geological description. CRNN is a hybrid architecture combining Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN), which can be used to process sequential and image data. It can simultaneously capture the spatial structural features of the data (such as local patterns in images) and sequence dependencies (such as temporal dependencies in text).
[0068] The detection output includes the bounding box coordinates of each text region; then, the Convolutional Recurrent Neural Network (CRNN) algorithm is used to perform optical character recognition on the detected text region image to accurately extract the text information and obtain the original geological description text. Through convolution operation, pooling operation, and ReLU activation as shown in formulas (5)-(7), the extracted text is output through the text extraction formula as shown in formula (8).
[0069] (5)
[0070] In the formula, This represents the pixel value at depth k at position (i, j) in the input feature map of the convolutional layer. This represents the pixel value at depth p at position (i+m, j+n) in the input data. This represents the weight value at position (m, n) in the convolution kernel with depth p and at depth k in the output feature map. represents the bias term of depth k on the output feature map, and f represents the kernel size.
[0071] Furthermore, in the above: (6)
[0072] in, Represents the pooling function (such as max pooling and average pooling). Indicates the input tensor All values within the pooling window.
[0073] Furthermore, in the above: (7)
[0074] in, This represents the output of the activation function. This indicates input.
[0075] Furthermore, in the above: , (8)
[0076] in, This represents the convolutional feature sequence extracted using formulas (5)-(7). This represents the hidden state of an RNN. Represents the output text sequence; RNN stands for Neural Network (Continuous Function Mapping). It is a probabilistic sequence decoding algorithm.
[0077] In steps S21-S22 above, the YOLOv5 object detection model, which has been trained with annotations, is first used to accurately detect and locate the text description areas in the geological sketch of the tunnel face, ensuring that the text box areas can be accurately identified even in complex line backgrounds. Then, the Convolutional Recurrent Neural Network (CRNN) algorithm is used to perform character-level recognition and sequence modeling on the detected text areas, so as to realize the automatic extraction of handwritten or printed text, thereby automatically obtaining the original geological description text information.
[0078] Optionally, during the standardization process of geological descriptions, differences in terminology, expressions, and formats used by different construction personnel in geological sketching lead to inconsistencies in the style of the original geological description text. This embodiment utilizes a large language model (such as DeepSeek) to standardize the extracted original text, such as... Figure 2As shown, the geological description is standardized into a standardized format, including standardized expressions of elements such as lithology, color, weathering degree, structural plane occurrence, groundwater conditions, surrounding rock structural characteristics, and stability, thereby ensuring the consistency and comparability of data in subsequent analysis and modeling.
[0079] S3 preprocesses the images of each tunnel face and extracts high-dimensional features from the preprocessed tunnel face images using a deep learning algorithm.
[0080] Due to factors such as insufficient lighting, dust interference, and camera shake at the construction site, the acquired images of the tunnel face may suffer from low resolution, insufficient brightness, and color distortion. Therefore, preprocessing is required. See the schematic diagrams of the images before and after processing. Figures 3-4 ,exist Figure 3 In the image, (a) and (c) are low-quality images, while (b) and (d) correspond to (a) and (c) respectively and are super-resolution enhanced images. Figure 4 In the image, (a) and (c) are low-light images, and (b) and (d) correspond to (a) and (c) respectively and are light-enhanced images.
[0081] Optionally, the above preprocessing of each tunnel face image includes image quality enhancement and brightness / color restoration. Image quality enhancement is performed using the Real-ESRGAN model; brightness / color restoration is performed using the Retinex-Net algorithm. Real-ESRGAN is an open-source image / video super-resolution algorithm, which can be used to perform image restoration and resolution enhancement by downloading pre-compiled programs or GitHub source code. The Retinex-Net algorithm combines the ideas of deep learning and Retinex theory, and can effectively improve the visual quality of images. The RetinexNet algorithm can enhance and dehaze images by simulating the processing method of the human visual system.
[0082] In this embodiment, the Real-ESRGAN model can be used to perform super-resolution reconstruction of the image resolution, improve the image detail performance, and make the micro-cracks and rock texture clearer; at the same time, Retinex-Net is used to repair the brightness and color of the image, improve the uneven lighting and color deviation caused by construction lights, dust, water mist, etc., and restore the real color and texture features of the tunnel face under natural lighting conditions as much as possible; these two models still use convolution operation, pooling operation, fully connected operation, ReLU activation, as shown in formula (9)-(12), and finally obtain the distribution of each pixel point through the Sigmoid function, as shown in formula (13).
[0083] (9)
[0084] in, This represents the pixel value at depth k at position (i, j) in the input feature map of the convolutional layer. This represents the pixel value at depth p at position (i+m, j+n) in the input data. This represents the weight value at position (m, n) in the convolution kernel with depth p and at depth k in the output feature map. represents the bias term of depth k on the output feature map, and f represents the kernel size.
[0085] (10)
[0086] in, Represents the pooling function (such as max pooling and average pooling). Indicates the input tensor All values within the pooling window.
[0087] (11)
[0088] in, denoted by , where W represents the output of the fully connected layer, W represents the weight matrix, and b represents the bias term.
[0089] (12)
[0090] in, This represents the output of the activation function. This indicates input.
[0091] (13)
[0092] in, It is the input value. These are the pixel values output by the model, and their values are between 0 and 1.
[0093] Optionally, the high-dimensional features of the images mentioned above are extracted using the Compact Convolutional Transformer model, which combines convolution and Transformer. The Compact Convolutional Transformer (CCT) is a hybrid architecture model that combines the advantages of Convolutional Neural Networks (CNN) and Transformers, which can improve the efficiency of computer vision tasks and reduce the requirements for data scale. The Transformer model architecture uses a Self-Attention structure to replace the RNN network structure commonly used in NLP tasks, which can perform parallel computation compared to the RNN network structure.
[0094] Transformer is essentially an Encoder-Decoder architecture; therefore, the middle part of Transformer can be divided into two parts: the encoding component and the decoding component.
[0095] In this embodiment, the Compact Convolutional Transformer (CCT) network can be used to extract features from the enhanced tunnel face image to generate a high-dimensional feature vector that can characterize key information such as lithology, structural surface features, color texture, etc. This feature vector is used as the deep feature input of the image end for subsequent image-text joint modeling. It is still extracted by convolution operation, pooling operation, fully connected operation, ReLU activation, and multi-head attention, as shown in formulas (14)-(18).
[0096] (14)
[0097] in, This represents the pixel value at depth k at position (i, j) in the input feature map of the convolutional layer. This represents the pixel value at depth p at position (i+m, j+n) in the input data. This represents the weight value at position (m, n) in the convolution kernel with depth p and at depth k in the output feature map. represents the bias term of depth k on the output feature map, and f represents the kernel size.
[0098] (15)
[0099] in, Represents the pooling function (such as max pooling and average pooling). Indicates the input tensor All values within the pooling window.
[0100] (16)
[0101] in, denoted by , where W represents the output of the fully connected layer, W represents the weight matrix, and b represents the bias term.
[0102] (17)
[0103] in, This represents the output of the activation function. This indicates input.
[0104] , ;
[0105] (18)
[0106] Where Z is the feature vector extracted by the fully connected layer. The model can learn vectors; head is the self-attention calculated between Q, K, and V; Softmax is a function operation. As a multi-head mechanism, Concat is a function operation where h represents the number of self-attention heads. It is a linear transformation.
[0107] S4 is used to construct training samples by using high-dimensional features of each image and geological descriptions of the corresponding mileage that have been standardized.
[0108] The training samples include bimodal datasets for each mileage. These bimodal datasets are formed by pairing high-dimensional image features of each mileage with standardized geological descriptions of the corresponding mileage. Specifically, the standardized geological text descriptions of the working face are paired with corresponding enhanced high-definition images of the working face to form a bimodal dataset with corresponding images and text. This dataset can also be divided into training, validation, and test sets according to a certain ratio for model training, validation, and testing, respectively.
[0109] S5. Use training samples to train the preset initial model until the initial model reaches the preset training termination condition. The initial model that reaches the preset training termination condition is determined as the Bert tunnel face image description model.
[0110] In this embodiment, the BERT model is used to construct a descriptive model of the tunnel face image; during the training process, such as Figure 5 As shown, the Bert model receives image feature vectors and corresponding text descriptions from the CCT network to achieve the fusion learning of image and text features, enabling the model to directly generate standardized geological description texts based on the face image. The model generates geological description texts through image and text feature fusion, forward computation, and predicted text encoding as shown in formulas (19)-(20).
[0111] (19)
[0112] in, The extracted image features (i.e., the high-dimensional image features mentioned above). For word vector embedding in geological descriptions, Concat is a function operation. Representing image features Word vector embeddings for geological descriptions The concatenated fused feature vector.
[0113]
[0114] ; (20)
[0115] Here, TransformerLayer represents the computation of a neural network layer. It is the hidden state matrix of the (l-1)th layer Transformer encoder, where This is a type of encoding operation used to encode text in the model's extracted results. The hidden state matrix of the l-th layer Transformer encoder. This indicates the text output generated after being encoded by the Decoder.
[0116] Specifically, once the model is trained, it can be applied in real time at the construction site to achieve automatic conversion from face images to standardized geological descriptions, thereby improving the efficiency and accuracy of geological information recording during the construction process.
[0117] The method for constructing a descriptive model of tunnel face images provided in this embodiment is based on images collected at the tunnel face construction site and corresponding geological sketches. Through steps such as image quality enhancement, text recognition, text standardization processing, and image-text feature fusion, it constructs a model that can automatically generate standardized geological descriptions. At the same time, this method realizes the technical path of directly generating geological sketch text from tunnel face images, which significantly improves the efficiency and accuracy of geological information recording during construction. It overcomes the problems of traditional manual recording, which relies on experience, is inefficient, and is prone to subjective bias. It avoids the complex process of constructing multiple geological identification models separately and has the characteristics of simple deployment and strong adaptability.
[0118] Example 2: This application provides a tunnel face image descriptive model construction system, applied to the tunnel face image descriptive model construction method of Example 1, and may include:
[0119] The construction data acquisition module is used to acquire geological sketches of the tunnel face at various mileages during tunnel construction, as well as images of the tunnel face at corresponding mileages. In this module, geological descriptions are obtained through the following methods:
[0120] The trained YOLOv5 object detection model was used to detect the text description region of the geological sketch at the working face, and the detected region was obtained.
[0121] The Convolutional Recurrent Neural Network algorithm is used to extract text from the detected area to obtain a geological description.
[0122] The geological sketching processing module is used to identify and extract geological descriptions from geological sketches of various working faces using deep learning algorithms, and to standardize the geological descriptions.
[0123] The high-dimensional feature extraction module is used to preprocess the images of each tunnel face and extract high-dimensional features from the preprocessed images using a deep learning algorithm. The preprocessing of the tunnel face images in the high-dimensional feature extraction module includes image quality enhancement and brightness / color restoration.
[0124] Image quality enhancement is performed using the Real-ESRGAN model;
[0125] Brightness and color restoration processing was performed using the Retinex-Net algorithm.
[0126] The training sample construction module is used to construct training samples by using the high-dimensional features of each image and the corresponding mileage and standardized geological descriptions.
[0127] The description model determination module is used to train a preset initial model using training samples until the initial model reaches a preset training termination condition. The initial model that reaches the preset training termination condition is then determined as the Bert tunnel face image description model.
[0128] Example 3: This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for constructing a descriptive model of a tunnel face image as described in Example 1.
[0129] Example 4: This application provides a non-transitory computer-readable storage medium that stores computer instructions that cause a computer to execute the tunnel face image descriptive model construction method of Example 1.
[0130] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0131] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0134] Those skilled in the art will understand that all or part of the steps in the above facts and methods can be implemented by a program instructing related hardware. The program or the program described therein can be stored in a computer-readable storage medium. When the program is executed, it includes the following steps: at this time, the corresponding method steps are introduced. The storage medium can be ROM / RAM, magnetic disk, optical disk, etc.
[0135] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for constructing a descriptive model of a tunnel face image, characterized in that, The specific steps include the following: Obtain geological sketches of the tunnel face at each mileage during the tunnel construction process, as well as images of the tunnel face at each corresponding mileage. Geological descriptions are identified and extracted from geological sketches of each working face using deep learning algorithms, and these geological descriptions are then standardized. Each of the tunnel face images is preprocessed, and high-dimensional features of the images are extracted from the preprocessed tunnel face images using a deep learning algorithm. Training samples are constructed using the high-dimensional features of each image and the corresponding mileage and standardized geological descriptions. The training samples are used to train the preset initial model until the initial model reaches the preset training termination condition. The initial model that reaches the preset training termination condition is determined as the Bert tunnel face image description model.
2. The method for constructing a descriptive model of a tunnel face image according to claim 1, characterized in that, The geological description was obtained through the following methods: The trained YOLOv5 object detection model was used to detect the text description region of the geological sketch at the working face, and the detected region was obtained. The Convolutional Recurrent Neural Network algorithm is used to extract text from the detected area to obtain the geological description.
3. The method for constructing a descriptive model of a tunnel face image according to claim 1, characterized in that, The images of each tunnel face are preprocessed, including image quality enhancement and brightness / color restoration. The image quality enhancement is performed using the Real-ESRGAN model, and the brightness / color restoration is performed using the Retinex-Net algorithm.
4. The method for constructing a descriptive model of a tunnel face image according to claim 1, characterized in that, The training samples include bimodal datasets for each mileage, which are formed by pairing high-dimensional image features of each mileage with standardized geological descriptions of the corresponding mileage.
5. The method for constructing a descriptive model of a tunnel face image according to claim 1, characterized in that, The high-dimensional features of the image are extracted using a Compact Convolutional Transformer model, combined with convolution and Transformer.
6. A system for constructing a descriptive model of a tunnel face image, characterized in that, include: The construction data acquisition module is used to acquire geological sketches of the tunnel face at each mileage during the tunnel construction process, as well as images of each tunnel face at the corresponding mileage. The geological sketch processing module is used to identify and extract geological descriptions from geological sketches of each working face using deep learning algorithms, and to standardize the geological descriptions. The high-dimensional feature extraction module is used to preprocess each of the tunnel face images and extract high-dimensional features from the preprocessed tunnel face images using a deep learning algorithm. The training sample construction module is used to construct training samples by using the high-dimensional features of each image and the corresponding mileage and standardized geological description. The description model determination module is used to train a preset initial model using the training samples until the initial model reaches a preset training termination condition, and then determine the initial model that has reached the preset training termination condition as the Bert tunnel face image description model.
7. The tunnel face image descriptive model construction system according to claim 6, characterized in that, In the construction data acquisition module, the geological description is obtained through the following methods: The trained YOLOv5 object detection model was used to detect the text description region of the geological sketch at the working face, and the detected region was obtained. The Convolutional Recurrent Neural Network algorithm is used to extract text from the detected area to obtain the geological description.
8. The tunnel face image descriptive model construction system according to claim 6, characterized in that, The high-dimensional feature extraction module preprocesses each of the tunnel face images, including image quality enhancement and brightness / color restoration, wherein: The image quality enhancement process is performed using the Real-ESRGAN model; The brightness and color restoration process is completed using the Retinex-Net algorithm.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for constructing a descriptive model of a tunnel face image as described in any one of claims 1-5.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the method for constructing a descriptive model of a tunnel face image as described in any one of claims 1-5.
Citation Information
Patent Citations
Intelligent geological sketching method and device for tunnel drilling and blasting method construction tunnel face
CN116934691A
Intelligent identification method, device and equipment for lithology of tunnel face and medium
CN118334505A