Tunnel lining crack identification method and system based on panoramic space image splicing

By employing a two-stage stitching method based on cylindrical projection and a global image space model, combined with Laplace pyramid fusion and LightGlue-DISK feature matching, the problems of mismatch and discontinuous stitching in tunnel lining crack detection are solved, achieving high-precision identification of tunnel lining cracks.

CN121767322APending Publication Date: 2026-03-31SHANDONG JIZAO HIGH-SPEED RAILWAY CO LTD +1
View PDF 0 Cites 4 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In tunnel lining crack detection, existing technologies are limited by traditional methods that struggle to accurately identify cracks in tunnel environments with uneven lighting and uniform textures. Furthermore, panoramic stitching is prone to issues such as mismatch, misalignment, and discontinuous stitching.

Method used

A local-global image stitching model based on cylindrical projection and a global image space model are adopted, combined with the Laplace pyramid fusion algorithm and the LightGlue-DISK feature matching method to perform local-global two-stage image stitching, and tunnel lining cracks are detected by spatial local differences.

Benefits of technology

It improves the accuracy and stability of tunnel lining crack identification, achieves panoramic stitching continuity and naturalness in weak texture environment, and enhances the pixel-level positioning capability of tunnel lining defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767322A_ABST
    Figure CN121767322A_ABST
Patent Text Reader

Abstract

The invention discloses a tunnel lining crack identification method and system based on panoramic space image splicing, and relates to the technical field of tunnel visual identification. The method comprises the following steps: acquiring a tunnel lining image array, and performing local splicing on tunnel lining images by using a splicing model based on cylindrical surface projection to obtain a local spliced image; fusing and splicing the local spliced images by using the global image space model to obtain a panoramic spliced data graph; and abnormal region detection is carried out on the panoramic splicing data graph according to the spatial local difference, and a tunnel lining crack identification result is obtained. According to the method, the lining panoramic space model based on image two-stage fusion is used for splicing, so that the visual coherence and naturalness of a global spliced image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tunnel visual recognition technology, and in particular to a method and system for identifying tunnel lining cracks based on panoramic spatial image stitching. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] During long-term service, tunnels are susceptible to various factors such as changes in geological conditions and repeated vehicle loads, which can easily lead to cracks and other defects in the lining structure. Timely and accurate identification of cracks is necessary to prevent them from expanding further and causing danger. Therefore, the technology for accurate segmentation and identification of tunnel lining cracks has always been a research hotspot in the field of tunnel engineering health monitoring.

[0004] As tunnel engineering develops towards more complex geology, longer distances, and higher demands for intelligent monitoring, current segmentation methods are unable to accurately handle small targets with long, narrow, and indistinct cracks, and are difficult to meet the accuracy requirements in practical applications.

[0005] In panoramic image stitching, weak textures and lighting effects result in monotonous textures on the plain concrete surface within tunnels and uneven lighting in transitional areas between strong and weak light. Traditional feature detection algorithms such as SIFT and SURF reduce the number of feature points extracted in weak texture areas, leading to insufficient matching accuracy and a tendency for mismatches and stitching misalignments. Traditional global stitching methods suffer from image-by-image registration error propagation, especially noticeable in long tunnels, causing significant distortion in long-distance panoramic images. While local stitching can reduce error propagation, it lacks global geometric constraints, leading to cumulative offsets between stitched blocks and affecting the overall stitching effect. Furthermore, the overlapping areas of multi-camera arrays and motion sequence images exhibit poor photometric consistency, resulting in noticeable seams and blurred high-frequency details even after multi-band fusion processing, failing to accurately reflect the global distribution characteristics of cracks.

[0006] Furthermore, in the field of panoramic anomaly detection, traditional anomaly detection methods often face the problem of missing spatial locality when dealing with surface defects in images. Existing methods typically rely on global feature similarity calculations during feature matching, neglecting the local spatial correlation of defect regions. For example, when detecting cracks on an image surface, relying solely on global feature comparison may lead to misclassification of similar textures in other normal areas as defects due to background texture interference. In addition, training with mixed multi-angle data introduces viewpoint difference noise, and multi-angle interference can make it difficult for the model to capture subtle anomaly features at specific angles, further reducing the model's detection capability. Summary of the Invention

[0007] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for identifying tunnel lining cracks based on panoramic spatial image stitching. By utilizing a panoramic spatial model of the lining based on two-stage image fusion, the visual coherence and naturalness of the global stitched image are improved.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solution: The first aspect of this invention provides a method for identifying tunnel lining cracks based on panoramic spatial image stitching, comprising the following steps: An array of tunnel lining images was acquired, and the tunnel lining images were locally stitched together using a stitching model based on cylindrical surface projection to obtain a locally stitched image. A global image space model is used to fuse and stitch together local stitched images to obtain a panoramic stitched data map. Anomaly detection was performed on the panoramic stitched data map based on local spatial differences to obtain the identification results of tunnel lining cracks.

[0009] Furthermore, the specific steps for locally stitching tunnel lining images using a stitching model based on cylindrical surface projection are as follows: Construct a projection stitching model based on a cylindrical surface; The tunnel lining image is registered using spatial prior constraints to obtain the registered image; Perform consistency optimization on camera parameters; The registered images are then merged and stitched together.

[0010] Furthermore, the specific steps for fusing and stitching locally stitched images using a global image space model are as follows: Perform anti-distortion remapping and global geometric alignment on locally stitched images; A Laplace pyramid-based image fusion algorithm is used to fuse locally stitched images to obtain a preliminary global stitched image; The initial global stitched image is optimized using an image matching method for weakly textured surfaces.

[0011] Furthermore, the specific steps for optimizing the initial global stitched image using an image matching method for weakly textured surfaces are as follows: Feature extraction is performed on the preliminary global stitched image using a weak texture feature matching method based on LightGlue; The extracted features are detected and matched using a fusion method of DISK and LightGlue.

[0012] Furthermore, the specific steps for detecting abnormal areas in panoramic stitched data images using spatial local differences are as follows: Image representation of panoramic stitched data is performed using an image anomaly and defect detection method based on spatial local correlation. Anomalies in panoramic stitched data images are extracted using a multi-scale fusion strategy.

[0013] Furthermore, the specific steps for image representation of panoramic stitched data using an image anomaly and defect detection method based on spatial local correlation are as follows: An anomaly detection model is constructed based on the PatchCore anomaly detection algorithm; A spatial locality enhancement memory is constructed, and the anomaly detection model is trained using data within the memory. We set up a spatial constraint training and testing strategy, and then used this strategy to further optimize the anomaly detection model.

[0014] Furthermore, the specific steps for extracting anomalous features from panoramic stitched data using a multi-scale fusion strategy are as follows: Multi-level feature extraction was performed on panoramic stitched data using ResNet-50; The extracted multi-level features are aggregated across local neighborhoods to obtain aggregated features; Dimensionality reduction and optimization are performed on the aggregated features.

[0015] A second aspect of the present invention provides a tunnel lining crack identification system based on panoramic spatial image stitching, comprising: The local stitching module is configured to acquire an array of tunnel lining images and use a stitching model based on cylindrical surface projection to locally stitch the tunnel lining images to obtain a locally stitched image. The global stitching module is configured to use a global image space model to fuse and stitch local stitched images to obtain a panoramic stitched data map; The anomaly detection module is configured to detect abnormal areas in the panoramic stitched data map based on local spatial differences, and obtain the identification results of tunnel lining cracks.

[0016] A third aspect of the present invention provides a computer-readable storage medium storing a computer program adapted to be loaded by a processor and to execute the steps of the tunnel lining crack identification method based on panoramic spatial image stitching as described in the first aspect of the present invention.

[0017] A fourth aspect of the present invention provides a computer device comprising: A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program, which, when executed by the processor, implements the tunnel lining crack identification method based on panoramic spatial image stitching as described in the first aspect of the present invention.

[0018] The above one or more technical solutions have the following beneficial effects: This invention discloses a method and system for identifying tunnel lining cracks based on panoramic spatial image stitching. In scenarios with uneven tunnel lighting and uniform surface texture of the tunnel lining structure, and addressing the problem of misalignment and deformation due to error accumulation in long-distance tunnel lining image stitching, a two-stage "local-global" image stitching method based on image feature learning and matching is proposed to construct a complete panoramic spatial image of the lining surface. Based on a stitching model using cylindrical projection, an image registration method based on spatial prior constraints is constructed. Anti-distortion remapping and global geometric alignment strategies are established. Based on the Laplace pyramid image fusion algorithm, stitching seams are eliminated, achieving visual continuity and naturalness in the global stitched image.

[0019] This paper proposes an anomaly detection method based on spatial locality constraints and multi-scale fusion. A spatial locality-enhanced memory model is constructed, and a spatial constraint training and testing strategy addresses the performance limitations of conventional methods under conditions of low sensitivity to local features and multi-angle interference. A cross-level local neighborhood aggregation, feature dimensionality reduction, and optimization method is established, achieving pixel-level accurate localization of tunnel lining defect anomaly regions in small-sample scenarios. This overcomes the limitation of limited tunnel lining defect sample data, effectively improving the initial detection accuracy and the ability to locate real defect anomalies in complex lining image backgrounds.

[0020] This invention proposes a two-stage local-global image stitching method, establishes a stitching model based on cylindrical surface projection, and constructs an image registration method based on spatial prior constraints, effectively reducing the accumulation and propagation of errors in image stitching. The image registration method based on LightGlue-DISK feature matching effectively improves the stability of image stitching on weakly textured and repetitive structural surfaces. An anti-distortion remapping and global geometric alignment strategy is established, and based on the Laplace image fusion algorithm, stitching seams are eliminated, achieving visual coherence and naturalness in the globally stitched image.

[0021] This invention proposes an image defect and anomaly detection method based on spatial local correlation. By constructing a spatial locality-enhanced memory model and combining it with a spatially constrained training and testing strategy, the performance bottleneck of conventional methods under low sensitivity to local features and multi-angle interference is overcome. Based on the ResNet-50 model, an innovative multi-level cross-scale fusion anomaly feature extraction method is proposed. Through cross-level local neighborhood aggregation, feature dimensionality reduction, and optimization, pixel-level localization accuracy of defect and anomaly regions is achieved.

[0022] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of the tunnel lining crack identification method based on panoramic spatial image stitching in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram illustrating the two-stage image stitching principle in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the splicing model of the cylindrical surface projection in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the image array matching search path in Embodiment 1 of the present invention. Detailed Implementation

[0025] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0026] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0027] Example 1: This invention provides a method for identifying tunnel lining cracks based on panoramic spatial image stitching. High-quality tunnel lining images are crucial data for locating abnormal defects. With the rise of machine vision technology, image detection based on industrial cameras has become the mainstream approach, enabling rapid acquisition and storage of tunnel lining images within effective window periods. Image stitching provides key image support for identifying cracks and other defects and for long-term monitoring. However, due to uneven lighting, interference from structural appendages, and other factors, feature point registration failures during large-scale image stitching lead to distortion, posing significant challenges to the efficiency and accuracy of the stitching process. To address this issue, this embodiment proposes an improved stitching method based on a segmented processing strategy and deep learning. Through steps such as local spatial prior constraints and anti-distortion remapping global optimization, it enhances the panoramic feature matching capability of complex images with weak textures, achieving visual coherence and naturalness in the global stitched image, and providing a high-quality image foundation for accurate location of abnormal areas in panoramic tunnel images.

[0028] First, the large-scale dataset is partitioned as a whole, and a sliding window is used for local stitching, processing only a few consecutive frames of images from different camera angles at a time to reduce error propagation. Second, image remapping is performed based on the tunnel structure characteristics to ensure that the stitched image has no obvious distortion. After geometric alignment, the transition part of the stitched region is processed by the Laplacian fusion method to ensure a natural transition in image color and brightness between overlapping areas, eliminating obvious stitching traces and obtaining a clean and intuitive stitched result. Finally, a feature detection and matching method based on the fusion of DISK and Lightglue is used to improve the feature matching capability of weak texture images.

[0029] Compared to conventional stitching methods, the method in this embodiment is suitable for image sequences of tunnel lining structures and can better adapt to their structural characteristics and imaging conditions. If the registration accuracy is insufficient, this method can easily cause blurring of high-frequency details and introduce errors. Therefore, this embodiment employs a multi-band mixing algorithm, adaptively determining the mixing weight for each image. Low-frequency details are smoothly fused over a larger spatial range, while high-frequency details are finely fused over a smaller range, thereby maintaining detail clarity and effectively suppressing stitching errors.

[0030] like Figure 1 and Figure 2 As shown, the specific steps include: Step 1: Obtain the tunnel lining image array, and use the stitching model based on cylindrical surface projection to locally stitch the tunnel lining images to obtain a locally stitched image.

[0031] Step 1.1: Obtain the tunnel lining image array.

[0032] In one specific implementation, local stitching involves precisely stitching together images within the same sliding window, laying the foundation for subsequent global stitching. Assuming the acquisition system contains N cameras, each capturing M images, this forms an N×M image array. Known spatial relationships between images in the array are used to improve the reliability of image registration. During local stitching, each sliding window contains T consecutive images. Therefore, each local stitching operation requires processing N×T input images and outputting the stitched result. If T is large, local stitching may fail due to excessive distortion. Therefore, to improve the feasibility of the output image, this embodiment uniformly selects T=3, performing M-2 local stitching operations on the entire image sequence.

[0033] Step 1.2: Use a stitching model based on cylindrical projection to perform local stitching of the tunnel lining image.

[0034] Step 1.2.1: Construct a projection stitching model based on a cylindrical surface.

[0035] The first step in panoramic image stitching is to calculate the registration relationship between images. Then, the registered pixels are projected onto a common stitching surface. Finally, image fusion technology is used to eliminate seams, thereby generating a complete panoramic image. The common stitching surface can typically be a plane, sphere, or cylinder. Since the tunnel lining surface can often be locally approximated as a cylinder, to ensure the realism and continuity of the stitching effect, this embodiment uses a projection stitching model based on a cylindrical surface, such as... Figure 3 As shown.

[0036] Assuming the camera rotates around its optical center, the geometric relationships between images can be represented as a special type of homography transformation. This can be achieved by using a rotation vector. and focal length By parameterizing each camera, the correspondence between image points can be obtained: The homography matrix is ​​defined as follows: (1).

[0037] in, It is a homography matrix. and These are the homogeneous coordinates of the corresponding points on the image. ,in Represents a two-dimensional image label. and These are the intrinsic parameter matrices of the camera. and These are the camera's rotation matrices. The camera's intrinsic parameter matrix contains parameters such as the camera's focal length. If we assume the image center is the origin, it can be simplified to: (2).

[0038] in, Let be the camera focal length. The rotation matrix is ​​expressed in exponential mapping form as: (3).

[0039] In theory, image matching methods aim to utilize features that are invariant to geometric transformations. If there are small positional differences between images, they can be approximated using Taylor expansion: (4).

[0040] Its equivalent form can be expressed as: (5).

[0041] in, Through in The affine transformation obtained by linearizing homography means that each small image patch undergoes an affine transformation. This represents the pixel coordinates at the reference position or unfolding point selected during Taylor expansion. For images where there is no significant parallax due to camera rotation, affine transformation is used. Coarse alignment is performed to handle minor changes in image position, thereby preserving the basic geometric properties of the image. This is achieved through a homography matrix. Images containing camera movement and tilt Transform to image In the coordinate system, pixel-level alignment between images is achieved, and then the overlapping areas after mapping are processed through methods such as beam adjustment and fusion.

[0042] According to formula (1), if the homography transformation between two images is known... and intrinsic parameter matrix and Then the rotation transformation between the two images can be obtained. Assuming the optical centers of the cameras are at the same point, meaning the translation between the cameras is 0, therefore... That is, an image To image Three-dimensional spatial transformations between them. If the first frame is selected ( If a world coordinate system is defined for a reference frame, then the first coordinate system can be defined. The transformation from the frame image coordinate system to the world coordinate system is as follows: Known , Then, using formulas (6) to (8), we can calculate the mapping of a pixel coordinate (x, y) on the image to the corresponding coordinate (0, y) on the stitching cylinder. ): (6), (7), (8).

[0043] Where Pw represents the three-dimensional coordinates of a point in the world coordinate system, denoted as Pw = (Xw, Yw, Zw). and These are the cylindrical coordinates (longitude and latitude) on a cylinder with a radius of 1. After appropriate scaling and translation, they can be converted into pixel coordinates on the panoramic image. Since the depth of each pixel is unknown, therefore... Transformation, projecting pixel coordinates onto the image In the coordinate system On the plane, and then through Transform to the world coordinate system.

[0044] Step 1.2.2: Register the tunnel lining image using spatial prior constraints to obtain the registered image.

[0045] Based on the above stitching model, the key to stitching lies in estimating the parameters corresponding to each image. For concentric camera arrays, this parameter can be obtained through homography transformation between images. The basic process for calculating the homography transformation between images is as follows: first, establish the sparse point correspondence between images through feature point matching, and then estimate the transformation matrix based on the correspondence. Traditional feature point matching methods often used have poor matching stability in weakly textured regions. Therefore, this embodiment proposes a feature matching method for weakly textured surfaces, which can significantly improve the stability of matching weakly textured tunnel surfaces. It is tentatively assumed that a feature point matching algorithm already exists, capable of matching any image pair (…). , Output a set of feature point correspondences based on spatial prior constraints.

[0046] (9).

[0047] Due to the instability of local feature matching, the original feature point correspondence set Homography often contains a large number of erroneous matching point pairs. This embodiment uses RANSAC for robust homography estimation and feature selection. The algorithm is implemented through cyclic random sampling: in each iteration, four sets of corresponding points are randomly selected, and the initial homography matrix is ​​calculated using the Direct Linear Transform (DLT) method. The model is evaluated based on the number of interior points, and the result with the most interior points is ultimately adopted as the optimal estimate.

[0048] The RANSAC algorithm robustly handles a high proportion of mismatches, achieving reliable estimation of the homography matrix. Simultaneously, this algorithm effectively distinguishes between correct matches (inliers) and incorrect matches (outliers), providing direct evidence for evaluating the reliability of image registration results. If the set of correctly matched corresponding points is the image, then... The reliability score of registration is defined as: (10).

[0049] in, For reliability score, The cardinality of the set is denoted by C, which is a constant used to represent the cardinality of the set. and A penalty is applied when the values ​​are relatively small. In this embodiment, C=20 is set.

[0050] If the input images are unordered, then pairwise matching of the images is required, followed by threshold filtering. This method aims to determine whether a matching relationship exists between images. However, this approach has two problems: First, pairwise matching itself is computationally intensive, and due to the lack of spatial constraints, it is prone to false matches in regions with repetitive structures or textures; second, Setting a fixed threshold is difficult, and using a fixed threshold often fails to yield stable matching results. To address these issues, this embodiment proposes a matching method based on spatial prior constraints and a dynamic adaptive threshold to effectively improve the accuracy and stability of the matching.

[0051] The input images in this embodiment are arranged in an array, such as... Figure 4 As shown. Based on this arrangement, the matching relationship between images should satisfy the inherent spatial constraints of the array. If the overlap area between adjacent images is small, theoretically each image will only have a matching relationship with its directly adjacent images. To improve the generalization ability of this method, this embodiment further performs a wider range of matching searches in the horizontal, vertical, and two diagonal directions to fully utilize the overlapping areas of the images and improve the stitching effect. (The image needs to be matched with the center image.) The objects being matched can be considered as those from The search proceeds sequentially along eight directions. Simultaneously, considering that images farther from the center should have lower matching scores, the search process terminates if this trend is violated. This strategy effectively suppresses mismatches caused by repetitive structures and textures on the tunnel surface. Furthermore, because... Since adjacent images are ensured to have overlapping regions during acquisition, they are always considered to have a matching relationship. Based on this, the matching threshold in a certain direction can be set as follows: (11).

[0052] in, Is with The matching score between adjacent images and its matching score. It is a fixed threshold that is set. The above method can effectively improve the efficiency and stability of image matching.

[0053] Step 1.2.3: Perform consistency optimization on camera parameters.

[0054] Preliminary camera parameters can be obtained through pairwise image registration, but each registration only considers two images, thus failing to guarantee the overall consistency of camera parameters. To address this, this embodiment introduces bundle adjustment to globally optimize all camera parameters after obtaining geometrically consistent matching points between images. Compared to traditional stitching methods that rely on paired homography matrices, bundle adjustment effectively avoids error propagation and accumulation while strengthening global constraints. During optimization, images are added incrementally based on their matching degree: each time, the image with the strongest match to the current set is selected, and its initial rotation parameters and focal length are set to the estimated values ​​of the best-matching image, thus providing a robust initial point for optimization and ensuring convergence efficiency. Finally, the global parameters are iteratively updated using the Levenberg-Marquardt (LM) algorithm, resulting in more accurate and stable camera pose and intrinsic parameter estimates.

[0055] The optimization objective of bundle adjustment is to minimize the reprojection error, which is the difference between the position of a feature point projected onto each image based on camera parameters and its actual observed matching position. Specifically, for each feature point, it is projected onto all images in which it was observed, and the squared Euclidean distance between the projected point and the corresponding matching point in each image is calculated. The final objective function is the sum of all these squared distances, which is minimized by adjusting the camera parameters. Given a correspondence... → The residual is (12).

[0056] in, Represents the residual. Representing an image No. The location of each feature It is an image The middle corresponds to Points in the image The projection onto the surface, according to the projection stitching model, is calculated using the following formula: (13).

[0057] in, This represents the projection coordinates of the k-th feature point in image j onto image i after the camera's intrinsic and extrinsic parameters have been transformed (i.e., the transformed pixel coordinates mapped from image j to image i).

[0058] The overall error function is the sum of the residuals of all feature points in the image: (14).

[0059] in, Indicates the number of images. Representation and Image Matched image matching set, Representing an image and images The feature matching set between them. It is a robust function used to suppress the effects of errors: (15).

[0060] The above formula combines the fast convergence of the L2 norm optimization scheme for interior points with the robustness of the L1 norm for exterior points. Outlier distance is used during initialization. In the final solution, .

[0061] The objective function of formula (14) is optimized to obtain a nonlinear least squares problem. In this embodiment, the Levenberg-Marquardt algorithm is used, and the form of each iteration step is as follows: (16).

[0062] Where φ is a column vector containing all the parameters to be solved. It is the parameter increment obtained in each iteration. For residuals, .matrix Includes prior knowledge of regularization weights with different parameters: (17).

[0063] Wherein, the standard deviation of the angle is The standard deviation of focal length is , Set a uniform step size parameter for the average focal length. This accelerates the convergence speed and improves stability.

[0064] The solution process uses the parameters obtained from image registration as initial values, and calculates the parameter increment in each iteration according to formula (16) until the parameter change is less than the preset threshold and the iteration is terminated.

[0065] Step 1.2.4: Merge and stitch the registered images together.

[0066] In the image fusion stage, this embodiment adopts the same method as OpenCV Stitcher. First, the geometrically registered images are mapped to the panoramic coordinate system through a distortion transformation; then, exposure compensation is performed to eliminate overall brightness differences between images; next, a graph cutting algorithm is used to optimize the seam positions of overlapping areas; finally, a seamless stitching result is generated through fusion processing. Exposure compensation is achieved by assigning a global gain coefficient to each image and scaling its pixel values; its objective function is... It can be defined as: (18).

[0067] in, and Images and The gain parameter, i.e., the scaling factor of the pixel value. This represents the overlapping region between two images. The objective function aims to minimize the difference in corresponding pixel values ​​within the overlapping region, simplifying computation and enhancing the robustness of the stitching system to changes in illumination. Finally, the optimal gain parameter is obtained by solving the objective function using the least squares method.

[0068] Step 2: Use the global image space model to fuse and stitch the local stitched images to obtain a panoramic stitched data map.

[0069] In one specific implementation, when the number of images is large or the stitching range is wide, local stitching can easily produce significant image distortion. In actual tunnel lining appearance inspection tasks, thousands to tens of thousands of images are typically required. If a local stitching method is directly used, it is obviously difficult to meet the accuracy and stability requirements of practical applications. Therefore, this embodiment proposes a global stitching method based on local stitching: by optimizing the global consistency of the local results obtained from segmented stitching and by using anti-distortion processing, seamless stitching of tunnels of arbitrary length can be achieved, ultimately generating a visually straight and geometrically consistent panoramic result.

[0070] Step 2.1: Perform anti-distortion remapping and global geometric alignment on the locally stitched image.

[0071] In the local stitching stage, a sliding window strategy is used to perform M-2 local stitching operations of size N×3 sequentially, resulting in a local stitching result covering all adjacent images. The local stitching algorithm uses the local stitching result as input and generates a complete and globally consistent panoramic image by remapping and fusing the stitching rows of each local panoramic image. However, the local stitching process inevitably introduces image distortion. If global stitching is performed directly on this basis, distortion errors will accumulate continuously, especially in long image sequences, potentially leading to final stitching failure. To address this issue, this embodiment proposes an anti-distortion remapping method tailored to the characteristics of tunnel scenes. This method can reverse-map distorted local panoramic images into distortion-free, fixed-width regular image blocks, thereby effectively suppressing error propagation and achieving robust global stitching of image sequences of arbitrary length.

[0072] Each stitched panoramic image can be divided into three regions: (1) Region 1 contained in the first frame image.

[0073] (2) Region 2 is contained in the middle frame image, but the image in Region 2 does not belong to Region 1.

[0074] (3) Only included in the image region three of the last frame.

[0075] In this context, any valid pixel in the local panoramic image must belong to one of the aforementioned regions. Let all local stitching results be... For any Its region one also appeared In the middle, Region 3 also appeared In the panorama, region two can be considered as a newly added part due to the introduction of the t-th frame image. Therefore, global stitching should include... Region 2 and and The corresponding region 2 in the image is spliced ​​together, and the anti-distortion remapping process is also mainly carried out in this region.

[0076] The boundaries of regions one through three can be determined based on the mapping relationship between the original image and the local panoramic image. These boundaries serve as the stitching lines between different locally stitched images in global stitching. Provided the registration is accurate, The upper suture and The lower stitching lines all correspond to the last row of pixels in the t-th frame of the original image. Therefore, the upper and lower stitching lines between adjacent anti-distortion image blocks are geometrically aligned, facilitating global geometric consistency. In addition, besides the main area between the stitching lines, each frame image retains an extra area of ​​a certain width at both the top and bottom. This part will be used to fuse the locally stitched images using an image fusion algorithm based on the Laplace pyramid.

[0077] The anti-distortion remapping process resamples each column in the anti-distortion block. The width of the anti-distortion block is a fixed value, consistent with the width of the panoramic image. Uniform sampling is performed along the upper and lower stitching lines at this width to obtain sample point sets. and Linear interpolation is used during sampling to ensure uniform distribution of sampling points. The height of the anti-twist block can be estimated by calculating the average distance between the upper and lower sampling points. As shown in Equation (19), it provides geometric constraints for subsequent image remapping.

[0078] (19).

[0079] in, This indicates the number of sample point pairs on the two concatenated rows. and, These are a pair of upsampling points and downsampling points, This represents the Euclidean distance. The height estimate of the stitched row can be calculated based on the Euclidean distances of all corresponding point pairs, and then the image region can be remapped. Based on the obtained sampling point information, a smooth mapping relationship is constructed using linear interpolation, thereby achieving accurate adjustment of the image geometry, as shown in Equation 20: (20).

[0080] in, and The coordinates in the target image and the mapping matrix are respectively. The size is This ensures that the stitched blocks have sufficient transition space for image fusion, where To splice the row height, To splice the line width, To determine the width of the merged region. Finally, through... of The function remaps the local stitching result to the image space of the anti-distorted block.

[0081] Step 2.2: Use the Laplace pyramid-based image fusion algorithm to fuse the local stitched images to obtain a preliminary global stitched image.

[0082] The global stitching process starts from Starting with the anti-twist block, sequentially... Subsequent image patches are then stitched sequentially into the panoramic image. To eliminate visible seams that may arise near the stitching lines due to differences in light intensity and registration errors, this embodiment employs the Laplacian pyramid fusion method to fuse the final stitching result. This method is a classic algorithm in the field of image fusion, and its core idea is based on the multi-resolution representation and reconstructable properties of the Laplacian pyramid: first, a Laplacian pyramid is constructed for each image to be fused; then, stitching and fusion are performed at each pyramid level; finally, a seamless output image is generated through pyramid reconstruction.

[0083] Anti-distortion image patches between two adjacent frames Assuming They are respectively The first of the Laplace Pyramid If the layers are combined, the resulting Laplace pyramid can be represented as: (twenty one).

[0084] in, The fused Pyramid of Laplace, This is a single-layer fusion operation, where the images are directly stitched together based on the image registration results. The resulting Laplacian pyramid is shown below. The final fused image is obtained through layer-by-layer reconstruction. The specific process can be viewed as the reverse of the construction process of the Laplace pyramid: (twenty two).

[0085] in, This is the reverse process of the construction of the Pyramid of Laplace. This indicates an upsampling operation, resulting in a final fused image that is the bottom layer of the Gaussian pyramid. The Laplacian pyramid fusion method smooths the transition of overlapping areas during line-by-line stitching, achieving a natural appearance. This method significantly reduces the impact of lighting differences, better preserves the details of the original image, and significantly reduces stitching artifacts. The image after Laplacian pyramid fusion exhibits natural light and shadow transitions in overlapping areas, coherent structure, and a good overall visual effect with no obvious stitching artifacts.

[0086] Step 2.3: Optimize the initial global stitched image using an image matching method for weakly textured surfaces.

[0087] Feature points are a key factor in achieving high-quality image stitching, and their stability and reliability directly affect the stitching effect. However, due to factors such as insufficient lighting, shadow interference, and reflection, tunnel stitched images often contain large areas of weak texture, which can easily lead to unstable feature matching and thus stitching failure. Due to feature matching errors, obvious defects such as distortion, ghosting, and misalignment appear in panoramic images. To address this, this embodiment proposes a weak texture matching method based on LightGlue, aiming to improve the robustness of panoramic stitching in weak texture scenes such as tunnel surfaces. LightGlue is a feature matching method based on the Transformer architecture, which uses an attention mechanism to model the global dependencies between feature points. Compared to traditional matching methods, LightGlue performs better under challenging conditions such as complex backgrounds, weak texture regions, and large angular transformations, making it particularly suitable for feature matching tasks in complex imaging environments such as tunnel surfaces.

[0088] Step 2.3.1: Use the feature matching method based on LightGlue weak texture to extract features from the preliminary global stitched image.

[0089] Unlike traditional image matching methods, LightGlue enhances feature representation and performs cross-image feature aggregation by using a self-attention mechanism when processing features of two images simultaneously. It is also supplemented by overall matching optimization based on spatial constraints, thereby significantly improving matching accuracy in challenging scenarios such as weak textures and repetitive structures.

[0090] For the input image and Assuming that its local features have been pre-extracted. Each feature point From a two-dimensional coordinate system normalized to the image size and a visual descriptor Composition. Image and Each contains *and * feature points, whose index sets are respectively and .

[0091] LightGlue's goal is to infer a set of matching relationships. Each pair of matches corresponds to the same point in three-dimensional space, where, and All represent feature point indices. It's important to note that not all feature points can be successfully matched; some points may fail to establish a reliable correspondence due to occlusion or low feature repetition. To achieve this, LightGlue first computes a soft assignment matrix between the two sets of features. Then, by combining spatial consistency constraints, the final matching pairs are extracted. This model is composed of... It consists of a hierarchical structure with the same composition, and jointly processes the feature set. and Each layer contains self-attention units and cross-attention units: self-attention units enhance the contextual representation between features within the same image; cross-attention units capture the relationships between features from two images. Furthermore, each layer includes a classifier to determine if a sufficient match has been achieved, allowing for early termination of inference and saving computational resources. Finally, a lightweight prediction head decodes the updated feature representations, outputting a partial match result between the two feature sets.

[0092] The core idea of ​​the LightGlue framework lies in leveraging context-rich visual descriptors to improve matching accuracy. Its key advantage is its adaptive computational capability: if the input image pairs are simple, with high overlap and small appearance variations, reliable descriptors allow the network to generate high-confidence matches consistent with deeper layer results at shallow layers, and the system terminates inference early. This mechanism significantly improves computational efficiency while maintaining accuracy. At each layer... In the middle, if the confidence score of a certain point Exceeding the current layer threshold If a point is deemed trustworthy, it is considered reliable. If a sufficiently high proportion (more than α) of points are deemed trustworthy, the algorithm terminates the inference process. This mechanism allows LightGlue to adaptively adjust the computation depth based on the matching difficulty of image pairs, significantly reducing computational overhead while maintaining accuracy.

[0093] (twenty three).

[0094] in, This is a Boolean condition indicating that LightGlue is currently in the [current position]. A quantifier for determining whether a layer triggers an early exit / early termination of inference.

[0095] In LightGlue, classifiers often lack confidence in early layers. Therefore, a decay threshold λ is introduced for each layer based on the validation accuracy of the classifiers at each layer. The exit threshold α allows for a flexible balance between matching accuracy and inference efficiency. If the current layer does not meet the exit condition, feature points that have been predicted as reliable but not matched (usually located in occluded or invisible areas) often fail to provide effective matching information in subsequent layers. Based on this observation, the algorithm actively removes such points at each layer, only passing the remaining features to the next layer. Since the computational complexity of the attention mechanism increases quadratically with the number of tokens, this strategy significantly reduces computational overhead while having almost no impact on overall matching accuracy. LightGlue demonstrates significant advantages in training speed, convergence stability, and matching accuracy. Compared to SuperGlue, LightGlue significantly reduces the computational resources and costs required for model training, making deep learning-based matching methods more readily adopted. During the training phase, LightGlue is compatible with traditional local features such as SuperPoint and SIFT, and employs specific data segmentation strategies during fine-tuning to avoid overfitting in certain scenarios. Addressing the common issues of weak and repetitive textures in the appearance images of tunnel lining structures, this embodiment introduces LightGlue as the basis for feature matching, combined with the deep learning-based feature detection method DISK. DISK can generate enhanced artificial features, while LightGlue combines deep representation learning with geometric optimization reasoning capabilities. The combination of the two demonstrates good robustness and stitching effect in tunnel image matching.

[0096] Step 2.3.2: Use the DISK and LightGlue fusion method to perform feature detection and matching on the extracted features.

[0097] Extracting and matching feature points from all images is the primary task of image stitching algorithms. At each feature location, its scale and orientation need to be determined. Using invariant feature descriptors ensures reliable matching of image sequences even when images undergo rotation, scaling, and illumination changes. By modeling image stitching as a multi-image matching problem, matching relationships between images can be automatically discovered, even identifying connection structures in unordered image sets. This embodiment uses the DISK feature detection method for initial feature extraction. This method, based on the U-Net structure, can automatically learn local structural information of images and capture detailed features at different network layers, thus extracting stable and discriminative feature points even in low-texture scenes. The LightGlue feature matching method, optimized for weak textures, is used for matching. By modeling the global dependencies between feature points through a self-attention mechanism, the matching accuracy is significantly improved.

[0098] Based on the fundamental idea of ​​LightGlue, the algorithm can predict and obtain partial matching relationships between local feature sets extracted from two images. This embodiment sets the maximum number of keypoints extracted from each image to 2048 and integrates this extractor into the matcher. In the attention unit, each image... Information is aggregated from all feature points, and feature representations within the same image and between different images are fused using self-attention and cross-attention units, thereby simultaneously capturing relative and absolute positional relationships. Subsequently, a matching score is calculated for each feature point, combining appearance similarity and geometric consistency into a single assignment matrix. Ground value matching is defined based on reprojection error and depth consistency: if a pair of feature points has low reprojection error and consistent depth in two images, it is considered a positive sample; if a point has a significantly larger reprojection error or depth difference than all other points, it is marked as a mismatch. When two points are predicted to be a match and their similarity is higher than any other candidate point in the two images, the algorithm matches the point pair. The match was determined to be valid.

[0099] This embodiment innovatively proposes a two-stage "local-global" image stitching method. Based on the stitching model of cylindrical projection, it constructs an image registration method based on spatial prior constraints, which greatly reduces the accumulation and propagation of errors in local image stitching. It establishes an anti-distortion remapping and global geometric alignment strategy, and based on the Laplace pyramid image fusion algorithm, it eliminates stitching seams and achieves visual continuity and naturalness of the global stitched image.

[0100] Step 3: Based on the spatial local differences, perform anomaly detection on the panoramic stitched data map to obtain the tunnel lining crack identification results.

[0101] In one specific implementation, the lining panoramic spatial model stitching method based on two-stage image fusion proposed in this embodiment can present a high-quality panoramic image of the tunnel lining structure detection area, providing a high-quality image foundation for anomaly detection in tunnel lining structures. However, with the increase in tunnel length, the scale of image data increases dramatically, and problems such as the dark and dusty environment of tunnels and the complex surface texture of tunnel lining structures make it difficult for manual and conventional intelligent detection methods to effectively characterize the morphological diversity of defects. Conventional anomaly detection methods use global feature distribution of images for comparison, which requires massive learning data support, and faces problems such as misjudgment due to local detail interference and large global processing workload when identifying large-scale anomaly defects on the surface of tunnel linings. In response, this embodiment addresses the challenges of large data scale, complex changes, and poor multi-scale defect recognition in tunnel surface anomaly detection by proposing an image anomaly feature extraction method based on multi-scale fusion of spatial local differences. This method effectively reduces the search range for subsequent crack annotation and accurate detection segmentation at a relatively low cost, achieving the goal of accurate detection of image anomaly areas in small sample scenarios.

[0102] Anomaly detection was performed on a panoramic spatial mosaic image of the tunnel lining structure. The input image was processed in blocks, and each block was analyzed... Extract the corresponding image patch and The defect region is predicted by local spatial constraints, and then the initial mask of the defect region is obtained. If the anomaly detection result is normal, then the subsequent processing will be terminated.

[0103] Specifically, the following steps are included: Step 3.1: Use an image anomaly and defect detection method based on spatial local correlation to perform image representation on the panoramic stitched data map.

[0104] Step 3.1.1: Construct an anomaly detection model based on the PatchCore anomaly detection algorithm.

[0105] PatchCore is an anomaly detection method based on image representation. Its core idea is to use a pre-trained network to extract features, construct a comprehensive and compact feature memory for normal samples, and then detect anomalies by comparing test samples with this memory. This method makes significant innovations in feature selection and feature library construction, achieving significant improvements in both detection accuracy and computational efficiency.

[0106] PatchCore employs a unique intermediate layer feature selection strategy. Unlike SPADE and PaDiM, which directly utilize deep high-level semantic features from pre-trained networks, this method focuses on extracting feature representations from the network's intermediate layers. While high-level features pre-trained on natural image datasets like ImageNet contain rich semantic information, their semantic concepts often deviate from the actual needs of anomaly detection tasks, easily introducing irrelevant semantic interference. In contrast, the network's intermediate layer features achieve a better balance between "spatial detail" and "abstract semantics," preserving sufficient spatial detail to support accurate defect localization while also possessing appropriate semantic abstraction capabilities, making it more suitable for image anomaly detection tasks in engineering scenarios.

[0107] For each input image, PatchCore extracts feature maps from the two middle layers of a pre-trained CNN (such as WideResNet-101), adjusts them to a uniform spatial resolution (typically 28×28) using bilinear interpolation, then concatenates them along the channel dimension, and flattens the feature vectors at each spatial location to generate a set of feature descriptors that represent local image information. The construction of the feature library is the most innovative part of PatchCore, but simply aggregating the feature vectors of all training images would result in an extremely large memory bank, consuming a lot of storage space and making subsequent nearest neighbor searches very time-consuming, which is difficult to meet the needs of engineering applications. To address this, PatchCore introduces a "core set" subsampling technique, selecting a smaller subset from the original feature set to maximize the preservation of the original feature space's distribution characteristics. During algorithm initialization, a feature point is randomly selected as the starting point of the core set. Subsequently, in each iteration, the feature point currently furthest from the core set is added to it, repeating this process until the core set reaches a predetermined size. This farthest-point sampling strategy ensures that the core set uniformly covers the entire feature space, avoiding the omission of any clusters in multimodal distributions and achieving good spatial coverage in uniform distributions. Using this method, PatchCore significantly compresses the feature library size with almost no loss in detection performance. For each test image, PatchCore first extracts its mid-level feature descriptors using the same method as in the training phase. Then, for each local feature descriptor, it searches for its nearest neighbor in the compressed core set memory, using the distance between them as the anomaly score for that location. The anomaly scores of all local locations are further reconstructed into a spatial map, which is then Gaussian smoothed to generate the final anomaly heatmap, clearly and intuitively displaying the degree of anomaly in each region of the image. Simultaneously, the system takes the maximum value among all local anomaly scores as the image-level anomaly score for overall anomaly determination.

[0108] Step 3.1.2: Construct a spatial locality enhancement memory and use the data in the memory to train the anomaly detection model.

[0109] To address the challenges of low energy efficiency, sensitivity to local features, and multi-angle interference in anomaly detection against complex texture backgrounds, this embodiment innovatively incorporates training data constraints and spatial coordinate encoding based on PatchCore. Through dual-dimensional optimization, it improves the model's detection accuracy and environmental robustness in engineering scenarios.

[0110] During the training phase, the model uses only normal sample images I∈R from the training set. H×W×3 In this dataset, all input images are considered normal, with a label value of 1 (in this dataset, 0 represents an abnormal image and 1 represents a normal image; this label is used for subsequent comparison with the model's prediction results; the model's output anomaly score is a continuous value between 0 and 1, representing the probability that the image is abnormal). This embodiment uses ResNet-50 as the backbone network for image feature extraction. This network has been pre-trained on ImageNet and can extract multi-level deep features, focusing on the features output by the level 2 and level 3 residual layers. Through a sliding window operation, the feature representations of all local image patches are further extracted and aggregated to construct a standard memory library representing the features of normal images, thereby constructing the feature representation space of normal samples.

[0111] While using only level 3 single-layer features can achieve basic feature extraction, it's difficult to fully preserve detailed information. In test environments where the anomaly type is unknown, especially in anomaly detection models trained only with normal features, this approach can easily lead to highly unstable model performance. Therefore, this model employs a multi-scale feature fusion strategy, preserving both fine structures and local features while leveraging intermediate layer features from ResNet-50 to construct a general representation suitable for image anomaly detection. This representation has a moderate receptive field size (5×5 pixels), capable of capturing semantic context information while avoiding loss of localization ability due to excessive abstraction. Furthermore, by embedding coordinate information into image annotations, a location code P∈R is generated for each local region. H×W×2 Here, H and W represent the height and width of the image, respectively, and each image patch contains normalized coordinate values ​​(x / H, y / W). To prevent the coordinate information from being overwhelmed by the main features, a nonlinear transformation is required using learnable parameters: (twenty four).

[0112] in, W represents the coordinates after nonlinear transformation. p ∈R 2×2 Let b be the weight matrix. p ∈R2 The bias vector ensures that the coordinate encoding values ​​are within the range [0,1] and aligned with the image feature dimensions. The multi-layer image features F extracted from the pre-trained network are used as the bias vector. img After being concatenated with the coordinate encoding P and fused through a convolutional layer, the spatial augmentation features F required by the model can be output. fused .

[0113] Step 3.1.3: Set up a spatial constraint training and testing strategy, and use the spatial constraint training and testing strategy to further optimize the anomaly detection model.

[0114] The dataset is divided into several groups according to the shooting angle, which serve as the input to the backbone network. Each group contains only normal and abnormal samples from the same angle. The feature f of the test image is... test (x,y) is only related to the standard feature blocks contained within a certain range of the region (x,y) in the memory bank. Perform a similarity comparison. This represents the set of all features to be compared, determined by the spatial location (x, y) of the test feature and the search range parameter d. A smaller d results in stronger spatial locality constraints, which not only reduces the size of the training and testing datasets and effectively avoids interference from image features at different angles, but also reduces interference from certain textured local regions on the detection of small cracks and defects. On the other hand, a smaller d requires higher spatial alignment accuracy between the test and training images; otherwise, alignment errors may lead to detection mistakes. A separate configuration is constructed for each (x, y). This requires significant overhead; therefore, for large-scale panoramic images of tunnel linings, the surface area can be pre-indexed into blocks, allowing pixels within the same block to share the same set of standard features. Pre-calculation can significantly improve the efficiency of search matching. Given a standard feature set, the anomaly score of a test sample, i.e., the degree of deviation from the distribution of normal features, can be calculated as follows: (25).

[0115] in, outlier score of the test sample An initial anomaly score map is generated by calculating the anomaly score at each spatial location (x, y). This score map is then upsampled to the resolution of the input image using bilinear interpolation to restore its original size. Subsequently, a Gaussian smoothing kernel (σ=4) is applied to suppress edge artifacts and noise, ultimately generating a visualized anomaly heatmap. Traditional visualization methods typically normalize the score range of a single image, stretching the display of differences based on the score interval of the entire image. While this method enhances the contrast between anomalous and normal regions, it also has a significant drawback: even when an image is determined to be normal and without any anomalous regions, it still forcibly amplifies score differences, leading to the display of false anomaly areas. To address this, this model employs a normalization mapping strategy based on global maximum and minimum values ​​to improve the interpretability of the heatmap, thereby more reliably reflecting the actual distribution of anomalies.

[0116] Anomaly detection methods based on spatial local dissimilarity enhancement effectively alleviate the performance limitations of conventional methods in the face of low sensitivity to local features and multi-angle interference by using fixed-angle training and coordinate information embedding. This method significantly improves the detection accuracy and robustness of the model in complex scenes in image surface defect detection, especially in complex texture backgrounds, maintaining a high detection capability for subtle defect areas such as small cracks and micro-peeling.

[0117] Step 3.2: Extract abnormal features from the panoramic stitched data map using a multi-scale fusion strategy.

[0118] Conventional anomaly detection methods generally suffer from significant problems of losing microscopic details and confusing macroscopic semantics when dealing with multi-scale defects. On the one hand, feature extraction using only deep or shallow convolutional features at a single scale is insufficient to address both small defects and large-area anomalies. On the other hand, existing multi-scale fusion methods lack dynamic weight adjustment mechanisms, leading to shallow noise interfering with deep semantic features, especially in complex engineering scenarios where the false alarm rate increases significantly, seriously affecting the reliability of the detection results. In response, this embodiment proposes a dynamic multi-scale feature fusion framework to address the following two core issues: (1) achieving cross-level feature adaptation to reduce the spatial information misalignment caused by differences in the output resolution of different convolutional layers; (2) suppressing feature redundancy to avoid directly fusing unfiltered multi-layer features and introducing irrelevant information or noise.

[0119] Step 3.2.1: Use ResNet-50 to extract multi-level features from the panoramic stitched data.

[0120] This embodiment uses a pre-trained ResNet-50 backbone network to extract features from the detection image. By registering forward hooks, it captures the features of the specified intermediate layers, namely level2 (output size 56×56, 512 channels) and level3 (output size 28×28, 1024 channels). These two layers of features are then concatenated along the channel dimension to construct a composite feature representation that combines detailed and semantic information.

[0121] ResNet-50 consists of four layers, from level 1 to level 4, each containing multiple residual convolutional layers. With each layer, the spatial resolution of the output feature map decreases by half, but the number of channels in the feature map doubles, thus increasing the amount of local information contained in a single feature. Since level 1 is the shallowest layer, it mainly contains low-level image features and is susceptible to alignment errors, lighting variations, and noise, making it unsuitable for training models. Level 4 mainly contains high-level semantic information of the image, but its spatial resolution is low. Although it possesses strong semantic abstraction capabilities, after multiple rounds of convolution and pooling operations, spatial information is severely lost, making it unable to distinguish the contours and edges of defects. Therefore, levels 2 and 3, with their moderate resolution and level of abstraction, are used as the core feature extraction layers. Level 2 can preserve high-frequency details, while level 3 can output larger contour features. Level 2 feature maps have a high spatial resolution (56×56), which can capture relatively small surface defects such as cracks and scratches. Their receptive field covers an area of ​​about 5×5 pixels, making them suitable for detecting local texture anomalies. Level 3 feature maps extract regional semantic information such as wall peeling and local structural shifts through deep convolution. Their receptive field is expanded to 14×14 pixels, which can identify structural defects across pixels and obtain the contour features of abnormal areas.

[0122] Step 3.2.2: Aggregate the extracted multi-level features across local neighborhoods to obtain aggregated features.

[0123] Traditional CNN deep features are susceptible to the characteristics of ImageNet classification tasks, causing features to be biased towards detailed texture features while ignoring local structural information. To enhance the robustness of features, adaptive average pooling needs to be performed on each spatial location (x, y) to achieve contextual feature fusion. (26).

[0124] in, For context-fusion features, Np(x,y) is defined as a 3×3 neighborhood window centered at (x,y). This window size achieves an optimal balance between noise suppression and defect edge preservation. This operation, on the one hand, incorporates local contextual information into the feature representation without reducing resolution, thereby expanding the local receptive field to a 5×5 pixel range while maintaining the original feature map resolution (56×56) to avoid spatial information loss due to downsampling; on the other hand, it suppresses minor positional shifts caused by lighting or shooting angle, improving the model's adaptability to industrial scenarios.

[0125] When fusing features across different levels, a single level of features struggles to accommodate defects at different scales. While the high resolution of level 2 (56×56) is beneficial for locating minute defects, and the deep semantics of level 3 (28×28) excels at identifying regional anomalies, direct fusion of the two can lead to spatial mismatch due to differences in allocation rates. To address this, this embodiment proposes a progressive alignment strategy: using bilinear interpolation to increase the resolution of level 3 features from 28×28 to 56×56. After alignment with level 2, these features can be stitched and fused along the channel dimension. (27).

[0126] in, The feature transformation operation represents the upscaling from level 3 to level 3, and ⊕ represents feature concatenation, ultimately yielding the composite feature tensor F. fused ∈R 1536×56×56 It retains the high-resolution features of level 2 for locating pixel-level anomalies, and also retains the mid-level features of level 3 for identifying regional defect patterns.

[0127] Step 3.2.3: Perform dimensionality reduction and optimization on the aggregated features.

[0128] High-dimensional feature data (1536 dimensions) can lead to a huge computational burden, easily causing a surge in computational complexity. To reduce computational complexity, this embodiment designs a learnable projection matrix W. P ∈R 1536×256 The concatenated high-dimensional features are compressed to a uniform, moderate dimension, a process achieved through a fully connected layer: (28).

[0129] in, For the features after dimensionality reduction, The learnable projection matrix can effectively preserve key information about defects across scales while significantly reducing memory resource usage.

[0130] The proposed multi-scale feature fusion-based anomaly detection framework for tunnel lining structures exhibits significant advantages in three aspects: feature layer construction, local feature capture, and lightweight computation. These advantages can be summarized as follows: First, in feature layer construction, a ResNet-50-based dual-stream feature extraction architecture is constructed. Bilinear interpolation and channel concatenation achieve spatial alignment between level 2 and level 3 feature maps, eliminating errors caused by cross-layer spatial misalignment and providing geometric consistency for subsequent fusion. Second, in local feature capture, a local neighborhood adaptive pooling method is proposed. Through a dynamic receptive field enhancement mechanism, the receptive field of each feature point is expanded to a 5×5 pixel range while maintaining the original resolution, suppressing illumination sensitivity and improving the local signal-to-noise ratio of micro-crack detection. Third, in terms of lightweight computation, feature dimensionality reduction technology reduces computational complexity. Without losing core feature information, it significantly reduces the amount of feature data and model computational complexity, improving operational efficiency.

[0131] This embodiment innovatively proposes an image anomaly detection method based on panoramic spatial locality correlation, constructs a spatial locality enhancement memory bank model, and solves the performance limitations of conventional methods under low sensitivity to local features and multi-angle interference through spatial constraint training and testing strategies. It also establishes cross-level local neighborhood aggregation, feature dimensionality reduction and optimization methods to achieve pixel-level accurate localization of tunnel lining defect anomaly areas in small sample scenarios.

[0132] Example 2: Embodiment 2 of the present invention provides a tunnel lining crack identification system based on panoramic spatial image stitching, comprising: The local stitching module is configured to acquire an array of tunnel lining images and use a stitching model based on cylindrical surface projection to locally stitch the tunnel lining images to obtain a locally stitched image. The global stitching module is configured to use a global image space model to fuse and stitch local stitched images to obtain a panoramic stitched data map; The anomaly detection module is configured to detect abnormal areas in the panoramic stitched data map based on local spatial differences, and obtain the identification results of tunnel lining cracks.

[0133] Example 3: Embodiment 3 of the present invention provides a computer-readable storage medium storing a computer program adapted for loading by a processor and executing the steps in the tunnel lining crack identification method based on panoramic spatial image stitching as described in Embodiment 1 of the present invention.

[0134] Example 4: Embodiment 4 of the present invention provides a computer device, the device comprising: A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program, which, when executed by the processor, implements the steps in the tunnel lining crack identification method based on panoramic spatial image stitching as described in Embodiment 1 of the present invention.

[0135] The steps and methods involved in Examples 2, 3 and 4 above correspond to those in Example 1. For specific implementation details, please refer to the relevant description section of Example 1.

[0136] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, an optical medium, or a semiconductor medium, etc. The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for tunnel lining crack identification based on panoramic space image stitching, characterized in that, The method comprises the following steps: An image array of a tunnel lining is acquired, and a local stitching of the tunnel lining images is performed by using a stitching model based on cylindrical projection to obtain a local stitching image; A global image space model is used to perform fusion stitching on the local stitching image to obtain a panoramic stitching data graph; An abnormal area detection is performed on the panoramic stitching data graph according to spatial local difference to obtain a tunnel lining crack identification result.

2. The tunnel lining crack identification method based on panoramic space image stitching according to claim 1, wherein, The specific steps of performing local stitching on the tunnel lining images by using the stitching model based on cylindrical projection are as follows: A projection stitching model based on a cylindrical surface is constructed; The tunnel lining images are registered by using a spatial priori constraint to obtain registered images; Consistency optimization is performed on the camera parameters; The registered images are fused and integrally stitched. 3.The tunnel lining crack identification method based on panoramic space image stitching according to claim 2, wherein, The specific steps of performing fusion stitching on the local stitching image by using the global image space model are as follows: Anti-warping remapping and global geometric alignment are performed on the local stitching image; An image fusion algorithm based on a Laplace pyramid is used to fuse the local stitching image to obtain a preliminary global stitching image; An image matching method for weak texture surfaces is used to optimize the preliminary global stitching image. 4.The tunnel lining crack identification method based on panoramic space image stitching according to claim 2, wherein, The specific steps of optimizing the preliminary global stitching image by using the image matching method for weak texture surfaces are as follows: A weak texture feature matching method based on LightGlue is used to extract features from the preliminary global stitching image; A DISK and LightGlue fusion method is used to perform feature detection and matching on the extracted features. 5.The tunnel lining crack identification method based on panoramic space image stitching according to claim 4, wherein, The specific steps of performing abnormal area detection on the panoramic stitching data graph according to spatial local difference are as follows: An image abnormal defect detection method based on spatial local correlation is used to perform image representation on the panoramic stitching data graph; A multi-scale fusion strategy is used to extract abnormal features of the panoramic stitching data graph. 6.The tunnel lining crack identification method based on panoramic space image stitching according to claim 1, wherein, The specific steps of performing image representation on the panoramic stitching data graph by using the image abnormal defect detection method based on spatial local correlation are as follows: An abnormal detection model is constructed based on a PatchCore abnormal detection algorithm; A memory bank enhanced in spatial locality is constructed, and the abnormal detection model is trained by using data in the memory bank; A spatial constraint training and testing strategy is set, and the abnormal detection model is further optimized by using the spatial constraint training and testing strategy.

7. The tunnel lining crack identification method based on panorama space image stitching according to claim 6, wherein, The specific steps of extracting abnormal features of the panoramic stitching data graph by using the multi-scale fusion strategy are as follows: ResNet-50 is used to extract multi-level features from the panoramic stitching data graph; The extracted multi-level features are aggregated in a cross-level local neighborhood to obtain aggregated features; The aggregated features are reduced in dimension and optimized. 8.A tunnel lining crack identification system based on panoramic space image stitching, characterized in that, The method comprises: A local stitching module configured to acquire an image array of a tunnel lining, and perform local stitching on the tunnel lining images by using a stitching model based on cylindrical projection to obtain a local stitching image; A global stitching module configured to perform fusion stitching on the local stitching image by using a global image space model to obtain a panoramic stitching data graph; An abnormal detection module configured to perform abnormal area detection on the panoramic stitching data graph according to spatial local difference to obtain a tunnel lining crack identification result.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by the processor to implement the tunnel lining crack identification method based on panoramic spatial image stitching according to any one of claims 1-7.

10. A computer device, comprising: Comprise: a processor suitable for executing a computer program; a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the tunnel lining crack identification method based on panoramic spatial image stitching according to any one of claims 1-7.

Citation Information

Cited By

  • Multi-scale lining image splicing method under tunnel detection parallax scene and computer program product

    CN122089564A

  • A multi-scale lining image stitching method in a tunnel detection parallax scene and a computer program product

    CN122089564B

  • UAV image stitching method based on LightGlue and adaptive multi-band fusion

    CN122175780A

  • Unmanned aerial vehicle image splicing method based on LightGlue and adaptive multi-band fusion

    CN122175780B