Pipeline construction monitoring method and system based on image recognition

By using spot array analysis of motion parameters and time-domain analysis technology, the problems of dust and vibration interference in image recognition technology during shield tunnel construction have been solved, enabling accurate identification and location of minute defects in welds and meeting the needs of real-time monitoring and control of construction quality.

CN121982000BActive Publication Date: 2026-08-04CCCC SECOND HIGHWAY ENG BUREAU RAILWAY CONSTR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CCCC SECOND HIGHWAY ENG BUREAU RAILWAY CONSTR CO LTD
Filing Date
2026-01-27
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing image recognition technology is subject to dual interference from high concentrations of metal dust and high-frequency vibrations of equipment during shield tunnel construction, resulting in insufficient reliability in identifying minute defects in welds and making it difficult to meet the requirements for real-time monitoring and precise control of construction quality.

Method used

By analyzing motion parameters through a projected light spot array and combining temporal analysis and texture reconstruction techniques, inverse jitter compensation is achieved to eliminate the influence of dust occlusion. Global motion flow field generation, physical calibration sequence, texture reconstruction, and defect marker map generation are used to accurately identify minute defects in the weld.

Benefits of technology

It effectively suppresses dust and vibration interference, improves the stability and effectiveness of image data, reduces the false defect misjudgment rate, realizes accurate identification and positioning of minor weld defects, and meets the real-time monitoring needs of construction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982000B_ABST
    Figure CN121982000B_ABST
Patent Text Reader

Abstract

This invention discloses a pipeline construction monitoring method and system based on image recognition, belonging to the field of pipeline construction technology. The method includes: projecting a light spot array onto a monitoring scene containing welds; analyzing the displacement of the light spot array in an image sequence to resolve affine transformation parameters and generate a global motion flow field; using the global motion flow field to perform inverse compensation on the image sequence to generate a physical calibration sequence; performing signal analysis on each pixel in the time dimension on the physical calibration sequence to generate a temporal modulation intensity map, and performing morphological dilation and connected component analysis on the temporal modulation intensity map. This method can suppress the dual interference of high-concentration metal dust and high-frequency vibration of equipment during underground shield tunnel construction. It achieves inverse jitter compensation by resolving motion parameters through the light spot array, and eliminates the influence of dust occlusion by combining temporal analysis and texture reconstruction techniques, avoiding the problem of traditional methods accidentally erasing defect features during noise reduction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pipeline construction technology, and more specifically, to a pipeline construction monitoring method and system based on image recognition. Background Technology

[0002] Image recognition technology has the advantages of being non-contact, highly efficient, and capable of automated detection. It has been widely used in the detection of defects such as weld cracks and porosity, gradually replacing traditional manual visual inspection methods. This technology has also been applied in the inspection of welds in pipeline construction. However, the construction environment of underground shield tunnels is special, and there are some interferences. First, the high concentration of metal dust generated can easily adhere to the lens of the acquisition equipment, resulting in images full of noise and obscuring the weld contour and subtle defect features. Second, the high-frequency vibration during the advancement of the shield machine can cause equipment shaking, resulting in frame misalignment and distortion of the image sequence.

[0003] Existing noise reduction solutions mostly employ strong spatial filtering algorithms. While these methods can effectively eliminate dust interference, they also tend to erase minute defect features such as cracks and pores, leading to missed detections. Regarding jitter compensation, mainstream techniques rely on inter-frame feature point matching for image alignment. However, dust noise is often misidentified as valid feature points, resulting in a significant decrease in alignment accuracy. This not only fails to eliminate the negative impact of jitter but also exacerbates the distortion effect caused by noise. Under this dual interference, simply increasing image resolution to capture micro-defects will not only amplify motion artifacts but also exacerbate alignment distortion caused by dust noise, ultimately causing the defect signal to be completely masked by noise. Therefore, existing methods are insufficiently reliable in identifying minute weld defects in the complex construction environment of shield tunnels, making it difficult to meet the real-time quality monitoring and precise control requirements during construction. Summary of the Invention

[0004] To address the problems existing in the prior art, the present invention aims to provide a pipeline construction monitoring method and system based on image recognition. This method can suppress the dual interference of high-concentration metal dust and high-frequency vibration of equipment during underground shield tunnel construction. It achieves reverse compensation of jitter by analyzing motion parameters through a spot array, and eliminates the influence of dust occlusion by combining time-domain analysis and texture reconstruction technology. This avoids the problem of accidentally erasing defect features during the noise reduction process of traditional methods.

[0005] To solve the above problems, the present invention adopts the following technical solution: Firstly, a pipeline construction monitoring method based on image recognition, the method comprising: A light spot array is projected onto the monitoring scene containing the weld seam, and the displacement of the light spot array in the image sequence is analyzed to resolve the affine transformation parameters and generate a global motion flow field. The image sequence is inversely compensated using the global motion flow field to generate a physical calibration sequence; Perform signal analysis on each pixel in the time dimension on the physical calibration sequence to generate a temporal modulation intensity map, and perform morphological dilation and connected component analysis on the temporal modulation intensity map to generate a binarized dust mask. Invalid data in the physical calibration sequence is marked using a binarized dust mask. Based on the valid pixel information of the unmarked area and the geometric prior of the weld, the texture of the marked area is reconstructed by solving the anisotropic diffusion equation, and the image sequence after texture reconstruction is output. Inter-frame difference maps are calculated on the image sequence after texture reconstruction. Based on historical data of global motion flow field, a normal deformation prediction model is established. The model is subtracted from the inter-frame difference map to obtain an abnormal motion residual map. The abnormal motion residual map is then segmented to obtain the preliminary defect area. Based on the historical distribution of the binarized dust mask and the topological relationship between the preliminary defect area and the weld, a defect marking map is output.

[0006] Furthermore, the generation of the global motion flow field includes: A composite light spot pattern is projected onto the monitoring scene. The composite light spot pattern is composed of a dominant frequency sine stripe and a moiré circular spot. Local micropattern analysis was performed on the moiré spot region, and the absolute phase value of each spot was calculated by two-dimensional phase unwrapping operation; Matching is performed between consecutive frames based on absolute phase values. The search area is predicted based on the motion information of the preceding frame and then matched within the search area. The matching point pairs are then verified based on phase continuity and motion consistency to filter out abnormal matches, thus obtaining the displacement observation set. Multiple motion models are fitted in parallel using a pure displacement observation set. The spatiotemporal distribution characteristics of the residuals of the multiple motion models are calculated. Based on the goodness of fit of each motion model and the spatiotemporal distribution characteristics of the residuals, the final model is selected from the multiple motion models to generate the global motion flow field.

[0007] Furthermore, the generation of the physical calibration sequence includes: An adaptive sampling grid is generated based on the model type of the global motion flow field. The grid points are then inversely mapped to the original image and interpolated according to the global motion flow field. Based on the residual distribution map corresponding to the global motion flow field, analyze the residual fluctuation of each position in the time series of the remapped image, calculate the pixel-level motion calibration confidence and mark the absolutely reliable anchor points; Using absolutely reliable anchor points and pixels with confidence levels higher than a first predetermined threshold as boundaries, an anisotropic diffusion process constrained by pixel-level motion calibration confidence is performed to repair regions with confidence levels lower than a second predetermined threshold. The pixel value difference in the overlapping area of ​​pixels with confidence scores higher than a first predetermined threshold in adjacent frames is compared, and temporal smoothing filtering is performed on areas with discontinuous jumps to output a physical calibration sequence.

[0008] Furthermore, the generation of the time-domain modulation intensity map includes: The pixel time history signal is preprocessed based on pixel-level motion calibration confidence to extract differential time history signals containing abnormal fluctuations, and a set of time-frequency atoms is obtained by matching and tracking the differential time history signals. Calculate the distribution entropy and energy concentration of time-frequency atoms in the time-frequency domain, and fuse them to generate a pixel-level modulation intensity scalar; Using the modulation intensity scalar as the anchor point, interpolation is performed under the constraints of the image gradient and confidence map to generate a temporal modulation intensity map.

[0009] Furthermore, the generation of the binarized dust mask includes: Based on the temporal modulation intensity map, and using pixel-level motion calibration confidence as the adjustment criterion, local threshold segmentation is performed to generate a preliminary binary map; Temporal persistence verification and morphological cleanup are performed on the connected regions in the preliminary binary graph; The cleaned area mask is fused across frames, and the fusion result is trimmed according to the confidence map to generate a binarized dust mask.

[0010] Furthermore, the texture of the labeled region is reconstructed by solving the anisotropic diffusion equation, and the resulting image sequence with reconstructed texture is output, including: The structural tensor field is calculated within the effective region of the physical calibration sequence and anisotropic smoothing is performed. Using the smoothed structure tensor field as the source, tensor voting is performed on the binarized dust mask region to generate an inference structure tensor field covering the entire image; Anisotropic diffusion equations are constructed based on the inference structure tensor field, and iterative solutions are performed using the effective region pixel values ​​as boundary conditions to reconstruct the texture base of the invalid region. Calculate the dense optical flow field of the current frame and adjacent frames in the effective region. Based on the dense optical flow field, project the texture details of adjacent frames to the invalid region of the current frame. Combine the pixel-level motion calibration confidence and the closure error of the optical flow projection to fuse the image sequence after texture reconstruction.

[0011] Furthermore, the generation of the anomalous motion residual map includes: Spatiotemporal decomposition of historical global motion flow field sequences is performed to extract spatial modes and their temporal evolution patterns. Periodic and trend-based predicted deformation fields are synthesized and superimposed to generate a predicted global deformation field. By utilizing the geometric priors of the weld, the predicted global deformation field is anisotropically modulated in the weld region to generate a physically constrained deformation prediction field. Multi-resolution pyramid analysis is performed on the physical constraint deformation prediction field and the inter-frame difference map. By fusing the residuals of each scale across scales, an abnormal motion residual map is generated.

[0012] Furthermore, the acquisition of the initial defect area includes: Anomaly seed points are determined based on the abnormal motion residual map, and competitive region growth is performed starting from the anomaly seed points to generate candidate defect regions and competition intensity scores. Based on the competition intensity score, and verifying the spatiotemporal coherence of each candidate defect region in the abnormal motion residual map of multiple consecutive frames, the preliminary defect regions are obtained after screening.

[0013] Furthermore, the generation of the defect marker map includes: Candidate defect regions are comprehensively verified from three dimensions: dust coverage history, geometric morphology and topological relationship, and temporal causal consistency. The comprehensive verification results are integrated with the competition intensity score to generate a defect label map labeled with defect type and detection confidence level.

[0014] Secondly, the present invention provides a pipeline construction monitoring system based on image recognition, specifically including: a flow field analysis module, used to project a light spot array onto a monitoring scene containing welds, analyze the displacement of the light spot array in an image sequence to resolve affine transformation parameters, and generate a global motion flow field; The image calibration module is used to perform inverse compensation on the image sequence using the global motion flow field to generate a physical calibration sequence; The mask creation module is used to perform signal analysis of each pixel in the time dimension on the physical calibration sequence, generate a temporal modulation intensity map, and perform morphological dilation and connected component analysis on the temporal modulation intensity map to generate a binarized dust mask. The texture restoration module is used to mark invalid data in the physical calibration sequence using a binarized dust mask. Based on the valid pixel information of the unmarked area and the geometric prior of the weld, the texture of the marked area is reconstructed by solving the anisotropic diffusion equation, and the image sequence after texture reconstruction is output. The defect screening module calculates the inter-frame difference map on the image sequence after texture reconstruction, establishes a normal deformation prediction model based on historical data of global motion flow field, subtracts the model from the inter-frame difference map to obtain the abnormal motion residual map, and segments the abnormal motion residual map to obtain the preliminary defect area. The defect screening module filters based on the historical distribution of the binary dust mask and the topological relationship between the preliminary defect area and the weld, and outputs a defect marking map.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This scheme can effectively suppress the dual interference of high concentration of metal dust and high frequency vibration of equipment in the construction of underground shield tunnels. It achieves reverse compensation of jitter by analyzing motion parameters through spot array, and eliminates the influence of dust occlusion by combining time domain analysis and texture reconstruction technology. It avoids the problem of accidentally erasing defect features in the noise reduction process of traditional methods, and improves the effectiveness and stability of image data in complex environments.

[0016] (2) This scheme constructs a normal deformation prediction model by decomposing the historical global motion flow field in time and space, accurately separates normal deformation from defect abnormal signals, and then verifies and screens through multiple dimensions such as geometric shape, temporal continuity, and dust history, which greatly reduces the false defect misjudgment rate, realizes the accurate identification and positioning of weld minor defects, and meets the precise control requirements of real-time monitoring of construction quality.

[0017] (3) Based on effective pixel information and weld geometry prior, this scheme reconstructs the texture of the dust-covered area through the structural tensor field and anisotropic diffusion equation, and optimizes the texture details by combining optical flow projection to ensure that the reconstructed texture is naturally connected with the actual weld texture and has strong authenticity, providing high-quality and high-reliability image data support for defect detection.

[0018] (4) This solution is suitable for complex construction scenarios of underground shield tunnels. The technical steps are closely connected and can be integrated with the existing image monitoring system without complex hardware upgrades. It can realize real-time monitoring of the construction process and output intuitive maps with marked defect types and confidence levels, thereby reducing the cost of construction quality control and improving the efficiency and scientific nature of control. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0020] Figure 1 This is a flowchart of the pipeline construction monitoring method based on image recognition according to the present invention; Figure 2 This is a flowchart illustrating the relationships between the various modules in the pipeline construction monitoring system based on image recognition of the present invention. Detailed Implementation

[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0022] Example 1: Please see Figure 1 The specific operation steps of the pipeline construction monitoring method based on image recognition are as follows: Step 1: Project a light spot array onto the monitoring scene containing the weld, analyze the displacement of the light spot array in the image sequence to resolve the affine transformation parameters, and generate a global motion flow field. The specific operations are as follows: When monitoring pipeline construction scenarios including welds, a pre-set array of light spots is first projected onto the monitoring area. This array must be stable and have clear phase characteristics to accurately capture its positional changes in subsequent image sequences. By acquiring a continuous image sequence containing this array of light spots, the pixel positions of the array of light spots in each frame of the sequence are located and tracked. The displacement of the array of light spots between adjacent frames and between consecutive frames is analyzed. Based on the displacement data of the array of light spots, combined with the mathematical properties of affine transformation, the affine transformation parameters describing the global motion of the image are obtained. These parameters can characterize the overall motion state of the image acquisition device during the monitoring process, including motion components such as translation, rotation, and scaling. Using the obtained affine transformation parameters, a global motion flow field covering the entire image area is constructed. This motion flow field can accurately reflect the motion trajectory and trend of each pixel in the image sequence.

[0023] The generation of the global motion flow field also includes the following steps: Step 11: Project a composite light spot pattern onto the monitoring scene. The composite light spot pattern is formed by superimposing the dominant frequency sine stripes and moiré circular spots. The specific operation is as follows: To achieve accurate capture of motion information in the monitoring scene, a composite light spot pattern is projected onto the monitoring scene containing the weld. The composite light spot pattern adopts a design of superimposing a dominant frequency sinusoidal stripe and a moiré circle. The dominant frequency sinusoidal stripe is a periodic stripe structure with a fixed spatial frequency, and its periodicity facilitates subsequent phase analysis and matching. The moiré circle is formed by the interference of two or more sets of stripes and has unique local micro-pattern features, which can serve as a carrier of stable feature points in the image. The superposition design of the two patterns can not only use the global periodicity of the sinusoidal stripe to achieve preliminary positioning of large-scale motion, but also use the local uniqueness of the moiré circle to achieve high-precision feature matching.

[0024] Step 12: Perform local micro-pattern analysis on the moiré circle region, and calculate the absolute phase value of each circle using two-dimensional phase unwrapping operation. The specific operation is as follows: For the acquired image sequence containing composite spot patterns, the moiré circle regions are first located. Local micro-pattern analysis is then performed on each moiré circle region to extract key information such as grayscale distribution, contour features, and fringe interference details. Because the phase information of the moiré circles exhibits encapsulation—that is, the phase values ​​are restricted to a certain range and cannot directly reflect the true positional relationship—two-dimensional phase unwrapping operations are needed to eliminate this encapsulation ambiguity. Based on the principle of phase continuity and combined with the spatial positional constraints of the spot regions, the two-dimensional phase unwrapping operation processes the encapsulated phase data to eliminate... Phase blurring is performed as an integer multiple, thereby calculating the absolute phase value of each moiré circle. This absolute phase value is unique and can accurately characterize the spatial position of the moiré circle in the image.

[0025] Step 13: Matching is performed between consecutive frames based on absolute phase values. The search area is predicted based on the motion information of the preceding frame, and matching is performed within the search area. Then, the matching point pairs are verified based on phase continuity and motion consistency to filter out abnormal matches, thus obtaining the displacement observation set. The specific operations are as follows: Using the absolute phase value of each moiré circle calculated in step 12 as a matching feature, corresponding matching of moiré circles is performed between consecutive frames of the image sequence. To improve matching efficiency and accuracy and reduce invalid search range, the possible search area of ​​the moiré circle in the current frame is predicted based on the motion information obtained from the previous frame, including parameters such as motion direction and motion speed. The matching operation is limited to this predicted search area to avoid redundant calculations and mismatches caused by global search. Within the predicted search area, the matching point pairs of moiré circles between frames are initially determined by comparing the similarity of absolute phase values. Subsequently, the preliminary matching point pairs are verified based on two dimensions: phase continuity and motion consistency. Phase continuity requires that the absolute phase value of the matching point pairs in adjacent frames changes smoothly without abrupt changes. Motion consistency requires that the motion trajectory of the matching point pairs is consistent with the overall motion trend of the image. Through the above verification process, abnormal matching point pairs caused by dust interference, image noise, and other factors are filtered out, and finally a clean displacement observation set is obtained, which contains reliable inter-frame moiré circle displacement data.

[0026] Step 14: Use the pure displacement observation set to fit multiple motion models in parallel, calculate the spatiotemporal distribution characteristics of the residuals of multiple motion models, and select the final model from multiple motion models based on the goodness of fit and the spatiotemporal distribution characteristics of the residuals to generate the global motion flow field. The specific operations are as follows: Using the clean displacement observation set obtained after filtering in step 13 as input data, this observation set contains the coordinates and corresponding displacement data of valid matching point pairs verified between consecutive frames, specifically represented by the pixel coordinates of the preceding frame for each matching point pair. Current frame pixel coordinates and displacement components Based on this observation set, multiple commonly used motion models are simultaneously fitted in parallel, including translation, rotation, and affine models. Parameter estimation for each model is performed using the least squares method. For the translation model, its parameterized expression is: , ,in , They are respectively , When solving for directional translation using the least squares method, the objective is to minimize the sum of squared residuals between the observed displacement values ​​and the model prediction values.

[0027] In the translation model, the objective function S is constructed as the sum of squares of the deviations between the displacement components at each observation point and the model parameters a and b, i.e., summing over all observation points. The deviation at each point includes the displacement in the x-direction. The squared difference with parameter a, and the displacement in the y direction Taking the squared difference between parameter a and parameter b, and then taking the partial derivatives with respect to parameters a and b respectively, and setting them to zero, we can solve for parameter a equal to all... The arithmetic mean of the observations, where parameter b equals all The arithmetic mean of the observed values.

[0028] For the rotational model, the displacement change is parameterized as a function of the rotation angle θ: displacement change in the x-direction equal to the initial coordinates Multiply by the cosine of the rotation angle θ and subtract the initial coordinates Multiply by the sine of the rotation angle θ, then subtract the reference displacement. ; displacement change in the y direction equal to the initial coordinates Multiply by the sine of the rotation angle θ and add the initial coordinates Multiply by the cosine of the rotation angle θ, then subtract the reference displacement. By constructing the objective function of residual sum of squares and solving it using an iterative numerical method, the optimal estimate of the rotation angle θ can be obtained.

[0029] For the affine model, displacement changes are represented by six parameters. to Linear representation: displacement change in the x-direction It is a parameter Multiply by the initial coordinates Add parameters Multiply by the initial coordinates Add parameters ; displacement change in the y direction It is a parameter Multiply by the initial coordinates Add parameters Multiply by the initial coordinates Add parameters An overdetermined system of equations is constructed based on the least squares method, and the estimated values ​​of these parameters are solved by numerical methods such as matrix inversion or QR decomposition.

[0030] After the model parameters are fitted, it is necessary to evaluate the model's fitting effect. The residual is defined as the difference between the actual displacement and the model's predicted displacement at each observation point, including the residual component in the x-direction. and y-direction residual components ,in Equal to actual Observed values ​​minus model predictions value, Similarly, the spatial distribution characteristics of residuals are quantified by dividing the image into regions according to a preset grid and calculating the mean, variance, extreme values, and distribution histogram of the residuals in each region, thereby assessing the degree of spatial clustering and uniformity of the residuals. Temporal distribution characteristics are characterized by analyzing the changes in the residual sequence at the same spatial location in consecutive frames, calculating the standard deviation, coefficient of variation, and trend term fitted by linear regression, to characterize the fluctuation pattern of residuals over time. Additionally, commonly used indicators for calculating the goodness of fit of the model include the coefficient of determination R0. 2 Root mean square error (RMSE) and coefficient of determination (R²) 2 Based on residual decomposition, the total sum of squares of deviation is decomposed into the regression sum of squares and the residual sum of squares. The formula is 1 minus the ratio of the residual sum of squares to the total sum of squares of deviation, where the total sum of squares of deviation is the sum of squares of the deviations of the observed displacement from its mean. R0 2 The closer the value is to 1, the stronger the model's ability to interpret displacement observation data and the better the fit. The root mean square error (RMSE) is the square root of the average of the sum of squared residuals. Specifically, it is calculated as the square root of the sum of squared residuals divided by twice the total number of observation points, and is used to quantify the overall magnitude of the prediction error.

[0031] The RMSE (Residual Squared Error) formula is derived from the square root of the mean of the sum of squared residuals. It quantifies the average deviation between model predictions and actual observations. A smaller RMSE indicates higher model prediction accuracy. In the model selection stage, considering both goodness of fit and the spatiotemporal distribution characteristics of the residuals, models with a lower RMSE are first selected. 2Candidate models with a residual spatial distribution uniformity of ≥0.8 and an RMSE in the minimum range are then ranked based on their residual spatial distribution uniformity and temporal fluctuation stability. Finally, the model with the best fit and uniform residual distribution with small fluctuations in the spatiotemporal dimension is selected as the final motion model. The uniformity of residual spatial distribution is judged by the spatial variation coefficient of the residual variance; the smaller the variation coefficient, the more uniform the distribution. Temporal fluctuation stability is judged by the standard deviation of the residual sequence; the smaller the standard deviation, the more stable the model. Based on the parameters of the final model, combined with the image resolution and pixel coordinate system of the monitoring scene, a global motion flow field is generated. The motion trajectory of each pixel in this flow field is calculated by the parameters of the final model, which can accurately reflect the motion laws of each pixel in each frame of the image, such as translation, rotation, scaling, or shearing.

[0032] In some embodiments of the present invention, step 2 is further included, which involves using the global motion flow field to perform inverse compensation on the image sequence to generate a physical calibration sequence. The specific operation is as follows: The global motion flow field has accurately described the motion trajectory and motion parameters of each pixel in each frame of the image. The reverse compensation process is based on these parameters to reverse the deduction and map the image pixels that are shaking due to the vibration of the tunnel boring machine back to their ideal positions in the state without shaking. This is to offset the inter-frame misalignment and distortion caused by the vibration. In the reverse compensation process, it is necessary to combine the grayscale information and pixel correlation of the image to reasonably correct and optimize the mapped pixel values ​​to ensure the texture continuity and authenticity of the compensated image. By performing this reverse compensation operation on each frame of the image sequence in sequence, the inter-frame motion interference is eliminated, and finally a physical calibration sequence with accurate inter-frame alignment, clear texture and no obvious shaking distortion is generated.

[0033] The generation of the physical calibration sequence also includes the following steps: Step 21: Generate an adaptive sampling mesh based on the model type of the global motion flow field. Then, inversely map the mesh points to the original image and perform interpolation based on the global motion flow field. The specific operations are as follows: First, based on the model type corresponding to the generated global motion flow field, the structure and density of the adaptive sampling mesh are determined. Different motion models have different motion patterns, and the sampling mesh needs to adapt to the model characteristics. For example, affine models contain scaling and shearing components, so a non-uniform mesh with adaptive density adjustment is required. The mesh nodes are densified in areas with drastic motion changes and appropriately sparsed in areas with stable motion to balance compensation accuracy and computational efficiency. After the mesh is generated, based on the motion parameters of each pixel in the global motion flow field, each node in the sampling mesh is inversely mapped from its coordinate position in the current frame image to its corresponding coordinate position in the original image. This original image is the uncompensated image. Since the coordinates of some mesh nodes after inverse mapping may not be integer pixel positions, an interpolation algorithm is used to obtain the pixel grayscale value at that position. The commonly used interpolation method is bilinear interpolation. By calculating the weighted average of the grayscale values ​​of the four integer pixels surrounding the node, the grayscale value of the mapped node is obtained, ensuring that the grayscale information of the sampling mesh is complete and continuous.

[0034] Step 22: Based on the residual distribution map corresponding to the global motion flow field, analyze the residual fluctuations at each location in the remapped image over time, calculate the pixel-level motion calibration confidence, and mark absolutely reliable anchor points. The specific operations are as follows: After the global motion flow field fitting is completed, the residual distribution map generated can be used to extract the residual data of each pixel position in the remapped image over a continuous time series. These data directly reflect the prediction accuracy of the global motion model for the motion state of each pixel: the smaller the fluctuation of the residual, the more accurate the model prediction, and the higher the reliability of motion calibration for that pixel. For each pixel, the characteristics of its residual sequence fluctuation over time are first calculated. The quantitative indicators are the standard deviation and coefficient of variation of the residual, which are used to measure the magnitude of the residual change. On this basis, a pixel-level motion calibration confidence calculation model is constructed. This model takes the fluctuation characteristics of the residual as input and integrates the pixel's own gray-level stability, such as gray-level gradient measurement and auxiliary information such as the motion consistency of neighboring pixels. Finally, a confidence value between 0 and 1 is generated for each pixel. The design logic is: the smaller the residual fluctuation of a pixel, the more stable its gray level, and the more consistent its motion with the surrounding pixels, the more reliable its motion calibration result, and the higher the corresponding confidence should be.

[0035] The specific formula for calculating confidence reflects the above logic, and the coordinates... Confidence of a pixel Defined as: 1 divided by a denominator consisting of 1 plus a weighted sum of two terms, where the first term is the residual coefficient of variation. Multiply by weighting factor , It can be obtained by the ratio of the residual standard deviation to the mean; the second term is the grayscale gradient value. Multiply by weighting factor , Reflecting the drasticness of local grayscale changes, the formula is derived based on the negative correlation between confidence level and these two factors. and The smaller the value, the closer the denominator is to 1, and the higher the confidence level. The closer the value is to 1, the higher the reliability, and the higher the weighting coefficient. and These are all positive constants used to balance the influence of residual fluctuations and grayscale variations on confidence levels. Their specific values ​​are, for example... It is usually between 0.3 and 0.5. Between 0.5 and 0.7, the confidence level can be determined by training based on data samples from the actual monitoring scenario. After calculating the confidence level of all pixels in the entire image, a threshold is set, usually 0.95. Pixels with a confidence level greater than or equal to this threshold and that remain stable in multiple consecutive frames, such as 3 to 5 frames, are marked as absolutely reliable anchors. These pixels have high and stable confidence in motion estimation.

[0036] Step 23: Using absolutely reliable anchor points and pixels with confidence levels higher than the first predetermined threshold as boundaries, perform an anisotropic diffusion process constrained by pixel-level motion calibration confidence to repair regions with confidence levels lower than the second predetermined threshold. The specific operations are as follows: In the image restoration stage, the absolutely reliable anchor points marked in the previous step and pixels with high confidence values ​​are used as fixed boundary constraints. The gray values ​​of these boundary pixels will remain unchanged during the restoration process. For areas with low confidence, that is, areas with confidence values ​​less than the second predetermined threshold of 0.8, an anisotropic diffusion process with pixel-level motion calibration confidence constraints is implemented to complete the restoration. The principle of this process is to guide gray-level information to be smoothly transmitted along the image texture direction, while suppressing diffusion in the direction of crossing texture edges to maintain the original gray-level differences, thereby avoiding detail blurring caused during the restoration process. In this process, pixel-level confidence values ​​are used to control the intensity of diffusion: high-confidence pixels have a stronger influence on the gray-level of their surrounding low-confidence pixels, while the mutual influence between low-confidence pixels is weaker, thus ensuring that the texture of the restored image is consistent with the real scene.

[0037] Anisotropic diffusion is achieved by numerically solving an anisotropic diffusion equation. This equation draws inspiration from the physical principle of heat conduction, analogizing the propagation of image grayscale values ​​to the diffusion of heat. A diffusion coefficient controls the rate and direction of propagation. The equation states that the rate of change of the grayscale value at a specific point in the image over time is equal to the divergence of the product of the diffusion coefficient and the grayscale gradient at that point. This diffusion coefficient is not constant but dynamically determined by two factors: first, the motion calibration confidence of the pixel itself; and second, the magnitude of the image gradient at that point, i.e., the drasticness of the grayscale change. Specifically, the diffusion coefficient equals the confidence value of the pixel multiplied by... A decay function with negative gradient magnitude squared as the exponent is used, where the decay strength of the exponent term is adjusted by a constant k, typically ranging from 0.01 to 0.05. The image gradient represents the direction and rate of change of the most significant grayscale change at that point. The diffusion process is solved numerically using an iterative method. In each iteration, the grayscale values ​​of all pixels are updated according to the above rules. The iteration terminates when the total grayscale change of all pixels in the image is less than a preset threshold in two adjacent iterations. Thus, the grayscale of the low-confidence region is repaired, and a natural connection is achieved with the texture of the surrounding high-confidence region. This preset threshold is usually 0.5.

[0038] Step 24: Compare the pixel value differences in the overlapping areas of pixels with confidence scores higher than a first predetermined threshold in adjacent frames, perform temporal smoothing filtering on areas with discontinuous transitions, and output a physical calibration sequence. The specific operations are as follows: For the image sequence repaired in step 23, high-confidence overlapping regions between adjacent frames are extracted. These regions are the intersections of pixels in both frames that are marked as high-confidence. A pixel with a confidence level greater than a first predetermined threshold of 0.8 is marked as high-confidence. The first predetermined threshold is equal to a second predetermined threshold. Such regions have high reliability in terms of grayscale information and can be used as a benchmark for inter-frame consistency verification. The grayscale values ​​of the high-confidence overlapping regions in adjacent frames are compared pixel by pixel, and the grayscale difference between corresponding pixels is calculated. By statistically analyzing the distribution characteristics of the grayscale difference, a threshold for judging discontinuous grayscale transitions is determined. This threshold is usually determined based on the image grayscale range, such as 0 to 255, and the noise level of the monitored scene. The threshold for judging discontinuous grayscale transitions is generally set between 15 and 25. When the grayscale difference of pixels in a certain region exceeds this threshold, and multiple consecutive pixels... If such differences exist in 3 to 4 pixels, it is determined that there is a discontinuous grayscale jump in the region. Such jumps are mostly caused by residual noise or local motion compensation deviations, and temporal smoothing filtering is required. The temporal smoothing filtering adopts a weighted average strategy based on the time dimension. Taking the grayscale value of the current frame pixel as the core, it combines the grayscale values ​​of the corresponding pixels in the previous and next frames to calculate the weighted average value as the corrected grayscale value of the current frame pixel. The weight coefficient is set according to the principle that the closer the time is, the greater the weight. Usually, the weight of the current frame is 0.6, and the weights of the previous and next frames are each 0.2. By performing this temporal smoothing filtering operation on all regions with discontinuous jumps, the grayscale abrupt changes between frames are eliminated, ensuring that the grayscale transition between adjacent frames is natural and continuous. After the above processing, the output is a physical calibration sequence with accurate inter-frame alignment, clear texture, and continuous grayscale.

[0039] In some embodiments of the present invention, step 3 is further included: performing signal analysis on each pixel in the time dimension on the physical calibration sequence to generate a temporal modulation intensity map, and performing morphological dilation and connected component analysis on the temporal modulation intensity map to generate a binarized dust mask. The specific operations are as follows: After obtaining the physical calibration sequence generated in step 2, for each pixel in the sequence, the grayscale change signal in the continuous frame time dimension is extracted. The fluctuation characteristics of pixel grayscale over time are mined through signal analysis to distinguish the abnormal fluctuations caused by dust adhesion from the normal grayscale changes in the weld area. Based on the analysis results, a temporal modulation intensity map is generated. The pixel grayscale value in the map represents the signal modulation intensity in the time dimension at the corresponding position. The dust-adhered area will show a significant difference in modulation intensity due to the grayscale fluctuation pattern being different from the normal area. Subsequently, a morphological dilation operation is performed on the temporal modulation intensity map to fill the small gaps in the dust area in the intensity map, making the dust area outline more complete. Then, through connected component analysis, connected regions that conform to the spatial distribution characteristics of dust are selected to eliminate the interference of isolated noise points. Based on the above processing, a binarized dust mask is finally generated. In the mask, a specific grayscale value, usually 255, is used to mark the invalid area covered by dust, and another grayscale value, usually 0, is used to mark the valid area without dust interference.

[0040] The generation of the time-domain modulation intensity map includes the following steps: Step 31: Based on pixel-level motion calibration confidence, preprocess the pixel time history signal to extract the differential time history signal containing abnormal fluctuations, and perform matching tracking on the differential time history signal to obtain a set of time-frequency atoms. The specific operations are as follows: First, the time-history signal of each pixel in the physical calibration sequence is extracted. This signal is a sequence of grayscale values ​​of the pixel in the time dimension of consecutive frames, denoted as a one-dimensional sequence composed of the grayscale data of the pixel in each frame. Based on the pixel-level motion calibration confidence generated in step 22, the pixel time-history signal is preprocessed to remove the grayscale data corresponding to frames with confidence scores below a preset threshold. At the same time, the retained grayscale data is linearly smoothed to eliminate slight random noise interference. After preprocessing, the time-history signal is differentially processed to calculate the difference between the grayscale values ​​of adjacent frames, resulting in a differential time-history signal. This signal can highlight abnormal fluctuation components in the time-history signal. Dust adhesion can cause sudden and continuous fluctuations in pixel grayscale values, and such fluctuations will be reflected in the differential time-history signal. Significant peaks or valleys are formed in the differential time history signal. Subsequently, the matching pursuit algorithm is used to decompose the differential time history signal. This algorithm is based on a preset time-frequency dictionary and successively selects the components with the highest matching degree with the time-frequency atoms in the dictionary from the differential time history signal. The iteration continues until the signal residual is less than a preset threshold, usually 5% of the original signal energy. Finally, a set of time-frequency atoms is obtained. This set of atoms can completely characterize the time-frequency features of the differential time history signal, especially the time-frequency features corresponding to the abnormal fluctuations caused by dust. The construction of the time-frequency dictionary combines the characteristics of the monitoring scenario and selects Gaussian wavelet atoms as basic atoms. Atom parameters, such as center frequency, bandwidth, and time width, are then adapted and adjusted according to the frequency range of the differential time history signal, such as usually 1 to 10 Hz.

[0041] Step 32: Calculate the distribution entropy and energy concentration of time-frequency atoms in the time-frequency domain, and fuse them to generate a pixel-level modulation intensity scalar. The specific operations are as follows: For the obtained set of time-frequency atoms, two indices need to be calculated in the time and frequency domains: distribution entropy and energy concentration. These two indices jointly characterize the distribution characteristics and energy concentration of time-frequency atoms in the time-frequency plane, and are ultimately used to fuse and generate a pixel-level modulation intensity scalar. Distribution entropy is used to quantify the degree of disorder or randomness in the energy distribution of time-frequency atoms in the time-frequency plane. Typically, time-frequency atoms corresponding to abnormal fluctuations caused by dust or other interference are more scattered and disordered, so the calculated distribution entropy value is also higher. Energy concentration is used to quantify the degree of concentration of energy of time-frequency atoms in the time-frequency plane. In areas affected by dust, the energy of time-frequency atoms is often more dispersed, so the calculated value of its energy concentration will be lower. Specifically, distribution entropy is calculated by statistically analyzing the probability distribution of time-frequency atoms in the divided time-frequency grid; energy concentration can be measured by calculating the variance of the energy of time-frequency atoms.

[0042] After calculating the distribution entropy and energy concentration separately, these two indicators need to be normalized, adjusting their numerical range to between 0 and 1. Then, through a linear weighted fusion process, these two normalized indicators are combined into a single pixel-level modulation intensity scalar. The logic of this fusion process is that modulation intensity should be positively correlated with energy concentration (i.e., the more concentrated the energy, the higher the modulation intensity) and negatively correlated with distribution entropy (i.e., the more disordered the distribution, the lower the modulation intensity). To unify this relationship in the calculation, the negative correlation of distribution entropy is first converted to a positive correlation, i.e., by subtracting the normalized value from 1. The distribution entropy value, the final modulation intensity scalar, is equal to the normalized energy concentration multiplied by a weighting coefficient, plus the transformed distribution entropy index, i.e., 1 minus the normalized distribution entropy, and then multiplied by another weighting coefficient, which is 1 minus the previous weighting coefficient. The weighting coefficient assigned to the energy concentration is usually between 0.6 and 0.7, determined through sample training. This reflects the relatively dominant role of energy concentration in distinguishing dust areas. The final generated modulation intensity scalar value also ranges from 0 to 1. The higher the value, the more significant the modulation feature caused by abnormal fluctuations in that pixel.

[0043] Step 33: Using the modulation intensity scalar as the anchor point, interpolation is performed under the constraints of the image gradient and confidence map to generate a temporal modulation intensity map. The specific operations are as follows: Using the pixel-level modulation intensity scalar generated in step 32 as the core anchor point, the anchor point is selected from the modulation intensity scalar corresponding to high-confidence pixels. The modulation intensity data of such anchor points has high reliability and can be used as the benchmark for interpolation. Before performing the interpolation operation, the image gradient information of the physical calibration sequence is extracted. The image gradient reflects the spatial change direction and amplitude of pixel grayscale. The gradient characteristics of the weld area and the background area are significantly different. During the interpolation process, the image gradient direction must be followed to ensure that the interpolated data fits the original texture structure of the image. At the same time, the pixel-level motion calibration confidence map is used as a constraint. The higher the confidence level, the higher the interpolation weight. The larger the value, the better the interpolation result fits the modulation intensity characteristics of the reliable region. The interpolation algorithm uses bicubic interpolation, which calculates the modulation intensity scalar of the 16 adjacent anchor points around the point to be interpolated, and combines the image gradient direction and confidence weight to fit the modulation intensity value of the point to be interpolated. Compared with other interpolation algorithms, it can better preserve the image's detailed features and edge information. By performing this interpolation operation on all pixels of the image, the discrete anchor point modulation intensity scalar is expanded into a continuous grayscale image, and finally a temporal modulation intensity map is generated. The grayscale values ​​in the image clearly reflect the differences in signal modulation characteristics of each pixel in the time dimension.

[0044] The generation of the binarized dust mask includes the following steps: Step 34: Based on the temporal modulation intensity map and using pixel-level motion calibration confidence as the adjustment criterion, perform local threshold segmentation to generate a preliminary binary map. The specific operations are as follows: Based on the temporal modulation intensity map generated in step 33, a local threshold segmentation algorithm is used to extract potential dust regions. The segmentation threshold is not a globally fixed value, but is adjusted based on the pixel-level motion calibration confidence level generated in step 22 to achieve adaptive local threshold setting. First, the temporal modulation intensity map is divided into multiple local windows according to a preset size, such as typically 15×15 pixels. For each local window, the mean and standard deviation of the modulation intensity of all pixels within the window are calculated, which serve as the basic threshold reference. Subsequently, the basic threshold is adjusted based on the average confidence level of the pixels within the window. The higher the average confidence level, the better. The higher the reliability of motion calibration of pixels within the window, the more accurate the modulation intensity data, and the threshold can be appropriately lowered to avoid missing small dust areas. The lower the average confidence level, the higher the threshold needs to be to reduce false detections caused by noise interference. The specific adjustment logic of the local threshold is as follows: the base threshold is added with an average confidence correction term, which is negatively correlated with the average confidence level. Based on the adjusted local threshold, the pixels within each local window are binarized. Pixels with modulation intensity higher than the threshold are marked as potential dust pixels, and pixels with modulation intensity lower than the threshold are marked as normal pixels. This operation generates a preliminary binary image.

[0045] Step 35: Perform temporal persistence verification and morphological cleansing on the connected regions in the initial binary graph. The specific operations are as follows: For the generated preliminary binary image, firstly, connected component analysis is performed to extract all connected regions in the image. Each connected region is composed of adjacent potential dust pixels. The connectivity determination adopts the 8-neighborhood criterion, that is, if the adjacent pixels in the 8 directions around a pixel are all potential dust pixels, they are classified into the same connected region. Subsequently, temporal persistence verification is performed on each connected region to extract the position and morphological information of the connected region in the preliminary binary image of consecutive frames, and to count its existence duration and morphological stability in consecutive frames. A persistence criterion is established, typically requiring connected regions to exist stably in the initial binary image for three or more consecutive frames, with a morphological change rate not exceeding 20%. The morphological change rate is determined by calculating the ratio of the overlapping area of ​​connected regions in adjacent frames to the total area. Connected regions that meet this criterion are considered stable regions and are retained; connected regions that do not meet the criterion are considered transient noise regions and are discarded. After temporal persistence verification, morphological cleansing operations are performed on the retained stable connected regions. First, erosion is used to remove small protrusions and isolated pixels at the edges of the connected regions. The erosion operation uses a 3×3 square structuring element, with all pixel values ​​within the structuring element being 1. Then, dilation is used to restore the main outline of the connected regions and fill the small gaps generated during the erosion process. The dilation operation also uses a 3×3 square structuring element. Through the combination of erosion and dilation operations, the morphological optimization of the connected regions is achieved, resulting in a cleaned region mask.

[0046] Step 36: Perform cross-frame fusion on the cleaned region mask, and trim the fusion result based on the confidence map to generate a binarized dust mask. The specific operations are as follows: Collect multiple consecutive frames, typically 5 to 7 frames. After cleaning the region mask in step 35, perform a cross-frame fusion operation. The purpose of fusion is to utilize the redundancy of multi-frame data to further improve the accuracy and completeness of dust region marking. Cross-frame fusion adopts a weighted voting strategy. For each pixel position in the image, count the number of frames in the multiple cleaned masks that it was marked as a dust region. Combine the motion calibration confidence of the corresponding pixels in each frame to set the weight. The higher the confidence, the greater the voting weight. Typically, the weight of a single-frame pixel is the ratio of the pixel's confidence value to the sum of the confidence values ​​of corresponding pixels in all frames. Calculate the weighted voting score of each pixel. When the score exceeds a preset threshold, the pixel is determined to be a dust region pixel. The preset threshold is usually 0.6; otherwise, it is determined to be a normal region pixel. An initial fusion mask is generated through this voting mechanism. Subsequently, the initial fusion mask is trimmed according to the pixel-level motion calibration confidence map generated in step 22. For pixels marked as dust areas in the initial fusion mask, if their motion calibration confidence is lower than the preset threshold, the labeling result is determined to be unreliable and they are corrected to normal region pixels. For pixels marked as normal regions, if their confidence is high and they are not marked as dust areas in multiple surrounding frames, their normal region attribute is confirmed. Through the above cross-frame fusion and confidence trimming operations, a binarized dust mask is finally generated. This mask can accurately mark all invalid dust-covered areas in the physical calibration sequence.

[0047] In some embodiments of the present invention, step 4 is further included: marking invalid data in the physical calibration sequence using a binarized dust mask; reconstructing the texture of the marked region based on the unmarked valid pixel information and the geometric prior of the weld seam by solving the anisotropic diffusion equation; and outputting the texture-reconstructed image sequence. The specific operations are as follows: After obtaining the binarized dust mask generated in step 3, it is matched pixel by pixel with the physical calibration sequence output in step 2. The invalid data areas covered by dust in the physical calibration sequence are marked by the mask, while the unmarked areas are valid data areas that retain complete texture information. In order to restore the weld and background textures in the invalid areas, based on the pixel grayscale information of the valid areas and combined with the geometric prior knowledge of the weld, including the typical width, direction, cross-sectional contour and other inherent features of the weld, an anisotropic diffusion equation adapted to texture reconstruction is constructed. By solving this equation, the texture structure of the valid areas is smoothly transferred to the invalid areas, realizing the accurate reconstruction of the texture in the invalid areas and making up for the information loss caused by dust occlusion. The entire reconstruction process needs to take into account the continuity and realism of the texture, ensuring that the reconstructed texture is naturally connected with the actual shape of the weld and the surrounding background texture, and finally outputting an image sequence with complete texture and no dust interference.

[0048] Step 4 also includes the following sub-steps: Step 41: Calculate the structure tensor field within the effective region of the physical calibration sequence and perform anisotropic smoothing. The specific operations are as follows: First, the effective region in the physical calibration sequence is defined as the unmarked portion of the two-dimensional dust mask. Calculations are performed only on this region. The structure tensor field is used to quantify the local orientation, intensity, and coherence of the image texture, providing a foundation for texture reconstruction. Its calculation relies on the gray-level gradient information of pixels within the effective region. For each pixel within the effective region, a 5×5 pixel neighborhood window is extracted centered on it. The gray-level gradients of each pixel within this window in the x and y directions are calculated, typically using the Sobel operator. Next, a local tensor is constructed through the outer product operation of the gradient vectors, and these tensors within the entire window are then processed. Gaussian weighted summation yields the structure tensor matrix of the center pixel. Specifically, the structure tensor matrix is ​​a 2x2 symmetric matrix whose four elements are obtained by summing the squares or products of the gradient components of each pixel within the window after Gaussian weighting: the top-left element is the weighted sum of the squares of the gradients in the x-direction, the top-right and bottom-left elements are the weighted sum of the products of the gradients in the x and y directions, and the bottom-right element is the weighted sum of the squares of the gradients in the y-direction. Here, the Gaussian weight function acts on each relative coordinate within the neighborhood window, and its standard deviation is usually set to 1.0. The purpose is to make the gradient contribution of the center pixel greater, while smoothing the gradient information of the neighborhood to reduce the influence of noise.

[0049] After the structural tensor field is calculated, it needs to be anisotropically smoothed. This smoothing operation is performed along the texture direction, which aims to maintain the continuity of the texture direction while suppressing the interference of gradient noise on the tensor field, thereby ensuring that the tensor field can stably and accurately reflect the texture structure features of the effective region.

[0050] Step 42: Using the smoothed structure tensor field as the source, perform tensor voting on the binarized dust mask region to generate an inference structure tensor field covering the entire image. The specific operation is as follows: Using the anisotropically smoothed structural tensor field from step 41 as a signal source, a tensor voting operation is performed on the invalid regions marked by the binarized dust mask. This transfers the texture structure information from the valid regions to the invalid regions, ultimately generating an inference structural tensor field covering all pixels in the image, which is the sum of the valid and invalid regions. The logic of tensor voting is to use the structural tensor of each pixel within the valid region as a voting source to transfer texture direction and intensity information to pixels in the surrounding invalid regions. The voting weight decreases as the distance between the voting source and the target pixel increases. Simultaneously, the weight is adjusted based on pixel-level motion calibration confidence; the higher the confidence level, the better. The larger the voting source, the greater its voting weight. During the voting process, each invalid pixel collects tensor information transmitted from the voting sources in its surrounding valid regions. Through tensor superposition and normalization, the inference structure tensor of that pixel is generated. For pixels in the invalid region that are far from the valid region, texture information is gradually transmitted through multiple rounds of iterative voting. The iteration terminates when the change in the tensor field after two adjacent rounds of voting is less than a preset threshold, usually 0.01. After the voting is completed, the structure tensor of the valid region remains unchanged, while the structure tensor of the invalid region is generated by voting fusion, ultimately forming an inference structure tensor field covering the entire image.

[0051] Step 43: Construct an anisotropic diffusion equation based on the inference structure tensor field, and iteratively solve it using the effective region pixel values ​​as boundary conditions to reconstruct the texture base of the invalid region. The specific operations are as follows: Based on the generated inference structure tensor field, an anisotropic diffusion equation adapted to texture reconstruction requirements is constructed. This equation ensures that the direction and intensity of diffusion are entirely constrained by the inference structure tensor field. This design aims to ensure that the diffusion process strictly follows the inferred main texture direction while suppressing diffusion at edges perpendicular to the texture, thus effectively preserving texture details and edge features during reconstruction. The form of this anisotropic diffusion equation draws inspiration from the physical principle of thermal diffusion. The equation shows that the rate of change of the gray value of a point in the image over time during diffusion iterations is equal to the divergence of the product of the diffusion tensor and the gray-level gradient at that point. Here, the diffusion tensor is a key variable. The adaptation parameter is calculated from the eigenvalues ​​of the inference structure tensor field corresponding to that point. Specifically, the magnitude of the diffusion tensor is related to the ratio of the maximum to the minimum eigenvalue of the structure tensor field at that point. The maximum eigenvalue represents the intensity of the main texture direction, while the minimum eigenvalue represents the intensity perpendicular to the texture direction. The diffusion coefficient is designed to be negatively correlated with this ratio. When the ratio is large, i.e., when the texture directionality is strong, diffusion along the texture direction is allowed, while diffusion perpendicular to the texture direction is suppressed. When the ratio is small, i.e., when the texture directionality is weak, isotropic smoothing is preferred. This relationship is usually achieved through an exponential function, the strength of which is controlled by an adjustment coefficient ranging from 0.5 to 1.0.

[0052] In the numerical solution of the equation, the pixel gray values ​​within the previously determined physical effective area are set as fixed boundary conditions and remain unchanged during iteration. The goal of the solution is to update the pixel gray values ​​of the invalid area, i.e. the area to be repaired. The finite difference method is used for iterative calculation: in each iteration, the gray value update amount of each pixel in the invalid area is calculated according to the above diffusion equation, and the gray value is refreshed accordingly. When the average change of the gray values ​​of all pixels in the entire image is less than the preset threshold after two adjacent iterations, the iteration terminates. Finally, the image after the iteration is completed is the result of the preliminary texture reconstruction. The texture of the invalid area is restored and naturally connected with the effective area.

[0053] Step 44: Calculate the dense optical flow field of the current frame and adjacent frames in the effective region. Based on the dense optical flow field, project the texture details of adjacent frames onto the invalid region of the current frame. Combine the pixel-level motion calibration confidence with the closure error of the optical flow projection to fuse the data and output the image sequence after texture reconstruction. The specific operations are as follows: First, the dense optical flow field in the effective region of the current frame and adjacent frames is calculated. This dense optical flow field characterizes the one-to-one correspondence between pixels within the effective region of adjacent frames, reflecting the texture motion trajectory between them. The optical flow field calculation employs the Pyramid LK optical flow algorithm. This algorithm constructs an image pyramid, iteratively optimizing the optical flow vector from the top layer (low resolution) to the bottom layer (high resolution) to ensure the accuracy of the optical flow calculation. During the calculation, pixel-level motion calibration confidence is used as a weight; pixels with higher confidence have a greater optimization weight for their optical flow vectors. Based on the calculated dense optical flow field, the texture details within the effective region of adjacent frames are projected onto the invalid region of the current frame. That is, the effective region pixels in adjacent frames corresponding to the invalid region pixels in the current frame are found based on the optical flow vector, and their grayscale values ​​are extracted as a reference for the texture details of the invalid region pixels in the current frame. Subsequently, combined with the image... The pixel-level motion calibration confidence and optical flow projection closure error are used to optimize the fusion of projected texture details. The optical flow projection closure error is used to quantify the degree of matching between the projected texture and the texture of the effective area around the current frame. The smaller the error, the higher the reliability of the projected texture. The fusion process adopts a weighted average strategy. The final gray value of the invalid area pixels in the current frame is obtained by weighting the gray value of the texture base, the gray value of the projected texture in the previous frame, and the gray value of the projected texture in the next frame. The weights are determined by the pixel-level motion calibration confidence and the optical flow closure error, respectively, to ensure that the fused texture fits the structural constraints and has realistic details. By performing this fusion operation on all invalid area pixels in the current frame and performing edge smoothing on the texture of the whole image, the final output is an image sequence after texture reconstruction. This sequence completely preserves the texture information of the weld and the background without dust interference and texture distortion.

[0054] In some embodiments of the present invention, step 5 is further included: calculating an inter-frame difference map on the image sequence after texture reconstruction; establishing a normal deformation prediction model based on historical data of the global motion flow field; subtracting the model from the inter-frame difference map to obtain an abnormal motion residual map; and segmenting the abnormal motion residual map to obtain a preliminary defect region. The specific operations are as follows: After obtaining the texture reconstruction image sequence output in step 4, the inter-frame difference map between adjacent frames is first calculated. By comparing the gray value differences between adjacent frames pixel by pixel, the gray value change area between frames is highlighted. This area includes both abnormal changes caused by weld defects and normal deformation changes caused by slight vibrations of equipment and minor environmental disturbances during construction. To accurately separate abnormal and normal changes, a normal deformation prediction model is constructed based on the historical global motion flow field data generated in step 1. This model specifically captures the inherent, non-defect-type normal motion deformation patterns during pipeline construction. The gray value change features corresponding to the normal deformation of the current frame are predicted by this model and subtracted from the inter-frame difference map. The remaining gray value change area is the abnormal motion residual map containing only defect information. Subsequently, a segmentation operation is performed on the abnormal motion residual map to filter out the areas that meet the gray value characteristics and spatial distribution characteristics of defects, and finally, the preliminary defect area is obtained.

[0055] The generation of the abnormal motion residual map also includes the following steps: Step 51: Perform spatiotemporal decomposition on the historical global motion flow field sequence, extract spatial modes and their temporal evolution patterns, and synthesize periodic and trend-based predicted deformation fields based on these, then superimpose them to generate the predicted global deformation field. The specific operations are as follows: First, data preparation is required. This involves collecting historical global motion flow field sequences within a preset time period during pipeline construction. The duration should cover at least one complete construction vibration cycle. Typically, flow field data corresponding to 10 to 20 consecutive frames are selected to ensure that the data can fully reflect the periodicity and trend characteristics of normal deformation. At the same time, the historical data needs to be cleaned to remove abnormal flow fields caused by sudden interference. The residual threshold method is usually used for screening, and data with residuals higher than three times the global mean standard deviation are removed to ensure the purity of the input data.

[0056] Subsequently, a spatiotemporal decomposition operation is performed. This step combines principal component analysis (PCA) with temporal trend decomposition. Spatially, the pixel-level motion vectors in each frame's flow field are first flattened to construct a high-dimensional feature matrix. Then, PCA is used to calculate the eigenvalues ​​and eigenvectors of its covariance matrix. Based on the variance contribution rate of the eigenvalues, the top eigenvalues ​​with a cumulative contribution rate of over 95% are selected. A spatial mode, The value is usually between 5 and 8, and can be adjusted according to the data redundancy. Each mode corresponds to a typical normal motion spatial distribution pattern, which can accurately characterize the main spatial features of global motion.

[0057] Next, the temporal dimension pattern is extracted. For each selected spatial mode, its corresponding time coefficient sequence in the historical sequence is extracted, that is, the weight ratio of the mode in each frame. The moving average method is used to decompose the time coefficient sequence, separating it into periodic components and trend components. The periodic components usually correspond to the periodic vibration of construction equipment, which shows the characteristics of fluctuating with time. The trend components correspond to factors such as pipeline construction progress or slow environmental changes, which show the characteristics of linear or slow nonlinear changes with time.

[0058] Finally, the prediction fields are synthesized. Based on the extracted spatial modes and their temporal evolution patterns, periodic prediction fields and trend prediction fields are generated respectively. The periodic prediction field is generated by a linear combination of each spatial mode and its predicted periodic time coefficient. The predicted periodic time coefficient is obtained by fitting historical periodic components to a sine function and extrapolating it to the current frame. The trend prediction field is generated by a linear combination of each spatial mode and its predicted trend time coefficient. The predicted trend time coefficient is obtained by fitting historical trend components to a linear or quadratic function, selecting the best fit, and extrapolating it. The two prediction fields are superimposed to obtain the final global deformation prediction field, which is the complete normal deformation prediction model. The logic of this model is based on the principle of linear superposition of time-varying modes, decomposing normal deformation into independently predictable periodic and trend components. By extrapolating historical patterns, it achieves accurate prediction of normal deformation in the current frame, thereby effectively characterizing the deformation features of non-defect classes.

[0059] Step 52: Utilizing the geometric priors of the weld, anisotropically modulate the predicted global deformation field in the weld region to generate a physically constrained deformation prediction field. The specific operations are as follows: Based on the generated predicted global deformation field, anisotropic modulation is performed using prior geometric knowledge of the weld to generate a physically constrained deformation prediction field. This ensures that the predicted normal deformation conforms to the actual physical structural characteristics of the weld. The prior geometric knowledge of the weld includes its typical width (usually 3 to 10 mm, with the corresponding image pixel size determined by resolution); its orientation (mostly straight lines or regular curves); and its cross-sectional profile (mostly trapezoidal or arc-shaped), among other inherent features. These features determine that the deformation of the weld region has a significant direction dependence, with the deformation degree of freedom along the weld orientation being higher than that perpendicular to the orientation. The anisotropic modulation process uses the geometric profile of the weld region as a constraint. First, the weld region is accurately located using image segmentation technology, and the centerline and orientation vector of the weld are extracted. Based on the orientation vector, a direction weight matrix is ​​constructed for the weld region. For each pixel within the domain, based on its distance from the weld centerline and the directional angle of its location, modulation coefficients are determined along the weld direction and perpendicular to the direction. The modulation coefficient along the direction is larger, typically 0.8 to 1.0, allowing for sufficient representation of normal deformation in that direction. The modulation coefficient perpendicular to the direction is smaller, typically 0.2 to 0.4, suppressing unreasonable deformation predictions in that direction and avoiding exceeding the deformation range allowed by the weld's physical structure. The modulation operation is achieved by multiplying the predicted global deformation field element by element with the direction weight matrix, correcting the deformation prediction results for the weld area. The original predicted global deformation field remains unchanged for non-weld areas. Through this modulation process, the generated physical constraint deformation prediction field is made to better fit the normal deformation law in actual construction, eliminating false deformation predictions that do not conform to the weld's physical structure.

[0060] Step 53: Perform multi-resolution pyramid analysis on the physical constraint deformation prediction field and the inter-frame difference map. Generate anomaly motion residual map by fusing residuals from various scales across scales. The specific operations are as follows: First, a multi-resolution pyramid is constructed for the generated physical constraint deformation prediction field and the inter-frame difference map initially calculated in step 5. The number of pyramid layers is usually set to 4 to 5, with the bottom layer being the original resolution, 1 / 2 resolution, 1 / 4 resolution, and 1 / 8 resolution respectively. If necessary, 1 / 16 resolution can be added. Each layer image is obtained by Gaussian downsampling to ensure that each layer image retains the texture and deformation features of the corresponding scale. For each layer of the pyramid, the residual between the inter-frame difference map of that layer and the gray-level change map corresponding to the physical constraint deformation prediction field is calculated. That is, the residual map of that layer is the gray-level value of the inter-frame difference map minus the gray-level change value corresponding to the predicted deformation, thus obtaining the residual information at each scale. The bottom residual map retains the detailed information of small-sized defects, while the top residual map highlights the overall features of large-sized defects. Subsequently, a cross-scale fusion operation is performed. The fusion process adopts a top-down strategy, starting from the top residual map. It is then amplified to the next layer resolution through bilinear interpolation and fused with the next layer residual map using weighted fusion. The fusion weight is determined based on the inter-layer resolution difference and the residual reliability. The weight of the bottom residual map is usually 0.6 to 0.7, and the weight of the upper interpolated residual map is usually 0.3 to 0.4, ensuring that the fusion preserves both the details of the bottom layer and the global features of the top layer. The fusion operation of each layer residual is completed sequentially until the residual map of the original resolution is obtained. Then, the residual map is subjected to grayscale normalization processing, mapping the grayscale values ​​to the range of 0 to 255, and finally generating an abnormal motion residual map. This map only retains the abnormal grayscale changes caused by defects, and the grayscale changes corresponding to normal deformation have been effectively deducted.

[0061] The initial defect area acquisition process also includes the following steps: Step 54: Determine anomaly seed points based on the anomaly motion residual map, and perform competitive region growing starting from the anomaly seed points to generate candidate defect regions and competition intensity scores. The specific operations are as follows: Based on the generated abnormal motion residual map, abnormal seed points are first determined. Seed points are selected using a residual gray-level thresholding method. The threshold is determined by statistically analyzing the gray-level histogram of the residual map. Typically, the gray value corresponding to a cumulative gray-level frequency of 98% in the gray-level histogram is selected as the seed point threshold. Pixels with residual gray-level values ​​higher than this threshold are considered potential abnormal seed points. To avoid noise interference, a neighborhood verification is performed on potential abnormal seed points. Only potential seed points with at least two pixels in their surrounding 3×3 neighborhood that also have residual gray-level values ​​higher than the threshold are retained as the final abnormal seed points. Subsequently, a competitive region growing operation is performed starting from each abnormal seed point. The growth criterion for region growing is that the residual gray-level values ​​of adjacent pixels are higher than the growth threshold, typically 70% to 80% of the seed point threshold, and are also higher than the current... The difference in the mean grayscale value of the growth area is less than the preset value. The core of competitive growth is that multiple seed points grow simultaneously. When the growth areas of different seed points overlap, the ownership of the overlapping area is determined by the competition intensity score. The competition intensity score comprehensively considers the residual grayscale value of the seed point, the grayscale uniformity and compactness of the growth area. The higher the residual grayscale value, the more uniform the grayscale, and the more compact the area, the higher the score. The competition intensity score is calculated using a weighted summation method. The logic is that the residual grayscale value of the seed point accounts for 60% of the weight, the reciprocal of the standard deviation of the area grayscale, which represents uniformity, accounts for 30% of the weight, and the circularity of the area, which represents compactness, accounts for 10% of the weight. This scoring mechanism ensures that the growth area can accurately fit the actual range of the defect, and finally generates multiple candidate defect areas and corresponding competition intensity scores.

[0062] Step 55: Based on the competition intensity score, and verifying the spatiotemporal coherence of each candidate defect region in the abnormal motion residual map of consecutive frames, a preliminary defect region is obtained through screening. The specific operation is as follows: Based on the generated candidate defect regions and their competition intensity scores, and combined with the spatiotemporal coherence verification of continuous multi-frame abnormal motion residual maps, candidate defect regions are screened to obtain preliminary defect regions. The spatiotemporal coherence verification is conducted from two dimensions: temporal coherence and spatial coherence. Temporal coherence verification requires that the candidate defect region exists stably in continuous multi-frame abnormal motion residual maps, and that the center coordinate offset of the region in each frame is less than a preset threshold, with the region area change being less than 20%, ensuring that the selected region is not a false region caused by instantaneous noise. Spatial coherence verification requires that the shape of the candidate defect region conforms to the typical spatial characteristics of weld defects; for example, crack defects are usually elongated, and porosity defects are often shaped like pores. Defects typically appear circular or elliptical. By calculating morphological parameters such as the aspect ratio and roundness of the region, candidate regions whose shapes do not conform to typical defect characteristics are eliminated. For example, regions with an aspect ratio less than 2 and a roundness less than 0.6, i.e., non-cracks and non-porosity, are eliminated. The screening process first retains the top 30% of candidate defect regions in terms of competitive intensity score, then performs spatiotemporal coherence verification on these regions, eliminating regions that fail the verification. Finally, the regions that pass the verification are processed to smooth the edges, correct the region contour, and ensure that the region boundary fits the actual edge of the defect, thus obtaining the preliminary defect region. This region has effectively eliminated noise and normal deformation interference, accurately capturing the location and range of potential weld defects.

[0063] In some embodiments of the present invention, step 6 is further included: filtering based on the historical distribution of the binary dust mask and the topological relationship between the preliminary defect region and the weld, and outputting a defect marking map. The specific operation is as follows: After acquiring the historical distribution data of the binarized dust mask generated in step 3, as well as continuous multi-frame mask data (usually 5 to 7 frames), and the preliminary defect area obtained in step 5, a final screening process is conducted based on two criteria to eliminate false areas and accurately locate real weld defects. The historical distribution of the binarized dust mask is used to determine whether the preliminary defect area has been subject to long-term dust interference, avoiding misjudging false signals caused by dust obstruction as defects. The topological relationship between the preliminary defect area and the weld is used to verify whether the area is within the effective detection range of the weld and whether it meets the location correlation characteristics between the defect and the weld. By combining the screening results of the two criteria, false areas that do not meet the conditions are eliminated, and real defect areas are retained. Finally, a defect marker map with key defect information is generated, providing intuitive and accurate defect detection results for construction quality control. The entire process must consider both the temporal characteristics of the historical data and the rationality of the spatial topological relationship to ensure the reliability and accuracy of the screening results.

[0064] The generation of the defect marker map also includes the following steps: Step 61: The candidate defect region is comprehensively verified from three dimensions: dust coverage history, geometric morphology and topological relationship, and temporal causal consistency. The specific operation is as follows: From three dimensions—dust coverage history, geometric morphology and topological relationship, and temporal causal consistency—the candidate defect regions generated in step 54, such as the precursors of the initial defect regions, are retained after spatiotemporal coherence verification in step 55. Then, a comprehensive verification is performed to further eliminate false candidate regions. The dust coverage history verification is based on the historical distribution data of the binarized dust mask. The dust coverage duration and coverage ratio of each candidate defect region in multiple consecutive mask frames are statistically analyzed. If a region is covered by dust in more than 60% of the frames, and the coverage ratio exceeds 50% of the region area, the region is considered severely affected by dust interference and the verification fails. If the coverage duration is less than 30% and the coverage ratio is less than 30%, the verification passes, ensuring that the candidate region is not affected by false dust signals. The geometric morphology and topological relationship verification is divided into two parts. The geometric morphology verification continues the morphology judgment logic of step 55. By calculating parameters such as the region's aspect ratio and roundness, it is confirmed that the region conforms to typical weld defects, such as a thin, elongated crack. The morphological characteristics include: aspect ratio ≥ 2; pores are circular or elliptical with a circularity ≥ 0.6; topological relationship verification involves accurately locating the weld centerline and edge contour, calculating the distance from the center of the candidate defect area to the weld centerline, and the degree of overlap between the area and the weld edge. Only candidate areas whose center distance from the weld centerline does not exceed 1.5 times the weld width and whose overlap with the weld area is ≥ 80% are retained to ensure that the area is within the effective detection range of the weld; temporal causal consistency verification is combined with the pipeline construction process, requiring that the appearance and development of candidate defect areas conform to the construction timeline logic. For example, defects should persist after the weld is formed and should not suddenly disappear as construction progresses. At the same time, the morphological changes of the area should conform to the natural evolution law of defects, such as cracks not suddenly widening and pores not suddenly deforming. If there is a temporal logical contradiction in the area, such as appearing before construction and suddenly disappearing during construction, the verification fails. Only candidate areas that pass the verification results in all three dimensions can proceed to the subsequent processing stage.

[0065] Step 62: Integrate the comprehensive verification results with the competition intensity score to generate a defect marker map labeled with defect type and detection confidence level. The specific operation is as follows: In the final stage of defect detection, the comprehensive results of the candidate region across three verification dimensions are first quantified and assigned values: if all three aspects—dust coverage history, geometric morphology and topological relationship, and consistency of temporal factors—pass verification, the value is assigned 1.0; if two aspects pass, the value is assigned 0.6; if only one aspect passes or all aspects fail, the value is assigned 0.0. This value is called the comprehensive verification quantification value.

[0066] Subsequently, this quantified value is integrated with the competition intensity score generated in the previous steps to calculate the final comprehensive evaluation score for each candidate defect region. This score is obtained by weighted summation of two parts: one part is the competition intensity score, which has been normalized to the 0-1 interval, multiplied by a weighting coefficient. The other part is the comprehensive verification quantification value, multiplied by a weighting coefficient. Among them, the weighting coefficient and The scores are determined through sample training to ensure that the two indicators contribute equally to the final score. The comprehensive evaluation score itself also falls between 0 and 1. The logic is that the score is positively correlated with both the competition intensity and the verification result, thus more accurately reflecting the probability that the area is a real defect.

[0067] After calculating the comprehensive evaluation score for each region, the detection confidence level is determined based on the score: regions with a score greater than or equal to 0.8 are marked as high confidence, corresponding to a reliability between 90% and 100%; regions with a score between 0.5 and 0.8 are marked as medium confidence, corresponding to a reliability between 60% and 89%; regions with a score less than 0.5 are marked as low confidence, corresponding to a reliability less than 60%, and these regions will be directly eliminated.

[0068] Meanwhile, the defect type is determined based on the geometric morphological parameters of the region. The specific rules are as follows: when the aspect ratio of the region is greater than or equal to 3 and the roundness is less than 0.4, it is determined to be a crack; when the roundness is greater than or equal to 0.6 and the aspect ratio is less than 2, it is determined to be a pore; and regions with morphological characteristics between the two are determined to be other defects.

[0069] Finally, all the defective areas that pass the screening, i.e. areas with medium or high confidence, along with their corresponding defect types and detection confidence levels, will be superimposed and annotated on the original monitoring images to generate a clear defect marking map. This map intuitively presents the location, extent, category, and confidence level of each defect, providing a direct basis for precise control of construction quality.

[0070] Example 2: Please see Figure 2 Based on Example 1, this embodiment provides a pipeline construction monitoring system based on image recognition, including: The flow field analysis module is used to project a light spot array onto the monitoring scene containing the weld, analyze the displacement of the light spot array in the image sequence to resolve the affine transformation parameters, and generate a global motion flow field. The image calibration module is used to perform inverse compensation on the image sequence using the global motion flow field to generate a physical calibration sequence; The mask creation module is used to perform signal analysis of each pixel in the time dimension on the physical calibration sequence, generate a temporal modulation intensity map, and perform morphological dilation and connected component analysis on the temporal modulation intensity map to generate a binarized dust mask. The texture restoration module is used to mark invalid data in the physical calibration sequence using a binarized dust mask. Based on the valid pixel information of the unmarked area and the geometric prior of the weld, the texture of the marked area is reconstructed by solving the anisotropic diffusion equation, and the image sequence after texture reconstruction is output. The defect screening module calculates the inter-frame difference map on the image sequence after texture reconstruction, establishes a normal deformation prediction model based on historical data of global motion flow field, subtracts the model from the inter-frame difference map to obtain the abnormal motion residual map, and segments the abnormal motion residual map to obtain the preliminary defect area. The defect screening module filters based on the historical distribution of the binary dust mask and the topological relationship between the preliminary defect area and the weld, and outputs a defect marking map.

[0071] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concepts, should be covered within the scope of protection of the present invention.

Claims

1. A pipeline construction monitoring method based on image recognition, characterized in that, The method includes: A light spot array is projected onto the monitoring scene containing the weld seam, and the displacement of the light spot array in the image sequence is analyzed to resolve the affine transformation parameters and generate a global motion flow field. The image sequence is inversely compensated using the global motion flow field to generate a physical calibration sequence; Perform signal analysis on each pixel in the time dimension on the physical calibration sequence to generate a temporal modulation intensity map, and perform morphological dilation and connected component analysis on the temporal modulation intensity map to generate a binarized dust mask. Invalid data in the physical calibration sequence is marked using a binarized dust mask. Based on the valid pixel information of the unmarked area and the geometric prior of the weld, the texture of the marked area is reconstructed by solving the anisotropic diffusion equation, and the image sequence with reconstructed texture is output. The process of reconstructing the texture of the marked area by solving the anisotropic diffusion equation includes: calculating the structure tensor field in the valid area of ​​the physical calibration sequence and performing anisotropic smoothing; using the smoothed structure tensor field as the source, performing tensor voting on the binarized dust mask area to generate an inference structure tensor field covering the entire image; constructing the anisotropic diffusion equation based on the inference structure tensor field, and iteratively solving it with the pixel values ​​of the valid area as boundary conditions to reconstruct the texture base of the invalid area. Frame difference maps are calculated on the image sequence after texture reconstruction. Based on historical data of global motion flow field, a normal deformation prediction model is established. The model is used to predict the gray-level change features corresponding to the normal deformation of the current frame. The predicted gray-level change features are subtracted from the frame difference map to obtain the abnormal motion residual map. The abnormal motion residual map is then segmented to obtain the preliminary defect area. Based on the historical distribution of the binarized dust mask and the topological relationship between the preliminary defect area and the weld, a defect marking map is output.

2. The pipeline construction monitoring method based on image recognition according to claim 1, characterized in that, The generation of the global motion flow field includes: A composite light spot pattern is projected onto the monitoring scene. The composite light spot pattern is composed of a dominant frequency sine stripe and a moiré circular spot. Local micropattern analysis was performed on the moiré spot region, and the absolute phase value of each spot was calculated by two-dimensional phase unwrapping operation; Matching is performed between consecutive frames based on absolute phase values. The search area is predicted based on the motion information of the preceding frame and then matched within the search area. The matching point pairs are then verified based on phase continuity and motion consistency to filter out abnormal matches, thus obtaining the displacement observation set. Multiple motion models are fitted in parallel using a pure displacement observation set. The spatiotemporal distribution characteristics of the residuals of the multiple motion models are calculated. Based on the goodness of fit of each motion model and the spatiotemporal distribution characteristics of the residuals, the final model is selected from the multiple motion models to generate the global motion flow field.

3. The pipeline construction monitoring method based on image recognition according to claim 2, characterized in that, The generation of the physical calibration sequence includes: An adaptive sampling grid is generated based on the model type of the global motion flow field. The grid points are then inversely mapped to the original image and interpolated according to the global motion flow field. Based on the residual distribution map corresponding to the global motion flow field, analyze the residual fluctuation of each position in the time series of the remapped image, calculate the pixel-level motion calibration confidence and mark the absolutely reliable anchor points; Using absolutely reliable anchor points and pixels with confidence levels higher than a first predetermined threshold as boundaries, an anisotropic diffusion process constrained by pixel-level motion calibration confidence is performed to repair regions with confidence levels lower than a second predetermined threshold. The pixel value difference in the overlapping area of ​​pixels with confidence scores higher than a first predetermined threshold in adjacent frames is compared, and temporal smoothing filtering is performed on areas with discontinuous jumps to output a physical calibration sequence.

4. The pipeline construction monitoring method based on image recognition according to claim 3, characterized in that, The generation of the time-domain modulation intensity map includes: The pixel time history signal is preprocessed based on pixel-level motion calibration confidence to extract differential time history signals containing abnormal fluctuations, and a set of time-frequency atoms is obtained by matching and tracking the differential time history signals. Calculate the distribution entropy and energy concentration of time-frequency atoms in the time-frequency domain, and fuse them to generate a pixel-level modulation intensity scalar; Using the modulation intensity scalar as the anchor point, interpolation is performed under the constraints of the image gradient and confidence map to generate a temporal modulation intensity map.

5. The pipeline construction monitoring method based on image recognition according to claim 4, characterized in that, The generation of a binarized dust mask includes: Based on the temporal modulation intensity map, and using pixel-level motion calibration confidence as the adjustment criterion, local threshold segmentation is performed to generate a preliminary binary map; Temporal persistence verification and morphological cleanup are performed on the connected regions in the preliminary binary graph; The cleaned area mask is fused across frames, and the fusion result is trimmed according to the confidence map to generate a binarized dust mask.

6. The pipeline construction monitoring method based on image recognition according to claim 5, characterized in that, The texture of the labeled region is reconstructed by solving the anisotropic diffusion equation, and the output image sequence after texture reconstruction includes: Calculate the dense optical flow field of the current frame and adjacent frames in the effective region. Based on the dense optical flow field, project the texture details of adjacent frames to the invalid region of the current frame. Combine the pixel-level motion calibration confidence and the closure error of the optical flow projection to fuse the image sequence after texture reconstruction.

7. The pipeline construction monitoring method based on image recognition according to claim 6, characterized in that, The generation of anomalous motion residual maps includes: Spatiotemporal decomposition of historical global motion flow field sequences is performed to extract spatial modes and their temporal evolution patterns. Periodic and trend-based predicted deformation fields are synthesized and superimposed to generate a predicted global deformation field. By utilizing the geometric priors of the weld, the predicted global deformation field is anisotropically modulated in the weld region to generate a physically constrained deformation prediction field. Multi-resolution pyramid analysis is performed on the physical constraint deformation prediction field and the inter-frame difference map. By fusing the residuals of each scale across scales, an abnormal motion residual map is generated.

8. The pipeline construction monitoring method based on image recognition according to claim 7, characterized in that, The initial acquisition of defect areas includes: Anomaly seed points are determined based on the abnormal motion residual map, and competitive region growth is performed starting from the anomaly seed points to generate candidate defect regions and competition intensity scores. Based on the competition intensity score, and verifying the spatiotemporal coherence of each candidate defect region in the abnormal motion residual map of multiple consecutive frames, the preliminary defect regions are obtained after screening.

9. The pipeline construction monitoring method based on image recognition according to claim 8, characterized in that, The generation of the defect marker map includes: Candidate defect regions are comprehensively verified from three dimensions: dust coverage history, geometric morphology and topological relationship, and temporal causal consistency. The comprehensive verification results are integrated with the competition intensity score to generate a defect label map labeled with defect type and detection confidence level.

10. A pipeline construction monitoring system based on image recognition, applied to the pipeline construction monitoring method based on image recognition as described in any one of claims 1-9, characterized in that, include: The flow field analysis module is used to project a light spot array onto the monitoring scene containing the weld, analyze the displacement of the light spot array in the image sequence to resolve the affine transformation parameters, and generate a global motion flow field. The image calibration module is used to perform inverse compensation on the image sequence using the global motion flow field to generate a physical calibration sequence; The mask creation module is used to perform signal analysis of each pixel in the time dimension on the physical calibration sequence, generate a temporal modulation intensity map, and perform morphological dilation and connected component analysis on the temporal modulation intensity map to generate a binarized dust mask. The texture restoration module is used to mark invalid data in the physical calibration sequence using a binarized dust mask. Based on the valid pixel information of the unmarked area and the geometric prior of the weld, it reconstructs the texture of the marked area by solving the anisotropic diffusion equation and outputs the texture-reconstructed image sequence. The process of reconstructing the texture of the marked area by solving the anisotropic diffusion equation includes: calculating the structure tensor field within the valid area of ​​the physical calibration sequence and performing anisotropic smoothing; using the smoothed structure tensor field as the source, performing tensor voting on the binarized dust mask area to generate an inference structure tensor field covering the entire image; constructing the anisotropic diffusion equation based on the inference structure tensor field and iteratively solving it with the pixel values ​​of the valid area as boundary conditions to reconstruct the texture base of the invalid area. The defect screening module calculates the inter-frame difference map on the image sequence after texture reconstruction. Based on the historical data of the global motion flow field, it establishes a normal deformation prediction model. It uses this model to predict the gray-level change features corresponding to the normal deformation of the current frame and subtracts the predicted gray-level change features from the inter-frame difference map to obtain the abnormal motion residual map. The abnormal motion residual map is then segmented to obtain the preliminary defect area. The defect screening module filters based on the historical distribution of the binary dust mask and the topological relationship between the preliminary defect area and the weld, and outputs a defect marking map.