Crack image processing method based on neural network and plane geometry image transformation

By constructing a joint model that integrates image anchor points, viewing angle estimation and image texture information, the difficulty in crack identification caused by image geometric distortion in the prior art is solved, and high-precision three-dimensional crack positioning and structural recovery are achieved.

CN120047755AActive Publication Date: 2025-05-27UNIV OF JINAN +1

Patent Information

Application Number
CN202510520702.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the case of viewing angle transformation disturbances, target surface reflection interference and image geometric distortions in the image acquisition process, it is difficult to achieve the correspondence between the crack image and its three-dimensional geometric structure, which limits the ability to identify and spatial structure reduction.

Method used

By constructing a joint model of image geometric correction and three-dimensional backprojection that integrates image anchor points, viewing angle estimation and image texture information, the geometric structure correction of image frames and the position recovery of cracks in three-dimensional space are achieved.

Benefits of technology

The accuracy of crack area positioning under complex viewing angle disturbance is improved, structural consistency and texture continuity before image analysis is achieved, and subsequent three-dimensional structural damage modeling and quantitative analysis are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047755A_ABST
    Figure CN120047755A_ABST
Patent Text Reader

Abstract

The invention discloses a crack image processing method based on a neural network and plane geometry image transformation, and particularly relates to the field of general image data processing driven by plane geometry image transformation and neural network identification. Comprising the steps that an image collection platform projects a structured image coding pattern to the surface of a to-be-measured target through a preset image coding pattern template, and synchronously collects an image frame containing an image mark so as to construct a space anchor point used for image geometric reference in an image plane; view angle parameter change information in an image acquisition process is acquired through an image acquisition platform, and projection configuration parameters of an image coding pattern are combined. By constructing an image offset modeling mechanism jointly driven by an image coding pattern and a space anchor point, image in-plane correction of geometric distortion features in an image frame is realized, and structural consistency and texture continuity before subsequent image analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of general image data processing driven by plane geometric image transformation and neural network recognition. More specifically, the present invention relates to a method for processing crack images based on a neural network and plane geometric image transformation. Background Art

[0002] The identification and processing of cracks refer to using an image acquisition device to obtain image frame data of the surface of a target structure, and combining an image data processing algorithm or a deep learning model based on a neural network to automatically detect and spatially locate potential crack regions in the image, and then extract image structure features such as crack boundaries and centerlines to provide data support for subsequent structural safety assessment and repair.

[0003] The existing methods are traditional processing methods based on two-dimensional images or shallow neural network models. When facing complex factors such as perspective transformation disturbances, target surface specular reflections, and image geometric distortions during the image acquisition process, the projection of the crack region in the image is prone to misalignment, blurring, and even structural fracture, making it difficult to establish the correspondence between the crack image and its three-dimensional geometric structure, which limits the recognition and spatial structure restoration capabilities. Summary of the Invention

[0004] To overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for processing crack images based on a neural network and plane geometric image transformation. By constructing a joint model of image geometric correction and three-dimensional backprojection that integrates image anchor points, perspective estimation, and image texture information, the problems proposed in the above background art are solved.

[0005] To achieve the above object, the present invention provides the following technical solution: A method for processing crack images based on a neural network and plane geometric image transformation, including an image acquisition platform; S1. The image acquisition platform projects a structured image coding pattern onto the surface of the target to be measured through a preset image coding pattern template, and synchronously acquires an image frame containing image markers to construct spatial anchor points for image geometric reference in the image plane; S2. Obtain the perspective parameter change information during the image acquisition process through the image acquisition platform, and combine the projection configuration parameters of the image coding pattern to calculate the theoretical position distribution of the image markers in the image frame to establish a set of predicted positions of the anchor points in the image plane; S3. Extract and decode the actual image markers in the image frame, and register them with the predicted position points to generate a set of spatial offset vectors for describing the geometric distortion characteristics of the image; S4. Perform geometric image transformation in the image plane, construct a non-linear inverse transformation function of the image frame based on the set of spatial offset vectors, and perform pixel-level remapping on the image frame to generate a sequence of image frames with corrected geometric structures; S5. Jointly encode the image frame with corrected geometric structure, the set of image coordinates of the image markers, and the perspective estimation vector of the image acquisition platform into a multi-modal input tensor, and input it into the deep recognition network to extract the set of image coordinates of the crack area; S6. Use the set of image coordinates of the crack area, the imaging parameters corresponding to the image frame, and the perspective estimation vector to construct a spatial back-projection function, restore the actual position of the crack in the three-dimensional space, and generate the corresponding structured geometric expression and three-dimensional view spectrum.

[0006] In a preferred embodiment, S1 further includes: The image acquisition platform generates a structured image coding pattern through a preset image coding pattern template. The pattern types of the image coding pattern include a structural dot matrix pattern, a periodic stripe pattern, or a De Bruijn sequence pattern; the generation of the image coding pattern is based on the image coding parameters set by the image acquisition platform, and the image coding parameters include the image unit distribution density, the inter-unit spacing, and the redundancy and fault tolerance configuration; The image coding pattern template projects the structured image coding pattern onto the surface of the target to be measured according to the image coding parameters, and the image coding pattern and the target surface form a spatially distributed image marker, serving as a spatial anchor for image geometry reference; The image acquisition platform synchronously acquires an image frame containing the image marker through the mounted image acquisition device. The image frame is attached with the image coding parameters and the acquisition timestamp, constituting the image data required for image geometry reference; S2 further includes: The image acquisition platform obtains the perspective parameter change information during its image acquisition process. The perspective parameter change information includes accelerometer data, gyroscope data, and positioning data; by inputting the perspective parameter change information into the perspective solution algorithm, the perspective estimation vector corresponding to the current image frame is output; The perspective estimation vector and the projection configuration parameters of the image coding pattern are jointly input into the projection calculation module of the image acquisition platform to calculate the theoretical position distribution of the image markers in the image frame; the projection configuration parameters include the projection angle, the optical axis direction, and the pattern coverage range; Based on the theoretical position distribution, a set of predicted positions of the image markers is formed, and the set of predicted positions is used to define the search range of the spatial anchor points in the image frame.

[0007] In a preferred embodiment, S3 further includes: Input the image frame into the image preprocessing module mounted on the image acquisition platform, perform preprocessing operations such as gray normalization, edge enhancement, and noise suppression, and extract the set of image coordinates of the actual image markers in the preprocessed image frame; Perform image decoding on the set of image coordinates to identify the identity codes of each image marker, register the set of image coordinates of the actual image marker with the corresponding predicted position points in the theoretical position distribution, generate a set of spatial offset vectors of the image marker in the image frame based on the registration result, and the set of spatial offset vectors is used to characterize the geometric distortion characteristics of the image frame, providing an input basis for image inverse transformation; S4 further includes: Input the set of spatial offset vectors into the image geometric modeling module carried by the image acquisition platform to construct a non-linear inverse transformation function of the image frame, where the non-linear inverse transformation function is composed of an affine mapping function, a periodic perturbation function, and a high-order residual function, and fit the geometric distortion characteristics of the image frame through the non-linear inverse transformation function; Input the original image frame into the non-linear inverse transformation function to perform pixel-level remapping and complete the geometric structure correction of the image frame; the texture information of the image is kept intact during the correction process, and the geometric distortion error caused by the perspective transformation of the image acquisition platform is eliminated; All image frames with corrected geometric structures, their corresponding timestamps, and perspective estimation vectors are combined into a sequence of corrected image frames.

[0008] In a preferred embodiment, S5 further includes: Jointly encode the image frames with corrected geometric structures, the set of image coordinates of the image marker, and the perspective estimation vector into a multi-modal input tensor; Input the multi-modal input tensor into the deep recognition network to perform multi-channel feature encoding and spatial alignment processing. The multi-channel feature encoding and spatial alignment processing fuse the feature information of the image texture channel, the image marker coordinate channel, and the perspective estimation channel to perform spatial feature alignment and consistency modeling of crack region localization; After completing the multi-channel feature encoding and spatial alignment processing, the deep recognition network outputs an image mask of the crack region, a set of image coordinates of the crack centerline, and a set of image coordinates of the crack boundary; extract the set of image coordinates of the crack region based on the image mask as the spatial alignment result obtained after fusing the image texture channel, the image marker coordinate channel, and the perspective estimation channel, and input the set of image coordinates into the subsequent spatial back-projection function to construct a three-dimensional positioning relationship; S6 further includes: Use the set of image coordinates of the crack region, the imaging parameters of the image acquisition device carried by the image acquisition platform corresponding to the image frame, and the perspective estimation vector to construct a spatial back-projection function, and input the set of image coordinates of the crack region into the spatial back-projection function to restore the actual position of the crack in three-dimensional space; Extract the spatial geometric attributes of the crack based on its three-dimensional spatial position. The spatial geometric attributes include crack curvature, crack direction, and crack length. Write the three-dimensional spatial position and spatial geometric attributes of the crack, as well as the image frame index and recognition confidence information corresponding to the image frame, into the structured geometric expression recording unit of the image acquisition platform. Generate a three-dimensional visual spectrum of the crack based on the structured geometric expression, and output a visualization result file for inspection and diagnosis.

[0009] In the formula structure involved in this solution, the dimensionless term can be used as a proportional or structural adjustment factor. When combined with a quantity with a unit, it only plays a role in numerical scaling and does not introduce a new physical dimension. Therefore, it will not change or confuse the overall unit system of the expression. Such a combination of "dimensionless term and quantity with a unit" can be understood as a composite structure expression form commonly used in mathematical and physical modeling, which conforms to the principle of dimensional consistency and has a clear physical interpretation basis. Secondly, in the formula structure of this solution, if there are multiple variable terms with different physical units, including but not limited to time, mass, or energy variables, their combined appearance is to express the co-modeling relationship of multiple physical mechanisms. Each variable forms a unified structure through function mapping, ratio combination, or normalization adjustment, with clear units and definite meanings. The overall expression conforms to the principle of dimensional consistency and the common norms of engineering modeling. In a preferred embodiment, construct a non-linear inverse transformation function of the image frame by introducing a horizontal inverse transformation function and a vertical inverse transformation function. The horizontal inverse transformation function is expressed as: ; The vertical inverse transformation function is expressed as: ; Where is the horizontal remapping coordinate of the pixel point in the image frame after non-linear inverse transformation; is the vertical remapping coordinate of the pixel point in the image frame after non-linear inverse transformation; represents the horizontal and vertical coordinates of the pixel to be inverse-transformed in the original image frame; is the input horizontal coordinate for the scaling or rotation transformation coefficient of the horizontal output coordinate; is the input vertical coordinate for the shear or rotation coupling coefficient of the horizontal output coordinate; is the input horizontal coordinate for the shear or rotation coupling coefficient of the vertical output coordinate; is the input vertical coordinate Scaling or rotation transformation coefficient for the vertical output coordinate; The translation component that controls the horizontal output in the affine transformation; The translation component that controls the vertical output in the affine transformation; Represents the actual offset of the image marker in the horizontal direction in the image frame; Represents the actual offset of the image marker in the vertical direction in the image frame; , , Respectively represent the estimated vector components corresponding to the pitch angle, roll angle, and yaw angle of the image acquisition platform when acquiring the current image frame.

[0010] In a preferred embodiment, a multi-channel attention fusion tensor is constructed , through Represents the spatial alignment feature tensor after fusing the image texture channel, the image marker coordinate channel, and the viewing angle estimation channel; ; Among them Represents the image texture features extracted from the image frame whose geometric structure has been corrected; The structural feature channel constructed from the set of image marker coordinates; Represents the viewing angle estimation channel constructed from the viewing angle estimation vector ; among which each channel feature is fused after being gated by the viewing angle-driven weight ; Among them Represents The individual component in Is non-linear compression; Based on Construct the coordinate residual alignment kernel function; ; Among them Represents the coordinate residual alignment kernel function; Is the set of crack centerline image coordinates output by the deep network; Is the set of crack boundary image coordinates output by the deep network; Is the position of the previous image marker in the image; The function for extracting the set of crack region image coordinates under mask guidance is expressed as: ; Among them Is the crack region image mask output by the deep network; Is used to construct the periodic discrimination field; the response region is screened by mask weighting, and finally the set of image coordinates of the crack region is output .

[0011] In a preferred embodiment, according to the multi-channel fusion feature tensor a crack image mask is constructed , and is jointly modeled by a view-driven attention activation term and a boundary-sensitive inverse response term; ; wherein is a feature coupling adjustment factor; is a non-linear response compression factor.

[0012] In a preferred embodiment, when constructing the spatial back-projection function , the set of image coordinates of the crack region extracted in the image domain , combined with the imaging parameters of the image acquisition device and the view estimation vector of the image acquisition platform , is back-projected into the three-dimensional world coordinate system to obtain the actual spatial position of the crack ; based on three spatial geometric attribute extraction functions are constructed, which are respectively used to calculate the direction angle of the crack , the curvature of the crack , and the spatial length of the crack , providing an input basis for structured recording and three-dimensional visualization atlas; ; ; ; ; wherein represents the three-dimensional spatial coordinate point corresponding to the th pixel in the crack region; , represents the lateral difference component of the th spatial difference vector in the three-dimensional space; represents the longitudinal difference component of the th spatial difference vector in the three-dimensional space; represents the height difference component of the th spatial difference vector in the three-dimensional space; represents the th image coordinate extracted from the crack mask in the image frame; represents the component in the height direction of the three-dimensional spatial coordinate; represents the arc length parameter along the crack path; wherein represents the actual length of the entire crack in the three-dimensional space; represents the number of all back-projected points in the set of image coordinates of the crack region.

[0013] Technical effects and advantages of the present invention: By constructing an image offset modeling mechanism jointly driven by an image coding pattern and a spatial anchor point, in-plane correction of geometric distortion features in an image frame is achieved, improving the structural consistency and texture continuity before subsequent image analysis; Fusing image texture features, image marker coordinates, and perspective estimation vector coding into a multi-channel attention tensor to drive a deep neural network to perform feature alignment and robust recognition, thereby improving the crack region localization accuracy under complex perspective perturbations; Constructing an image alignment kernel function based on coordinate residuals and fused feature responses to achieve spatial consistency alignment of crack centerlines and boundary lines, effectively suppressing image boundary fracture problems caused by projection offset or depth changes; Introducing a dynamic perspective regulation mechanism, adaptively adjusting the response intensity of the perturbation term in combination with the steady-state information during the image acquisition process, reducing the error transmission effect of image deformation during geometric transformation and recognition; By constructing an image space back-projection function and a crack geometric attribute extraction function, the three-dimensional spatial pose of the crack is restored on the basis of image plane processing, and its structural parameters such as direction, curvature, and length are extracted to support subsequent three-dimensional structural damage modeling and quantitative analysis. Description of the Drawings

[0014] Figure 1 It is a schematic diagram of the method steps of the present invention. Detailed Embodiments

[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0016] Refer to the attached drawings of the specification Figure 1 , a crack image processing method based on a neural network and planar geometric image transformation according to an embodiment of the present invention, includes: an image acquisition platform; S1. The image acquisition platform projects a structured image coding pattern onto the surface of the target to be measured through a preset image coding pattern template, and synchronously acquires an image frame containing image markers to construct a spatial anchor point for image geometric reference; S2. Obtain the perspective parameter change information during the image acquisition process through the image acquisition platform, and calculate the theoretical position distribution of the image markers in the image frame in combination with the projection configuration parameters of the image coding pattern to establish a set of predicted positions of the anchor points; S3. By extracting and decoding the actual image markers in the image frame and registering them with the corresponding predicted position points in the theoretical position distribution, a set of spatial offset vectors of the image markers in the image frame is generated to characterize the image distortion features; S4. Based on the set of spatial offset vectors, a non-linear inverse transformation function of the image frame is constructed, and pixel-level remapping is performed on the image frame to generate a sequence of image frames with corrected geometric structures; S5. The image frames with corrected geometric structures, the set of image coordinates of the image markers, and the perspective estimation vector of the image acquisition platform are jointly encoded into a multi-modal input tensor and input into a deep recognition network to extract the set of image coordinates of the crack region; S6. Using the set of image coordinates of the crack region, the imaging parameters corresponding to the image frame, and the perspective estimation vector, a spatial back-projection function is constructed to restore the actual position of the crack in three-dimensional space, and a corresponding structured geometric expression and three-dimensional view spectrum are generated.

[0017] In the formula structure involved in this solution, dimensionless terms can serve as proportional or structural adjustment factors. When combined with quantities with units, they only play a role in numerical scaling and do not introduce new physical dimensions, so they do not change or confuse the overall unit system of the expression; such combinations of "dimensionless terms and terms with unit quantities" can be understood as the composite structure expression forms commonly used in mathematical and physical modeling, which conform to the principle of dimensional consistency and have a clear physical interpretation basis; Secondly, in the formula structure of this solution, if there are multiple variable terms with different physical units, including but not limited to time-type, mass-type, or energy-type variables, their joint appearance is to express the collaborative modeling relationship of multiple physical mechanisms. Each variable forms a unified structure through function mapping, ratio combination, or normalization adjustment, with clear units and clear meanings, and the overall expression conforms to the principle of dimensional consistency and the common norms of engineering modeling; For the above solution, it should be further supplemented and explained that: S1 further includes: The image acquisition platform generates a structured image coding pattern through a preset image coding pattern template. The pattern types of the image coding pattern include a structured dot matrix pattern, a periodic stripe pattern, or a DeBruijn sequence pattern; the generation of the image coding pattern is based on the image coding parameters set by the image acquisition platform, and the image coding parameters include the distribution density of image units, the spacing between units, and the redundant fault tolerance configuration; The image coding pattern template projects the structured image coding pattern onto the surface of the target to be measured according to the image coding parameters, and the image coding pattern and the target surface form a spatially distributed image marker, which serves as a spatial anchor for image geometry reference; The image acquisition platform synchronously acquires image frames containing image markers through the mounted image acquisition device. The image frames are attached with image coding parameters and acquisition timestamps, constituting the image data required for image geometric reference; Regarding S1, it should be further explained that according to the pattern type selection parameter set by the image acquisition platform, different structured image coding pattern prototypes are correspondingly generated. Each type of pattern corresponds to a different function generation mechanism, including a structure dot matrix function, a periodic stripe function, and a De Bruijn coding function, which are used for subsequent projection and registration in practical applications; ; In the above formula, is the image coding pattern prototype function generated in the two-dimensional space, and its output is used to construct the projection pattern; , is the virtual coordinate variable during the pattern generation process, which refers to the coordinate that has not been mapped to the target surface; is the image unit distribution density; is the spacing between units; is the redundant perturbation variable, and the redundant fault tolerance configuration is described by ; is the pattern type control variable, which is an integer variable used to select the pattern type: when represents the structure dot matrix, when represents the periodic stripe, and when represents the De Bruijn sequence coding; in addition, in the formula represents the floor symbol; represents the modulo operation, that is, taking the remainder after dividing the value in the parentheses by 3, which is used to construct the periodic coding logic; Based on S1, construct the structured image coding pattern function , which is used to introduce the geometric parameter regulation and redundancy enhancement strategy of the overall image structure on the basis of the selected pattern type, and construct a complete projectable image coding pattern, which includes a periodic structure, spatial variation, and fault tolerance ability; ; Among them is the pattern geometric structure strength, controlling the density through λ and the relative spacing through ξ; is the redundant fault tolerance function, which is used to insert redundant marker points or local perturbations in the coding pattern, and η controls the fault tolerance amplitude; Construct the spatial coordinate perturbation function after the pattern is projected onto the target surface, which is used to project the generated image coding pattern coordinates from the pattern plane into the target surface coordinate system, and add a perturbation factor to simulate the distortion effects brought by distance, angle, and viewing angle during the actual projection process; ; where is the actual projection coordinate formed by the image coding pattern on the target surface; is the pattern projection angle parameter; is the optical axis direction parameter; is used to simulate the non - linear perturbation caused by the local high - frequency fluctuation of the target surface during the projection process, reflecting the periodic distortion in the direction of the combined coordinate; is used to simulate the geometric deviation generated in the orthogonal axis direction, manifested as the pattern deformation caused by directional reflection or refraction; Then, according to the projection coordinates, the pattern function is mapped to obtain the image marker map value function on the target surface: ; where is the image marker intensity function value on the target surface, used for subsequent recognition; is the structured image coding pattern function; When the acquisition device obtains an image frame, it structurally records the image marker content, pattern parameters, pattern type, and acquisition timestamp of the image frame for subsequent coordinate decoding and view angle calculation, which is expressed as: ; where is the image frame data structure; is the image marker image data; is the image frame acquisition timestamp; is the transpose symbol; S2 also includes: The image acquisition platform obtains the view angle parameter change information during its image acquisition process. The view angle parameter change information includes accelerometer data, gyroscope data, and positioning data; by inputting the view angle parameter change information into the view angle calculation algorithm, the view angle estimation vector corresponding to the current image frame is output; The view angle estimation vector and the projection configuration parameters of the image coding pattern are jointly input into the projection calculation module of the image acquisition platform to calculate the theoretical position distribution of the image marker in the image frame; the projection configuration parameters include projection angle, optical axis direction, and pattern coverage; Based on the theoretical position distribution, a prediction position set of the image marker is formed, and the prediction position set is used to limit the search range of the spatial anchor points in the image frame; For S2, it should be further noted that according to the three types of view angle parameter change information collected in real - time during the UAV image acquisition process: accelerometer data, gyroscope data, and positioning data, through the composite mapping of the exponential function and the arctangent function, a view angle estimation vector for spatially describing the current state of the image frame is generated; the view angle estimation vector is used to correct the projection angle and optical axis offset during the projection process to achieve theoretical consistency in the position of the image marker in the image frame; ; where is the perspective estimation vector, which is used to describe the spatial perspective of the image tilt at the current acquisition moment; are the observed values of the accelerometer in three axes; are the angular velocity observed values of the gyroscope in three axes; are the three-dimensional spatial position or velocity observed values provided by the navigation system; is used to correct the image projection angle; is used to correct the image optical axis direction; is used to correct the pattern coverage perturbation; is the base of the natural logarithm; is the arctangent function; The above formula is used to calculate the spatial theoretical position distribution of the image coding pattern based on the known Ψ and the projection configuration parameters of the image coding pattern, that is, the coordinate positions where the pattern should appear in the image frame in the ideal state; ; where is the virtual point in the two-dimensional image coding pattern is the theoretical projection coordinate in the image frame; is the projection angle of the image coding pattern; is the pattern optical axis direction; is the pattern coverage range factor, which is used to control the perturbation amplitude; , are respectively used to introduce the periodic terms of asymmetric perturbation to simulate pattern deformation; For each coding point in the image coding atlas perform theoretical position projection to form a predicted position set; this set will be used as the range for subsequent anchor point search and perform registration with the actual image markers; ; where is the image marker predicted position set; is the integer set.

[0018] S3 also includes: Input the image frame into the image preprocessing module carried by the image acquisition platform, perform preprocessing operations such as gray normalization, edge enhancement, and noise suppression, and extract the image coordinate set of the actual image markers in the preprocessed image frame; Perform image decoding on the set of image coordinates to identify the identity codes of each image marker, register the set of image coordinates of the actual image marker with the predicted position points corresponding in the theoretical position distribution, generate a set of spatial offset vectors of the image marker in the image frame based on the registration result, and the set of spatial offset vectors is used to characterize the geometric distortion features of the image frame and provide an input basis for image inverse transformation; Formulate Denote the set of image marker coordinates detected in the image; ; where Denote the horizontal coordinate of the marker in the image; Is the vertical coordinate of the marker in the image; the set of theoretical predicted coordinates of the image marker is ; Is the horizontal coordinate of the theoretical predicted position; Is the vertical coordinate of the theoretical predicted position; By taking the And Perform element-wise difference to construct a set of spatial offset vectors, which is used to characterize the geometric distortion characteristics of the image frame and serves as the input for subsequent construction of the image inverse transformation function; ; Where Is the set of spatial offset vectors; each dimension represents the difference between the actual image coordinates and the predicted coordinates, quantifying the image distortion; Based on this, construct an image geometric distortion feature function and formulate the geometric distortion function value of the image frame , Fuses the spatial offset intensity and the image coordinate position information for use in the next stage of constructing the image inverse transformation function; ; Where Is the distortion feature value of the image frame; the first input part in the formula is the Euclidean norm of the offset vector, used to measure the offset intensity; the second input part in the formula Represents the sine modulation factor in the horizontal direction of the current position; the third input part in the formula Represents the cosine modulation factor in the vertical direction of the current position; their combination characterizes the non-uniform distortion trend of the image at different spatial positions; S4 further includes: Input the set of spatial offset vectors into the image geometric modeling module carried by the image acquisition platform to construct a non-linear inverse transformation function of the image frame, where the non-linear inverse transformation function is composed of an affine mapping function, a periodic perturbation function, and a high-order residual function, and fit the geometric distortion features of the image frame through the non-linear inverse transformation function; Input the original image frame into the non - linear inverse transformation function, perform pixel - level remapping, and complete the geometric structure correction of the image frame; during the correction process, the integrity of the image texture information is maintained, and the geometric distortion error of the image caused by the perspective transformation of the image acquisition platform is eliminated; All the image frames with corrected geometric structures, their corresponding timestamps, and the perspective estimation vectors are combined into a sequence of corrected image frames.

[0019] S5 also includes: Jointly encode the image frames with corrected geometric structures, the set of image coordinates of the image markers, and the perspective estimation vectors into a multi - modal input tensor; Input the multi - modal input tensor into the depth recognition network, perform multi - channel feature encoding and spatial alignment processing. The multi - channel feature encoding and spatial alignment processing fuse the feature information of the image texture channel, the image marker coordinate channel, and the perspective estimation channel, and perform spatial feature alignment and consistency modeling for crack region localization; After completing the multi - channel feature encoding and spatial alignment processing, the depth recognition network outputs the image mask of the crack region, the set of image coordinates of the crack centerline, and the set of image coordinates of the crack boundary; extract the set of image coordinates of the crack region based on the image mask as the spatial alignment result obtained after fusing the image texture channel, the image marker coordinate channel, and the perspective estimation channel, and input this set of image coordinates into the subsequent spatial back - projection function for constructing a three - dimensional positioning relationship; S6 also includes: Use the set of image coordinates of the crack region, the imaging parameters of the image acquisition device carried by the image acquisition platform corresponding to the image frame, and the perspective estimation vector to construct a spatial back - projection function, input the set of image coordinates of the crack region into the spatial back - projection function, and restore the actual position of the crack in three - dimensional space; Extract the spatial geometric attributes of the crack based on the three - dimensional spatial position of the crack. The spatial geometric attributes include crack curvature, crack direction, and crack length; write the three - dimensional spatial position and spatial geometric attributes of the crack, as well as the image frame index and recognition confidence information corresponding to the image frame, into the structured geometric expression recording unit of the image acquisition platform; Generate a three - dimensional visual spectrum of the crack based on the structured geometric expression, and output a visualization result file for inspection and diagnosis.

[0020] Construct the non - linear inverse transformation function of the image frame by introducing the horizontal inverse transformation function and the vertical inverse transformation function to eliminate the geometric distortion caused by spatial offset, periodic perturbation, non - linear residual, local modulation, and perspective change, and achieve pixel - level image correction; The horizontal inverse transformation function is expressed as: ; The vertical inverse transformation function is expressed as: ; where is the horizontal remapping coordinate of the pixel point in the image frame after the non - linear inverse transformation; is the vertical remapping coordinate of the pixel point in the image frame after the non - linear inverse transformation; represents the horizontal and vertical coordinates of the pixel to be inverse - transformed in the original image frame; is the sum of the squares of the pixel coordinates, used to represent the spatial amplitude and construct the position modulation term; is the cube of the horizontal coordinate, used to model the high - order non - linear residual perturbation in the horizontal direction; is the cube of the vertical coordinate, used to model the high - order non - linear residual perturbation in the vertical direction; is the sum of the fourth - powers of the pixel coordinates, used to enhance the scale response and perturbation amplification in the image edge region; is the input horizontal coordinate the scaling or rotation transformation coefficient for the horizontal output coordinate; is the input vertical coordinate the shear or rotation coupling coefficient for the horizontal output coordinate; is the input horizontal coordinate the shear or rotation coupling coefficient for the vertical output coordinate; is the input vertical coordinate the scaling or rotation transformation coefficient for the vertical output coordinate; is the translation component that controls the horizontal output in the affine transformation; is the translation component that controls the vertical output in the affine transformation; As the spatial offset vector term in the formula includes: represents the real offset in the horizontal direction of the image marker in the image frame; represents the real offset in the vertical direction of the image marker in the image frame; represents the non - linear high - order modulation of the pixel position based on the spatial offset; As the periodic perturbation function term in the formula includes: represents the periodic perturbation that constitutes the frequency parameter with the spatial offset, used to model device jitter or target surface reflection interference; is the composite periodic term, and the composite periodic term is used to constitute the two - dimensional spatial rotation perturbation kernel; As the high - order residual function term in the formula includes: is a non - linear activation mapping, which is used to control the convergence of distortion in the edge region of the image; represents the attenuation control based on the spatial offset energy, which is used to prevent the local perturbation from being too strong; represents amplifying the response of the image edge region, that is, modeling the distortion caused by surface curvature; As the spatial modulation kernel term in the formula includes: represents simulating the response difference of different regions in the image to the distortion, which is used to form a non - uniform perturbation kernel; is used to reflect the structural correction requirement when there is symmetric anti - perturbation in the image; The perspective regulation weight term in the formula includes: , , respectively represent the estimated vector components corresponding to the pitch angle, roll angle, and yaw angle of the image acquisition platform during the acquisition of the current image frame; is the perspective regulation weight coefficient for dynamically adjusting the action intensity of the perturbation function; where represents the perturbation adjustment factor constructed based on the perspective estimation vector When the perspective change is stable, that is, Ψ is close to 0, is close to 1, and the action of the perturbation function is enhanced; when the perspective fluctuates violently, decreases, which is used to suppress the distortion response intensity and reduce the introduction of geometric errors.

[0021] Construct a multi - channel attention fusion tensor , through represents the spatial alignment feature tensor after fusing the image texture channel, image marker coordinate channel, and perspective estimation channel, which is used to express the spatial consistency features of the crack region extracted by the depth recognition network under multi - modal input, and provides fusion feature support for subsequent mask generation, coordinate residual modeling, and region localization; ; where represents the image texture features extracted from the image frame with the geometric structure corrected; The structural feature channel constructed from the set of image marker coordinates; represents the perspective estimation channel constructed from the perspective estimation vector ; where each channel feature is fused after passing through the perspective - driven weight gating ; where represents a single component in is used as a regulation variable for the channel attention weight; when is relatively small, it indicates that the perspective transformation is stable, and this weight approaches 1, indicating that the features of this channel maintain a high confidence level in participating in the fusion; when is relatively large, it indicates that the perspective transformation changes drastically, and this weight approaches 0, indicating that the features of this channel reduce their influence in the fusion to suppress the noise caused by perspective jitter; is non-linear compression, and non-linear compression is used to control the degree of fusion activation to generate the final spatially aligned feature tensor ; ; Based on construct a coordinate residual alignment kernel function; ; Among them, represents the coordinate residual alignment kernel function; is the set of crack centerline image coordinates output by the deep network; is the set of crack boundary image coordinates output by the deep network; is the position of the previous image marker in the image, used as a spatial anchor point; represents the coordinate difference between the crack centerline and the image marker; represents the coordinate difference between the crack boundary and the image marker; and and the fusion feature energy term together constitute , which is used for subsequent region determination; The function for extracting the set of crack region image coordinates under the mask guidance is expressed as: ; Among them, is the crack region image mask output by the deep network; in the in the formula, is used to construct a periodic discrimination field to make the high-confidence regions form significant responses; by screening the response regions with mask weighting, finally output the set of image coordinates of the crack region , which is used as the input of the three-dimensional space back-projection function in the next stage.

[0022] According to the multi-channel fusion feature tensor construct a crack image mask , through the joint modeling of the perspective-driven attention activation term and the boundary-sensitive reverse response term, which is used for the discrimination of high-confidence pixels in the crack region and provides regional support for the subsequent extraction of the set of crack image coordinates; ; Among them, is the feature coupling adjustment factor, and the feature coupling adjustment factor is used to control the degree of gradient expansion of the response in channel fusion; is a non - linear response compression factor, which is used to adjust the response frequency of the reverse boundary judgment field; represents the gated attention activation guided by the alignment tensor; in the formula of is used to construct the reverse Gaussian response to the background and boundary regions, that is, to enhance the edge response separability; represents the sigmoid judgment function, which makes the encoded value converge in the interval [0, 1], and is used for region segmentation.

[0023] When constructing the spatial back - projection function , the set of image coordinates of the crack region extracted in the image domain , combined with the imaging parameters of the image acquisition device and the perspective estimation vector of the image acquisition platform are back - projected into the three - dimensional world coordinate system to obtain the actual spatial position of the crack; based on three spatial geometric property extraction functions are constructed, which are respectively used to calculate the direction angle of the crack, the curvature of the crack, and the spatial length of the crack, providing an input basis for structured recording and three - dimensional visualization atlas; ; ; where represents the three - dimensional space coordinate point corresponding to the th pixel in the crack region; , represents the transverse differential component of the th spatial differential vector in the three - dimensional space, which is used to reflect the displacement change of the crack in the X direction; represents the longitudinal differential component of the th spatial differential vector in the three - dimensional space, which is used to reflect the displacement change of the crack in the Y direction; represents the height differential component of the th spatial differential vector in the three - dimensional space, which is used to reflect the up - and - down undulation change of the crack in the Z direction; represents the th image coordinate extracted from the crack mask in the image frame; the imaging parameter represents the internal parameter matrix of the imaging parameters of the image acquisition device, and the internal parameter matrix includes focal length, principal point coordinates, and pixel scaling factor, It represents a 3x3 rotation matrix used to describe the rotation relationship between coordinate systems; In the square root term in the formula is the two-dimensional Euclidean distance, which is used to estimate the degree of direction change; It represents non-linear amplification of small direction changes to enhance the perception of directional changes; Where represents the component in the height direction of the three-dimensional space coordinates; represents the arc length parameter along the crack path; is the second derivative of the path, which is used to describe the degree of curvature in space; represents taking the absolute value; is used to enhance the differential expression of micro-curvature segments; Where represents the actual length of the entire crack in three-dimensional space; represents the number of all back-projected points in the crack area image coordinate set; In each term in the formula represents the Euclidean distance between adjacent points, and the sum of all terms gives the total length; it can be used for physical scale quantitative analysis and structural defect level assessment.

[0024] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A crack image processing method based on neural network and plane geometric image transformation, including an image acquisition platform; S1, the image acquisition platform projects the structured image coding pattern onto the surface of the target to be measured through a preset image coding pattern template, and synchronously acquires image frames containing image markers to construct a spatial anchor point for image geometric reference in the image plane; Features: S2. Obtaining the view parameter change information during the image acquisition process through the image acquisition platform, and calculating the theoretical position distribution of the image marker in the image frame in combination with the projection configuration parameters of the image coding pattern, so as to establish a predicted position set of anchor points on the image plane; S3, by extracting and decoding the actual image markers in the image frame and registering them with the predicted position points, a set of spatial offset vectors is generated to describe the geometric distortion characteristics of the image; S4, performing geometric image transformation in the image plane, constructing a nonlinear inverse transformation function of the image frame based on the set of spatial offset vectors, and performing pixel-level remapping on the image frame to generate a sequence of image frames with corrected geometric structures; S5, jointly encoding the geometrically corrected image frame, the image coordinate set of the image markers, and the view estimation vector of the image acquisition platform into a multimodal input tensor, and inputting it into a deep recognition network to extract the image coordinate set of the crack area; S6. A spatial back-projection function is constructed using the image coordinate set of the crack area, the imaging parameters corresponding to the image frame, and the view estimation vector to restore the actual position of the crack in the three-dimensional space and generate the corresponding structured geometric expression and three-dimensional visual atlas.

2. The crack image processing method based on neural network and plane geometric image transformation according to claim 1 is characterized in that: S1 also includes: The image acquisition platform generates a structured image coding pattern through a preset image coding pattern template, and the pattern types of the image coding pattern include a structured lattice pattern, a periodic stripe pattern or a DeBruijn sequence pattern; the generation of the image coding pattern is based on the image coding parameters set by the image acquisition platform, and the image coding parameters include image unit distribution density, inter-unit spacing, and redundant fault-tolerant configuration; The image coding pattern template projects the structured image coding pattern onto the surface of the target to be measured according to the image coding parameters, and the image coding pattern and the target surface form a spatially distributed image mark as a spatial anchor point for image geometric reference; The image acquisition platform synchronously acquires image frames containing image tags through the image acquisition device carried by the platform. The image frames are accompanied by image coding parameters and acquisition timestamps, which constitute the image data required for image geometric reference. S2 also includes: The image acquisition platform obtains the view parameter change information during the image acquisition process, and the view parameter change information includes accelerometer data, gyroscope data, and positioning data; the view parameter change information is input into the view calculation algorithm, and the view estimation vector corresponding to the current image frame is output; The view estimation vector and the projection configuration parameters of the image coding pattern are jointly input into the projection calculation module of the image acquisition platform to calculate the theoretical position distribution of the image marker in the image frame; the projection configuration parameters include the projection angle, the optical axis direction, and the pattern coverage range; A predicted position set of image markers is constructed based on the theoretical position distribution, and the predicted position set is used to limit the search range of the spatial anchor point in the image frame.

3. The crack image processing method based on neural network and plane geometric image transformation according to claim 2 is characterized in that: S3 also includes: The image frame is input into the image preprocessing module carried by the image acquisition platform, and the preprocessing of grayscale normalization, edge enhancement, and noise suppression is performed, and the image coordinate set of the actual image mark is extracted from the preprocessed image frame; Perform image decoding on the image coordinate set to identify the identity code of each image tag, align the image coordinate set of the actual image tag with the corresponding predicted position point in the theoretical position distribution, and generate a set of spatial offset vectors of the image tags in the image frame based on the alignment result. The set of spatial offset vectors is used to characterize the geometric distortion characteristics of the image frame and provide an input basis for image inverse transformation; S4 also includes: The spatial offset vector set is input into the image geometry modeling module carried by the image acquisition platform to construct a nonlinear inverse transformation function of the image frame, wherein the nonlinear inverse transformation function is composed of an affine mapping function, a periodic perturbation function and a high-order residual function, and the geometric distortion characteristics of the image frame are fitted through the nonlinear inverse transformation function; The original image frame is input into the nonlinear inverse transformation function, pixel-level remapping is performed, and the geometric structure correction of the image frame is completed; the correction process keeps the image texture information intact and eliminates the image geometric distortion error caused by the change of the image acquisition platform perspective; All geometrically corrected image frames and their corresponding timestamps and view estimation vectors are combined into a corrected image frame sequence.

4. The crack image processing method based on neural network and plane geometric image transformation according to claim 3 is characterized in that: The S5 also includes: Jointly encode the geometrically corrected image frame, the image coordinate set of the image label and the view estimation vector into a multimodal input tensor; Input the multimodal input tensor into the deep recognition network, perform multi-channel feature encoding and spatial alignment processing, and perform spatial feature alignment and crack area positioning consistency modeling by fusing feature information of image texture channel, image marker coordinate channel, and view estimation channel; After completing the multi-channel feature encoding and spatial alignment processing, the deep recognition network outputs the image mask of the crack area, the crack centerline image coordinate set and the crack boundary image coordinate set; based on the image mask, the image coordinate set of the crack area is extracted as the spatial alignment result obtained after fusing the image texture channel, the image marker coordinate channel and the view estimation channel, and the image coordinate set is input into the subsequent spatial back-projection function to construct a three-dimensional positioning relationship; The S6 also includes: A spatial back-projection function is constructed using an image coordinate set of the crack region, imaging parameters of an image acquisition device carried by the image acquisition platform corresponding to an image frame, and a viewing angle estimation vector, and the image coordinate set of the crack region is input into the spatial back-projection function to restore the actual position of the crack in three-dimensional space; Extracting the spatial geometric attributes of the crack based on the three-dimensional spatial position of the crack, the spatial geometric attributes include the crack curvature, crack direction and crack length; writing the three-dimensional spatial position and spatial geometric attributes of the crack, as well as the image frame index and recognition confidence information corresponding to the image frame into the structured geometric expression recording unit of the image acquisition platform; Generate a 3D visual map of cracks based on structured geometric expression, and output a visualization result file for inspection and diagnosis.

5. The crack image processing method based on neural network and plane geometric image transformation according to claim 4 is characterized in that: A nonlinear inverse transformation function of an image frame is constructed by introducing a horizontal inverse transformation function and a vertical inverse transformation function; The inverse lateral transformation function is expressed as: ; The longitudinal inverse transformation function is expressed as: ; in is the pixel point in the image frame Laterally remapped coordinates after nonlinear inverse transformation; is the pixel point in the image frame Longitudinal remapped coordinates after nonlinear inverse transformation; Represents the horizontal and vertical coordinates of the pixels to be inversely transformed in the original image frame; To enter the horizontal coordinate Scaling or rotation transformation coefficients for horizontal output coordinates; To enter the vertical coordinate Shear or rotational coupling coefficient to the transverse output coordinate; To enter the horizontal coordinate Shear or rotational coupling coefficient to the longitudinal output coordinate; To enter the vertical coordinate Scaling or rotation transformation factor for the vertical output coordinates; It is the translation component that controls the lateral output in the affine transformation; It is the translation component that controls the longitudinal output in the affine transformation; Indicates the real offset of the image marker in the horizontal direction in the image frame; Indicates the real offset of the image marker in the longitudinal direction of the image frame; , , They respectively represent the estimated vector components corresponding to the pitch angle, roll angle, and yaw angle of the image acquisition platform when the current image frame is acquired.

6. The crack image processing method based on neural network and plane geometric image transformation according to claim 5 is characterized in that: Constructing multi-channel attention fusion tensor ,pass Represents the spatially aligned feature tensor after fusing the image texture channel, image marker coordinate channel, and view estimation channel; ; in represents image texture features extracted from the geometrically corrected image frame; A structural feature channel constructed from a set of image marker coordinates; Represents the vector estimated by the view Constructed view estimation channel; each channel feature is gated by view-driven weight Then fusion is performed; in express A single component in ; It is nonlinear compression; based on Construct coordinate residual alignment kernel function; ; in represents the coordinate residual alignment kernel function; The image coordinate set of the crack centerline output by the deep network; The coordinate set of the crack boundary image output by the deep network; The position of the previous image mark in the image; The mask-guided crack region image coordinate set extraction function is expressed as: ; in The crack area image mask output by the deep network; Used to construct a periodic discrimination field; filter the response area through mask weighting, and finally output the image coordinate set of the crack area .

7. The crack image processing method based on neural network and plane geometric image transformation according to claim 6 is characterized in that: According to the multi-channel fusion feature tensor Constructing crack image mask , by jointly modeling a viewpoint-driven attention activation term and a boundary-sensitive inverse response term; ; in is the characteristic coupling adjustment factor; is the nonlinear response compression factor.

8. The crack image processing method based on neural network and plane geometric image transformation according to claim 7 is characterized in that: In constructing the spatial back-projection function , the image coordinates of the crack area extracted in the image domain are set , combined with the imaging parameters of the image acquisition device and the view estimation vector of the image acquisition platform , back-mapped to the three-dimensional world coordinate system to obtain the actual spatial position of the crack ;based on Construct three spatial geometric attribute extraction functions to calculate the direction angle of the crack , the curvature of the crack , the spatial length of the crack , providing input basis for structured records and three-dimensional visualization maps; ; ; ; ; in Indicates the crack area The three-dimensional space coordinate point corresponding to the pixel; , Indicates The lateral differential component of a spatial differential vector in three-dimensional space; Indicates The longitudinal differential component of a spatial differential vector in three-dimensional space; Indicates The height difference component of a spatial difference vector in three-dimensional space; represents the first image coordinates; Represents the component of the height direction in the three-dimensional space coordinate; represents the arc length parameter along the crack path; in Indicates the actual length of the entire crack in three-dimensional space; Represents the number of all back-projected points in the image coordinate set of the crack area.

Citation Information

Patent Citations

  • AI visual inspection method and system for concrete structure crack identification and analysis

    CN117876381A

  • Concrete bridge crack detection method based on computer vision

    CN119130986A

  • Bridge structure crack position identification method based on deep learning and computer vision

    CN119295660A

  • Generalized neural radiation field reconstruction method based on multi-modal information fusion

    CN119359934A

  • Deep learning based robot target recognition and motion detection method, storage medium and apparatus

    US11763485B1

Cited By

  • Dynamic follow-up visit system after maxillofacial deformity orthognathic operation based on intelligent nursing platform

    CN120376194A

  • Image processing method and system for beam crack depth based on MLP optimization

    CN120411113A

  • An image processing method and system for beam crack depth based on MLP optimization

    CN120411113B