Fracture Image Processing Method Based on Neural Network and Plane Geometry Image Transformation
The method integrates neural networks with plane geometry transformations to correct geometric distortions and align image features, improving the precision of crack detection and three-dimensional reconstruction in complex imaging conditions.
Patent Information
- Application Number
- CN202510520702.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing two-dimensional image processing methods and shallow neural networks struggle to accurately handle complex factors like viewpoint changes, reflections, and image geometric distortions, leading to misalignment and blurring of crack regions in images, hindering the correspondence between image and three-dimensional geometric structures.
A method combining neural networks with plane geometry transformations, using image encoding patterns to establish spatial anchors, estimate viewpoints, and perform geometric corrections and three-dimensional projections to align and reconstruct crack features.
Enhances the precision of crack region detection by correcting geometric distortions and aligning image features, enabling accurate three-dimensional reconstruction of crack structures.
Smart Images

Figure CN120047755B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of general image data processing driven by plane geometric image transformation and neural network recognition. More specifically, the present invention relates to a crack image processing method based on neural network and plane geometric image transformation. Background Art
[0002] The identification and processing of cracks refer to using an image acquisition device to obtain image frame data of the surface of a target structure, and combining image data processing algorithms or deep learning models based on neural networks to automatically detect and spatially locate potential crack areas in the image, and then extract image structure features such as crack boundaries and centerlines to provide data support for subsequent structural safety assessment and repair.
[0003] The existing methods are traditional processing methods based on two-dimensional images or shallow neural network models. When facing complex factors such as perspective transformation disturbances, target surface reflection interference, and image geometric distortion during the image acquisition process, the projection of the crack area in the image is prone to dislocation, blurring, and even structural fracture, making it difficult to establish the correspondence between the crack image and its three-dimensional geometric structure, which limits the recognition and spatial structure restoration capabilities. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a crack image processing method based on neural network and plane geometric image transformation, by constructing a joint model of image geometric correction and three-dimensional back-projection that integrates image anchor points, perspective estimation, and image texture information, to solve the problems proposed in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solution: A crack image processing method based on neural network and plane geometric image transformation, including an image acquisition platform;
[0006] S1. The image acquisition platform projects a structured image coding pattern onto the surface of the target to be measured through a preset image coding pattern template, and synchronously acquires an image frame containing image markers to construct spatial anchor points for image geometric reference in the image plane;
[0007] S2. Obtain the perspective parameter change information during the image acquisition process through the image acquisition platform, and combine the projection configuration parameters of the image coding pattern to calculate the theoretical position distribution of the image markers in the image frame, so as to establish a set of predicted positions of the anchor points in the image plane;
[0008] S3. Extract and decode the actual image markers in the image frame, and register them with the predicted position points to generate a set of spatial offset vectors for describing the geometric distortion characteristics of the image;
[0009] S4. Perform a geometric image transformation in the image plane, construct a non-linear inverse transformation function for the image frame based on the set of spatial offset vectors, and perform pixel-level remapping on the image frame to generate a sequence of image frames with corrected geometric structures;
[0010] S5. Jointly encode the image frames with corrected geometric structures, the set of image coordinates of the image markers, and the perspective estimation vector of the image acquisition platform into a multi-modal input tensor, and input it into the deep recognition network to extract the set of image coordinates of the crack region;
[0011] S6. Use the set of image coordinates of the crack region, the imaging parameters corresponding to the image frame, and the perspective estimation vector to construct a spatial back-projection function, restore the actual position of the crack in the three-dimensional space, and generate the corresponding structured geometric expression and three-dimensional view spectrum.
[0012] In a preferred embodiment, S1 further includes:
[0013] The image acquisition platform generates a structured image coding pattern through a preset image coding pattern template. The pattern types of the image coding pattern include a structure dot matrix pattern, a periodic stripe pattern, or a De Bruijn sequence pattern; the generation of the image coding pattern is based on the image coding parameters set by the image acquisition platform, and the image coding parameters include the distribution density of image units, the spacing between units, and the redundancy and fault tolerance configuration;
[0014] The image coding pattern template projects the structured image coding pattern onto the surface of the target to be measured according to the image coding parameters, and the image coding pattern and the target surface form a spatially distributed image marker, serving as a spatial anchor for image geometry reference;
[0015] The image acquisition platform synchronously acquires an image frame containing the image marker through the mounted image acquisition device. The image frame is attached with the image coding parameters and the acquisition timestamp, constituting the image data required for image geometry reference;
[0016] S2 further includes:
[0017] The image acquisition platform obtains the perspective parameter change information during its image acquisition process. The perspective parameter change information includes accelerometer data, gyroscope data, and positioning data; by inputting the perspective parameter change information into the perspective solution algorithm, the perspective estimation vector corresponding to the current image frame is output;
[0018] The perspective estimation vector and the projection configuration parameters of the image coding pattern are jointly input into the projection calculation module of the image acquisition platform to calculate the theoretical position distribution of the image marker in the image frame; the projection configuration parameters include the projection angle, the optical axis direction, and the pattern coverage range;
[0019] Based on the theoretical position distribution, a set of predicted positions for image markers is formed, and the set of predicted positions is used to define the search range of spatial anchor points in the image frame.
[0020] In a preferred embodiment, S3 further includes:
[0021] Input the image frame into the image preprocessing module carried by the image acquisition platform, perform preprocessing such as gray normalization, edge enhancement, and noise suppression, and extract the set of image coordinates of the actual image markers in the preprocessed image frame;
[0022] Perform image decoding on the set of image coordinates to identify the identity codes of each image marker, register the set of image coordinates of the actual image markers with the corresponding predicted position points in the theoretical position distribution, and generate a set of spatial offset vectors of the image markers in the image frame based on the registration result. The set of spatial offset vectors is used to characterize the geometric distortion characteristics of the image frame and provide an input basis for image inverse transformation;
[0023] S4 further includes:
[0024] Input the set of spatial offset vectors into the image geometric modeling module carried by the image acquisition platform to construct a non-linear inverse transformation function of the image frame. The non-linear inverse transformation function is composed of an affine mapping function, a periodic perturbation function, and a high-order residual function, and fits the geometric distortion characteristics of the image frame through the non-linear inverse transformation function;
[0025] Input the original image frame into the non-linear inverse transformation function to perform pixel-level remapping and complete the geometric structure correction of the image frame; the texture information of the image is kept intact during the correction process, and the geometric distortion error caused by the perspective transformation change of the image acquisition platform is eliminated;
[0026] All the image frames with corrected geometric structures, their corresponding timestamps, and the perspective estimation vectors are combined into a sequence of corrected image frames.
[0027] In a preferred embodiment, S5 further includes:
[0028] Jointly encode the image frames with corrected geometric structures, the set of image coordinates of the image markers, and the perspective estimation vectors into a multi-modal input tensor;
[0029] Input the multi-modal input tensor into the deep recognition network to perform multi-channel feature encoding and spatial alignment processing. The multi-channel feature encoding and spatial alignment processing fuse the feature information of the image texture channel, the image marker coordinate channel, and the perspective estimation channel, and perform spatial feature alignment and consistency modeling of crack region localization;
[0030] After completing the multi-channel feature encoding and spatial alignment processing, the deep recognition network outputs the image mask of the crack area, the crack centerline image coordinate set and the crack boundary image coordinate set; based on the image mask, the image coordinate set of the crack area is extracted as the spatial alignment result obtained after fusing the image texture channel, the image marker coordinate channel and the view estimation channel, and the image coordinate set is input into the subsequent spatial back-projection function to construct a three-dimensional positioning relationship;
[0031] The S6 also includes:
[0032] A spatial back-projection function is constructed using an image coordinate set of the crack region, imaging parameters of an image acquisition device carried by the image acquisition platform corresponding to an image frame, and a viewing angle estimation vector, and the image coordinate set of the crack region is input into the spatial back-projection function to restore the actual position of the crack in three-dimensional space;
[0033] Extracting the spatial geometric attributes of the crack based on the three-dimensional spatial position of the crack, the spatial geometric attributes include the crack curvature, crack direction and crack length; writing the three-dimensional spatial position and spatial geometric attributes of the crack, as well as the image frame index and recognition confidence information corresponding to the image frame into the structured geometric expression recording unit of the image acquisition platform;
[0034] Generate a 3D visual map of cracks based on structured geometric expression, and output a visualization result file for inspection and diagnosis.
[0035] In the formula structure involved in this scheme, dimensionless terms can be used as proportional or structural adjustment factors. When combined with quantities with units, they only play a role in numerical scaling without introducing new physical dimensions, so they will not change or confuse the unit system of the overall expression. This combination of "dimensionless terms and units" can be understood as a composite structural expression commonly used in mathematical and physical modeling, which conforms to the principle of dimensional consistency and has a clear physical interpretation basis.
[0036] Secondly, in the formula structure of this scheme, if multiple variables with different physical units are involved, including but not limited to time, mass or energy variables, their joint appearance is to express the collaborative modeling relationship of multiple physical mechanisms. Each variable forms a unified structure through function mapping, ratio combination or normalization adjustment, with clear units and meanings, and the overall expression conforms to the principle of dimensional consistency and the common formula of engineering modeling;
[0037] In a preferred embodiment, a nonlinear inverse transformation function of the image frame is constructed by introducing a horizontal inverse transformation function and a vertical inverse transformation function;
[0038] The inverse lateral transformation function is expressed as:
[0039] ;
[0040] The longitudinal inverse transformation function is expressed as:
[0041] ;
[0042] where is the horizontal remapping coordinate of the pixel point in the image frame after non - linear inverse transformation; is the vertical remapping coordinate of the pixel point in the image frame after non - linear inverse transformation; represents the horizontal and vertical coordinates of the pixel to be inverse - transformed in the original image frame; is the input horizontal coordinate and is the scaling or rotation transformation coefficient for the horizontal output coordinate; is the input vertical coordinate and is the shear or rotation coupling coefficient for the horizontal output coordinate; is the input horizontal coordinate and is the shear or rotation coupling coefficient for the vertical output coordinate; is the input vertical coordinate and is the scaling or rotation transformation coefficient for the vertical output coordinate; is the translation component controlling the horizontal output in the affine transformation; is the translation component controlling the vertical output in the affine transformation; represents the actual offset in the horizontal direction of the image marker in the image frame; represents the actual offset in the vertical direction of the image marker in the image frame; , , respectively represent the estimated vector components corresponding to the pitch angle, roll angle, and yaw angle of the image acquisition platform during the acquisition of the current image frame.
[0043] In a preferred embodiment, a multi - channel attention fusion tensor is constructed, and through it represents the spatial alignment feature tensor after fusing the image texture channel, the image marker coordinate channel, and the viewing angle estimation channel;
[0044] ;
[0045] where represents the image texture features extracted from the image frame with corrected geometric structure; is the structural feature channel constructed from the set of image marker coordinates; represents the viewing angle estimation channel constructed from the viewing angle estimation vector ; where each channel feature is fused after passing through the viewing - angle - driven weight gating ;
[0046] wherein denotes a single component in; is non - linear compression;
[0047] Based on construct a coordinate residual alignment kernel function; ;
[0048] wherein denotes the coordinate residual alignment kernel function; is the set of crack center - line image coordinates output by the deep network; is the set of crack boundary image coordinates output by the deep network; is the position of the previous image marker in the image;
[0049] The function for extracting the set of image coordinates of the crack area under the mask guidance is expressed as: ;
[0050] wherein is the crack area image mask output by the deep network; is used to construct a periodic discrimination field; by screening the response area with mask weighting, finally output the set of image coordinates of the crack area .
[0051] In a preferred embodiment, according to the multi - channel fusion feature tensor construct a crack image mask , jointly modeled by a view - driven attention activation term and a boundary - sensitive reverse response term;
[0052] ;
[0053] wherein is a feature coupling adjustment factor; is a non - linear response compression factor.
[0054] In a preferred embodiment, when constructing the spatial back - projection function , the set of image coordinates of the crack area extracted in the image domain , combined with the imaging parameters of the image acquisition device and the view - angle estimation vector of the image acquisition platform , is mapped back to the three - dimensional world coordinate system to obtain the actual spatial position of the crack ; Based on construct three spatial geometric property extraction functions, which are respectively used to calculate the direction angle of the crack , the curvature of the crack , and the spatial length of the crack , providing an input basis for structured records and three-dimensional visualization atlases;
[0055] ;
[0056] ;
[0057] ;
[0058] ;
[0059] where represents the three-dimensional spatial coordinate point corresponding to the th pixel in the crack area; , represents the horizontal differential component of the th spatial differential vector in three-dimensional space; represents the longitudinal differential component of the th spatial differential vector in three-dimensional space; represents the height differential component of the th spatial differential vector in three-dimensional space; represents the th image coordinate extracted from the crack mask in the image frame; represents the component in the height direction of the three-dimensional spatial coordinates; represents the arc length parameter along the crack path;
[0060] where represents the actual length of the entire crack in three-dimensional space; represents the number of all back-projected points in the set of image coordinates of the crack area.
[0061] Technical effects and advantages of the present invention:
[0062] By constructing an image offset modeling mechanism jointly driven by image coding patterns and spatial anchor points, in-plane correction of geometric distortion features in the image frame is achieved, improving the structural consistency and texture continuity before subsequent image analysis;
[0063] Fusing image texture features, image marker coordinates, and perspective estimation vectors into a multi-channel attention tensor to drive a deep neural network to perform feature alignment and robust recognition, thereby improving the crack area localization accuracy under complex perspective perturbations;
[0064] Constructing an image alignment kernel function based on coordinate residuals and fused feature responses to achieve spatial consistency alignment of the crack centerline and boundary line, effectively suppressing the image boundary fracture problem caused by projection offset or depth change;
[0065] Introduce a dynamic perspective adjustment mechanism, adaptively adjust the response intensity of the disturbance term by combining the steady-state information during the image acquisition process, and reduce the error transmission effect of image deformation during geometric transformation and recognition;
[0066] By constructing an image space back-projection function and a crack geometric attribute extraction function, restore the three-dimensional spatial pose of the crack on the basis of image plane processing, and extract its structural parameters such as direction, curvature and length to support subsequent three-dimensional structural damage modeling and quantitative analysis. Brief Description of the Drawings
[0067] Figure 1 It is a schematic diagram of the method steps of the present invention. Detailed Embodiments
[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0069] Refer to the attached drawings of the specification Figure 1 For a crack image processing method based on a neural network and planar geometric image transformation according to an embodiment of the present invention, it includes: an image acquisition platform;
[0070] S1. The image acquisition platform projects a structured image coding pattern onto the surface of the target to be measured through a preset image coding pattern template, and synchronously acquires an image frame containing image markers to construct spatial anchor points for image geometric reference;
[0071] S2. Obtain the perspective parameter change information during the image acquisition process through the image acquisition platform, and calculate the theoretical position distribution of the image markers in the image frame in combination with the projection configuration parameters of the image coding pattern to establish a set of predicted positions of the anchor points;
[0072] S3. Extract and decode the actual image markers in the image frame, and register them with the corresponding predicted position points in the theoretical position distribution to generate a set of spatial offset vectors of the image markers in the image frame, which are used to characterize the image distortion characteristics;
[0073] S4. Based on the set of spatial offset vectors, construct a non-linear inverse transformation function of the image frame, and perform pixel-level remapping on the image frame to generate a sequence of image frames with corrected geometric structures;
[0074] S5. Jointly encode the image frame with corrected geometric structure, the set of image coordinates of the image markers, and the perspective estimation vector of the image acquisition platform into a multi-modal input tensor, and input it into the deep recognition network to extract the set of image coordinates of the crack region;
[0075] S6. Use the set of image coordinates of the crack region, the imaging parameters corresponding to the image frame, and the perspective estimation vector to construct a spatial back-projection function, restore the actual position of the crack in three-dimensional space, and generate the corresponding structured geometric expression and three-dimensional visual spectrum.
[0076] In the formula structure involved in this solution, dimensionless terms can be used as proportional or structural adjustment factors. When combined with quantities with units, they only play a role in numerical scaling and do not introduce new physical dimensions, so they will not change or confuse the overall unit system of the expression; such combinations of "dimensionless terms and terms with unit quantities" can be understood as the composite structure expression forms commonly used in mathematical physics modeling, conform to the principle of dimensional consistency, and have a clear physical interpretation basis;
[0077] Secondly, in the formula structure of this solution, if there are multiple variable terms with different physical units, including but not limited to time-type, mass-type, or energy-type variables, their joint appearance is to express the collaborative modeling relationship of multiple physical mechanisms. Each variable forms a unified structure through function mapping, ratio combination, or normalization adjustment, with clear units and clear meanings. The overall expression conforms to the principle of dimensional consistency and the common norms of engineering modeling;
[0078] Regarding the above solution, it should be further supplemented and explained that:
[0079] S1 further includes:
[0080] The image acquisition platform generates a structured image coding pattern through a preset image coding pattern template. The pattern types of the image coding pattern include structural dot matrix patterns, periodic stripe patterns, or DeBruijn sequence patterns; the generation of the image coding pattern is based on the image coding parameters set by the image acquisition platform, and the image coding parameters include the distribution density of image units, the spacing between units, and the redundant fault tolerance configuration;
[0081] The image coding pattern template projects the structured image coding pattern onto the surface of the target to be measured according to the image coding parameters, and the image coding pattern and the target surface form a spatially distributed image marker, serving as a spatial anchor for image geometric reference;
[0082] The image acquisition platform synchronously acquires an image frame containing the image marker through the mounted image acquisition device. The image frame is attached with the image coding parameters and the acquisition timestamp, constituting the image data required for image geometric reference;
[0083] Regarding S1, it should be further noted that according to the pattern type selection parameters set by the image acquisition platform, different structured image coding pattern prototypes are generated correspondingly. Each type of pattern corresponds to a different function generation mechanism, including a structure dot matrix function, a periodic stripe function, and a De Bruijn coding function, which are used for subsequent projection and registration in practical applications;
[0084] ;
[0085] In the above formula, is the image coding pattern prototype function generated in the two-dimensional space, and its output is used to construct the projection pattern; , is the virtual coordinate variable in the pattern generation process, which refers to the coordinates that have not been mapped to the target surface; is the distribution density of image units; is the spacing between units; is the redundant perturbation variable, and the redundant fault tolerance configuration is described by ; is the pattern type control variable, which is an integer variable used to select the pattern type: when represents the structure dot matrix, when represents the periodic stripe, and when represents the De Bruijn sequence coding; in addition, in the formula represents the floor symbol; represents the modulo operation, that is, taking the remainder after dividing the value in the parentheses by 3, which is used to construct the periodic coding logic;
[0086] Based on S1, a structured image coding pattern function is constructed, which is used to introduce the geometric parameter regulation and redundancy enhancement strategy of the overall image structure on the basis of the selected pattern type, and construct a complete projectable image coding pattern, which includes a periodic structure, spatial variation, and fault tolerance ability;
[0087] ;
[0088] Among them, is the pattern geometric structure strength, which controls the density through λ and the relative spacing through ξ; is the redundant fault tolerance function, which is used to insert redundant marker points or local perturbations in the coding pattern, and η controls the fault tolerance amplitude;
[0089] A spatial coordinate perturbation function after the pattern is projected onto the target surface is constructed, which is used to project the generated image coding pattern coordinates from the pattern plane into the target surface coordinate system and add a perturbation factor to simulate the distortion effects caused by distance, angle, and viewing angle in the real projection process;
[0090] ;
[0091] wherein are the actual projection coordinates formed by the image coding pattern on the target surface; is the pattern projection angle parameter; is the optical axis direction parameter; is used to simulate the non - linear perturbation caused by the local high - frequency fluctuations on the target surface during the projection process, reflecting the periodic distortion in the combined coordinate direction; is used to simulate the geometric deviation generated in the orthogonal axis direction, manifested as the pattern deformation caused by directional reflection or refraction;
[0092] Immediately afterwards, the pattern function is mapped according to the projection coordinates to obtain the image marker map value function on the target surface: ;
[0093] wherein is the image marker intensity function value on the target surface, used for subsequent identification; is the structured image coding pattern function;
[0094] When the acquisition device obtains an image frame, the image marker content, pattern parameters, pattern type and acquisition timestamp are structurally recorded for the image frame, for subsequent coordinate decoding and perspective calculation use, which is expressed as:
[0095] ;
[0096] wherein is the image frame data structure; is the image marker image data; is the image frame acquisition timestamp; is the transpose symbol;
[0097] S2 further includes:
[0098] The image acquisition platform obtains the change information of the perspective parameters during its image acquisition process. The change information of the perspective parameters includes accelerometer data, gyroscope data, and positioning data; by inputting the change information of the perspective parameters into the perspective calculation algorithm, the perspective estimation vector corresponding to the current image frame is output;
[0099] The perspective estimation vector and the projection configuration parameters of the image coding pattern are jointly input into the projection calculation module of the image acquisition platform to calculate the theoretical position distribution of the image markers in the image frame; the projection configuration parameters include the projection angle, the optical axis direction, and the pattern coverage range;
[0100] Based on the theoretical position distribution, a prediction position set of the image markers is formed, and the prediction position set is used to define the search range of the spatial anchor points in the image frame;
[0101] Regarding S2, it should be further noted that according to the three types of perspective parameter change information collected in real time during the acquisition of UAV images: accelerometer data, gyroscope data, and positioning data, through the composite mapping of the exponential function and the arctangent function, a perspective estimation vector for spatially describing the current state of the image frame is generated; the perspective estimation vector is used to correct the projection angle and the deviation of the optical axis during the projection process, so that the position of the image marker in the image frame achieves theoretical consistency;
[0102] ;
[0103] where is the perspective estimation vector, and the perspective estimation vector is used to describe the spatial perspective of the image tilt at the current acquisition moment; is the observed value of the accelerometer in three axes; is the angular velocity observed value of the gyroscope in three axes; is the three-dimensional spatial position or velocity observed value provided by the navigation system; is used to correct the image projection angle; is used to correct the direction of the image optical axis; is used to correct the pattern coverage perturbation; is the base of the natural logarithm; is the arctangent function;
[0104] The above formula is used to calculate the spatial theoretical position distribution of the image coding pattern on the basis of known Ψ and the projection configuration parameters of the image coding pattern, that is, the coordinate position where the pattern should appear in the image frame in the ideal state;
[0105] ;
[0106] where is the virtual point in the two-dimensional image coding pattern is the theoretical projection coordinate in the image frame; is the projection angle of the image coding pattern; is the pattern optical axis direction; is the pattern coverage range factor, and the pattern coverage range factor is used to control the perturbation amplitude; , are respectively used to introduce the periodic terms of the asymmetric perturbation to simulate the pattern deformation;
[0107] For each coding point in the image coding atlas perform theoretical position projection to form a predicted position set; this set will be used as the range for subsequent anchor point search and registered with the actual image marker;
[0108] ;
[0109] where is a set of predicted positions for image markers; is a set of integers.
[0110] S3 also includes:
[0111] Input the image frame into the image preprocessing module carried by the image acquisition platform, perform preprocessing such as gray normalization, edge enhancement, and noise suppression, and extract the set of image coordinates of the actual image markers in the preprocessed image frame;
[0112] Perform image decoding on the set of image coordinates to identify the identity codes of each image marker, register the set of image coordinates of the actual image markers with the corresponding predicted position points in the theoretical position distribution, and generate a set of spatial offset vectors for the image markers in the image frame. The set of spatial offset vectors is used to characterize the geometric distortion features of the image frame and provides an input basis for image inverse transformation;
[0113] Formulate represents the set of image marker coordinates detected in the image; ; where represents the horizontal coordinate of the marker in the image; is the vertical coordinate of the marker in the image; the theoretical prediction coordinate set of the image marker is ; is the horizontal coordinate of the theoretical prediction position; is the vertical coordinate of the theoretical prediction position;
[0114] By taking and to do element-by-element difference, construct a set of spatial offset vectors, which is used to characterize the geometric distortion characteristics of the image frame and serves as the input for subsequent construction of the image inverse transformation function;
[0115] ;
[0116] where is the set of spatial offset vectors; each dimension represents the difference between the actual image coordinates and the predicted coordinates, quantifying the image distortion;
[0117] Based on this, construct an image geometric distortion feature function and formulate the geometric distortion function value of the image frame , which integrates the spatial offset intensity and image coordinate position information for use in the next stage of constructing the image inverse transformation function;
[0118] ;
[0119] where is the distortion feature value of the image frame; the first input part in the formula is the Euclidean norm of the offset vector, used to measure the offset intensity; the second input part in the formula represents the sine modulation factor in the horizontal direction of the current position; the third input part in the formula represents the cosine modulation factor in the vertical direction of the current position; their combination characterizes the non-uniform distortion trend of the image at different spatial positions;
[0120] S4 further includes:
[0121] Input the set of spatial offset vectors into the image geometric modeling module carried by the image acquisition platform to construct the non-linear inverse transformation function of the image frame, where the non-linear inverse transformation function is composed of an affine mapping function, a periodic perturbation function, and a high-order residual function, and fit the geometric distortion characteristics of the image frame through the non-linear inverse transformation function;
[0122] Input the original image frame into the non-linear inverse transformation function, perform pixel-level remapping, and complete the geometric structure correction of the image frame; the texture information of the image is kept intact during the correction process, and the geometric distortion error caused by the perspective transformation of the image acquisition platform is eliminated;
[0123] All the image frames with corrected geometric structures, their corresponding timestamps, and the perspective estimation vectors are combined into a sequence of corrected image frames.
[0124] S5 further includes:
[0125] Jointly encode the image frames with corrected geometric structures, the set of image coordinates marked on the image, and the perspective estimation vectors into a multi-modal input tensor;
[0126] Input the multi-modal input tensor into the depth recognition network, perform multi-channel feature encoding and spatial alignment processing. The multi-channel feature encoding and spatial alignment processing fuse the feature information of the image texture channel, the image marked coordinate channel, and the perspective estimation channel, and perform spatial feature alignment and consistency modeling of the crack region location;
[0127] After completing the multi-channel feature encoding and spatial alignment processing, the depth recognition network outputs the image mask of the crack region, the set of image coordinates of the crack center line, and the set of image coordinates of the crack boundary; extract the set of image coordinates of the crack region based on the image mask as the spatial alignment result obtained after fusing the image texture channel, the image marked coordinate channel, and the perspective estimation channel, and input this set of image coordinates into the subsequent spatial back-projection function to construct a three-dimensional positioning relationship;
[0128] S6 further includes:
[0129] Utilize the set of image coordinates of the crack region, the imaging parameters of the image acquisition device carried by the image acquisition platform corresponding to the image frame, and the perspective estimation vector to construct a spatial back-projection function, and input the set of image coordinates of the crack region into the spatial back-projection function to restore the actual position of the crack in three-dimensional space;
[0130] Extract the spatial geometric attributes of cracks based on their three-dimensional spatial positions. The spatial geometric attributes include crack curvature, crack direction, and crack length. Write the three-dimensional spatial positions and spatial geometric attributes of the cracks, as well as the image frame index and recognition confidence information corresponding to the image frame, into the structured geometric expression recording unit of the image acquisition platform.
[0131] Generate a three-dimensional visual spectrum of the cracks based on the structured geometric expression, and output a visualization result file for inspection and diagnosis.
[0132] Construct a non-linear inverse transformation function for the image frame by introducing a horizontal inverse transformation function and a vertical inverse transformation function, which is used to eliminate geometric distortions caused by spatial offset, periodic perturbation, non-linear residuals, local modulation, and perspective changes, and achieve pixel-level image correction.
[0133] The horizontal inverse transformation function is expressed as:
[0134] ;
[0135] The vertical inverse transformation function is expressed as:
[0136] ;
[0137] Where is the horizontal remapping coordinate of the pixel point in the image frame after non-linear inverse transformation; is the vertical remapping coordinate of the pixel point in the image frame after non-linear inverse transformation; represents the horizontal and vertical coordinates of the pixel to be inverse-transformed in the original image frame; is the sum of the squares of the pixel coordinates, which is used to represent the spatial amplitude and construct the position modulation term; is the cube of the horizontal coordinate, which is used to model the high-order non-linear residual perturbation in the horizontal direction; is the cube of the vertical coordinate, which is used to model the high-order non-linear residual perturbation in the vertical direction; is the sum of the fourth powers of the pixel coordinates, which is used to enhance the scale response and perturbation amplification in the image edge region; is the scaling or rotation transformation coefficient of the input horizontal coordinate for the horizontal output coordinate; is the shear or rotation coupling coefficient of the input vertical coordinate for the horizontal output coordinate; is the shear or rotation coupling coefficient of the input horizontal coordinate for the vertical output coordinate; is the scaling or rotation transformation coefficient of the input vertical coordinate for the vertical output coordinate; is the translation component for controlling the horizontal output in the affine transformation; is the translation component for controlling the vertical output in the affine transformation;
[0138] The spatial offset vector term in the formula includes:
[0139] represents the actual offset in the horizontal direction of the image marker in the image frame;
[0140] represents the actual offset in the vertical direction of the image marker in the image frame;
[0141] represents the non - linear high - order modulation of the pixel position based on the spatial offset;
[0142] The periodic perturbation function term in the formula includes:
[0143] represents the periodic perturbation with the spatial offset as the frequency parameter, used for modeling device jitter or target surface reflection interference;
[0144] is a composite periodic term, and the composite periodic term is used to form a two - dimensional spatial rotation perturbation kernel;
[0145] The high - order residual function term in the formula includes:
[0146] is a non - linear activation mapping, and the non - linear activation mapping is used to control the convergence of distortion in the image edge region;
[0147] represents the attenuation control based on the spatial offset energy, which is used to prevent excessive local perturbation;
[0148] represents the amplification of the response in the image edge region, that is, modeling the distortion caused by surface curvature;
[0149] The spatial modulation kernel term in the formula includes:
[0150] represents simulating the response difference of different regions in the image to distortion, which is used to form a non - uniform perturbation kernel;
[0151] is used to reflect the structural correction requirement when there is symmetric anti - perturbation in the image;
[0152] The perspective regulation weight term in the formula includes:
[0153] , , respectively represent the estimated vector components corresponding to the pitch angle, roll angle, and yaw angle when the image acquisition platform acquires the current image frame;
[0154] is the viewing angle regulation weight coefficient used to dynamically adjust the action intensity of the perturbation function; where represents the perturbation adjustment factor constructed based on the viewing angle estimation vector When the viewing angle change is stable, that is, Ψ is close to 0, is close to 1, and the action of the perturbation function is enhanced; when the viewing angle fluctuates violently, decreases, which is used to suppress the distortion response intensity and reduce the introduction of geometric errors.
[0155] Construct a multi-channel attention fusion tensor , through represents the spatially aligned feature tensor after fusing the image texture channel, image marker coordinate channel, and viewing angle estimation channel, which is used to express the spatial consistency features of the crack region extracted by the depth recognition network under multi-modal input, and provides fusion feature support for subsequent mask generation, coordinate residual modeling, and region localization;
[0156] ;
[0157] where represents the image texture features extracted from the image frame with corrected geometric structure; The structural feature channel constructed from the set of image marker coordinates; represents the viewing angle estimation channel constructed from the viewing angle estimation vector ; where each channel feature is fused after passing through the viewing angle-driven weight gating ;
[0158] where represents a single component in, in , is used as the adjustment variable for the channel attention weight; when is small, it indicates that the viewing angle transformation is stable, and this weight approaches 1, indicating that the channel feature participates in the fusion with high confidence; when is large, it indicates that the viewing angle transformation changes violently, and this weight then approaches 0, indicating that the channel feature reduces its influence in the fusion to suppress the noise caused by viewing angle jitter; is non-linear compression, and non-linear compression is used to control the fusion activation degree to generate the final spatially aligned feature tensor ;
[0159] Based on construct a coordinate residual alignment kernel function; ;
[0160] wherein represents the coordinate residual alignment kernel function; is the set of crack centerline image coordinates output by the deep network; is the set of crack boundary image coordinates output by the deep network; is the position of the previous image marker in the image, serving as a spatial anchor; represents the coordinate difference between the crack centerline and the image marker; represents the coordinate difference between the crack boundary and the image marker; and and the fusion feature energy term together constitute , which is used for subsequent region determination;
[0161] The function for extracting the set of crack region image coordinates under the mask guidance is expressed as: ;
[0162] wherein is the crack region image mask output by the deep network; in in the formula is used to construct a periodic discrimination field, so that high-confidence regions form significant responses; the response regions are screened by mask weighting, and finally the set of image coordinates of the crack region is output, which is used as the input of the three-dimensional space back-projection function in the next stage.
[0163] According to the multi-channel fusion feature tensor construct the crack image mask , through the joint modeling of the view-driven attention activation term and the boundary-sensitive reverse response term, which is used for the discrimination of high-confidence pixels in the crack region and provides region support for the subsequent extraction of the crack image coordinate set;
[0164] ;
[0165] wherein is the feature coupling adjustment factor, and the feature coupling adjustment factor is used to control the gradient expansion degree of the response in channel fusion; is the non-linear response compression factor, and the non-linear response compression factor is used to adjust the response frequency of the reverse boundary discrimination field; represents the gated attention activation guided by the alignment tensor; in in the formula of part is used to construct the reverse Gaussian response to the background and boundary regions, that is, to enhance the edge response separability; represents the sigmoid judgment function, which makes the coding value converge in the interval [0, 1], For region segmentation.
[0166] When constructing the spatial back-projection function , the set of image coordinates of the crack region extracted in the image domain , combined with the imaging parameters of the image acquisition device and the perspective estimation vector of the image acquisition platform , is mapped back to the three-dimensional world coordinate system to obtain the actual spatial position of the crack ; Based on , three spatial geometric property extraction functions are constructed, which are respectively used to calculate the direction angle of the crack , the curvature of the crack , and the spatial length of the crack , providing an input basis for structured recording and three-dimensional visualization atlas;
[0167] ;
[0168] ;
[0169] ;
[0170] ;
[0171] Where represents the three-dimensional space coordinate point corresponding to the th pixel in the crack region; , represents the horizontal difference component of the th spatial difference vector in the three-dimensional space, which is used to reflect the displacement change of the crack in the X direction; represents the longitudinal difference component of the th spatial difference vector in the three-dimensional space, which is used to reflect the displacement change of the crack in the Y direction; represents the height difference component of the th spatial difference vector in the three-dimensional space, which is used to reflect the up and down fluctuation change of the crack in the Z direction; represents the th image coordinate extracted from the crack mask in the image frame; The imaging parameter represents the internal parameter matrix of the imaging parameters of the image acquisition device. The internal parameter matrix includes focal length, principal point coordinates, and pixel scaling factor. represents a 3x3 rotation matrix, which is used to describe the rotation relationship between coordinate systems;
[0172] In , the square root term in the formula is the two-dimensional Euclidean distance, which is used to estimate the degree of direction change; represents the use of non-linear amplification of small direction changes to enhance the perception of directional changes;
[0173] wherein represents the component in the height direction of the three-dimensional space coordinates; represents the arc length parameter along the crack path; is the second derivative of the path, which is used to describe the degree of curvature in space; represents taking the absolute value; is used to enhance the differential expression of the micro-curvature segment;
[0174] wherein represents the actual length of the whole crack in the three-dimensional space; represents the number of all back-projected points in the crack area image coordinate set; in each term in the formula represents the Euclidean distance between adjacent points, and the sum of all terms gives the total length; it can be used for physical scale quantitative analysis and structural defect level assessment.
[0175] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A crack image processing method based on a neural network and planar geometric image transformation, including an image acquisition platform; S1. The image acquisition platform projects a structured image coding pattern onto the surface of the target to be measured through a preset image coding pattern template, and synchronously acquires an image frame containing image markers to construct spatial anchor points for image geometric reference in the image plane; It is characterized in that: S2. The image acquisition platform obtains the change information of the perspective parameter during its image acquisition process, and combines it with the projection configuration parameters of the image coding pattern to calculate the theoretical position distribution of the image markers in the image frame, so as to establish a set of predicted positions of the anchor points in the image plane; S3. By extracting and decoding the actual image markers in the image frame and registering them with the predicted position points, a set of spatial offset vectors is generated to describe the geometric distortion characteristics of the image; S4. Perform geometric image transformation in the image plane, construct a non-linear inverse transformation function of the image frame based on the set of spatial offset vectors, and perform pixel-level remapping on the image frame to generate a sequence of image frames with corrected geometric structures; S5. Jointly encode the image frame with corrected geometric structure, the set of image coordinates of the image markers, and the perspective estimation vector of the image acquisition platform into a multi-modal input tensor, and input it into the deep recognition network to extract the set of image coordinates of the crack area; S6. Use the set of image coordinates of the crack area, the imaging parameters corresponding to the image frame, and the perspective estimation vector to construct a spatial back-projection function, restore the actual position of the crack in three-dimensional space, and generate the corresponding structured geometric expression and three-dimensional view spectrum.
2. The crack image processing method based on a neural network and planar geometric image transformation according to claim 1, characterized in that: S1 further includes: The image acquisition platform generates a structured image coding pattern through a preset image coding pattern template. The pattern types of the image coding pattern include a structure dot matrix pattern, a periodic stripe pattern, or a DeBruijn sequence pattern; the generation of the image coding pattern is based on the image coding parameters set by the image acquisition platform, and the image coding parameters include the distribution density of image units, the spacing between units, and the redundant fault tolerance configuration; The image coding pattern template projects the structured image coding pattern onto the surface of the target to be measured according to the image coding parameters. The image coding pattern and the target surface form a spatially distributed image marker, which serves as a spatial anchor point for image geometric reference; The image acquisition platform synchronously acquires an image frame containing image markers through the mounted image acquisition device. The image frame is attached with image coding parameters and acquisition timestamps to form the image data required for image geometric reference; S2 further includes: The image acquisition platform obtains the change information of the perspective parameter during its image acquisition process. The change information of the perspective parameter includes accelerometer data, gyroscope data, and positioning data; by inputting the change information of the perspective parameter into the perspective solution algorithm, the perspective estimation vector corresponding to the current image frame is output; The perspective estimation vector and the projection configuration parameters of the image coding pattern are jointly input into the projection calculation module of the image acquisition platform to calculate the theoretical position distribution of the image markers in the image frame; the projection configuration parameters include the projection angle, the optical axis direction, and the pattern coverage range; Based on the theoretical position distribution, a set of predicted positions for image markers is formed, and the set of predicted positions is used to define the search range of spatial anchor points in the image frame.
3. The method for processing crack images based on neural network and plane geometric image transformation according to claim 2, characterized in that: S3 further includes: Input the image frame into the image preprocessing module carried by the image acquisition platform, perform preprocessing such as gray normalization, edge enhancement, and noise suppression, and extract the set of image coordinates of the actual image markers in the preprocessed image frame; Perform image decoding on the set of image coordinates to identify the identity codes of each image marker, register the set of image coordinates of the actual image markers with the corresponding predicted position points in the theoretical position distribution, and generate a set of spatial offset vectors of the image markers in the image frame based on the registration result. The set of spatial offset vectors is used to characterize the geometric distortion characteristics of the image frame and provide an input basis for image inverse transformation; S4 further includes: Input the set of spatial offset vectors into the image geometric modeling module carried by the image acquisition platform to construct a non-linear inverse transformation function of the image frame. The non-linear inverse transformation function is composed of an affine mapping function, a periodic perturbation function, and a high-order residual function, and fits the geometric distortion characteristics of the image frame through the non-linear inverse transformation function; Input the original image frame into the non-linear inverse transformation function to perform pixel-level remapping and complete the geometric structure correction of the image frame; the texture information of the image is kept intact during the correction process, and the geometric distortion error caused by the perspective transformation of the image acquisition platform is eliminated; All image frames with corrected geometric structures, their corresponding timestamps, and perspective estimation vectors are combined into a sequence of corrected image frames.
4. The method for processing crack images based on neural network and plane geometric image transformation according to claim 3, characterized in that: S5 further includes: Jointly encode the image frame with corrected geometric structure, the set of image coordinates of the image markers, and the perspective estimation vector into a multi-modal input tensor; Input the multi-modal input tensor into the deep recognition network to perform multi-channel feature encoding and spatial alignment processing. The multi-channel feature encoding and spatial alignment processing fuse the feature information of the image texture channel, the image marker coordinate channel, and the perspective estimation channel, and perform spatial feature alignment and consistency modeling of crack region localization; After completing the multi-channel feature encoding and spatial alignment processing, the deep recognition network outputs the image mask of the crack region, the set of image coordinates of the crack centerline, and the set of image coordinates of the crack boundary; extract the set of image coordinates of the crack region based on the image mask as the spatial alignment result obtained after fusing the image texture channel, the image marker coordinate channel, and the perspective estimation channel, and input this set of image coordinates into the subsequent spatial back-projection function to construct a three-dimensional positioning relationship; S6 further includes: Use the set of image coordinates of the crack region, the imaging parameters of the image acquisition device carried by the image acquisition platform corresponding to the image frame, and the perspective estimation vector to construct a spatial back-projection function, and input the set of image coordinates of the crack region into the spatial back-projection function to restore the actual position of the crack in three-dimensional space; Extract the spatial geometric attributes of the crack based on its three-dimensional spatial position, where the spatial geometric attributes include crack curvature, crack direction, and crack length; write the three-dimensional spatial position and spatial geometric attributes of the crack, as well as the image frame index and recognition confidence information corresponding to the image frame, into the structured geometric expression recording unit of the image acquisition platform. Generate a three-dimensional visual spectrum of the crack based on the structured geometric expression, and output a visualization result file for inspection and diagnosis.
5. The crack image processing method based on neural network and planar geometric image transformation according to claim 4, wherein: Construct a non-linear inverse transformation function of the image frame by introducing a horizontal inverse transformation function and a vertical inverse transformation function. The horizontal inverse transformation function is expressed as: ; The vertical inverse transformation function is expressed as: ; Wherein is the horizontal remapping coordinate of the pixel point in the image frame after non - linear inverse transformation; is the vertical remapping coordinate of the pixel point in the image frame after non - linear inverse transformation; represents the horizontal and vertical coordinates of the pixel to be inverse - transformed in the original image frame; is the input horizontal coordinate and is the scaling or rotation transformation coefficient for the horizontal output coordinate; is the input vertical coordinate and is the shear or rotation coupling coefficient for the horizontal output coordinate; is the input horizontal coordinate and is the shear or rotation coupling coefficient for the vertical output coordinate; is the input vertical coordinate and is the scaling or rotation transformation coefficient for the vertical output coordinate; is the translation component controlling the horizontal output in the affine transformation; is the translation component controlling the vertical output in the affine transformation; represents the actual offset of the image marker in the horizontal direction in the image frame; represents the actual offset of the image marker in the vertical direction in the image frame; , , respectively represent the estimated vector components corresponding to the pitch angle, roll angle, and yaw angle of the image acquisition platform when acquiring the current image frame.
6. The crack image processing method based on neural network and planar geometric image transformation according to claim 5, wherein: Construct a multi-channel attention fusion tensor , through represents the spatially aligned feature tensor after fusing the image texture channel, the image marker coordinate channel, and the view angle estimation channel; ; Among them represents the image texture features extracted from the geometrically corrected image frames; the structural feature channels constructed from the set of image marker coordinates; represents the view estimation channels constructed by the view estimation vector ; where each channel feature is fused after being gated by the view-driven weights ; wherein denotes a single component in; is non-linear compression; Based on Construct a coordinate residual alignment kernel function; ; Among them represents the coordinate residual alignment kernel function; is the set of crack centerline image coordinates output by the deep network; is the set of crack boundary image coordinates output by the deep network; is the position of the previous image marker in the image; The function for extracting the set of image coordinates of the crack area under the mask guidance is expressed as: ; wherein is the crack area image mask output by the deep network; is used to construct a periodic discrimination field; the response area is screened by mask weighting, and finally the image coordinate set of the crack area is output .
7. The crack image processing method based on neural network and planar geometric image transformation according to claim 6, wherein: According to the multi-channel fusion feature tensor Construct a crack image mask , which is jointly modeled by a view-driven attention activation term and a boundary-sensitive inverse response term; ; wherein is a feature coupling adjustment factor; is a non-linear response compression factor.
8. The crack image processing method based on neural network and planar geometric image transformation according to claim 7, wherein: When constructing the spatial back-projection function the set of image coordinates of the crack region extracted in the image domain is combined with the imaging parameters of the image acquisition device and the perspective estimation vector of the image acquisition platform and is mapped back to the three-dimensional world coordinate system to obtain the actual spatial position of the crack ; Based on three spatial geometric property extraction functions are constructed, which are respectively used to calculate the direction angle of the crack , the curvature of the crack , and the spatial length of the crack , providing an input basis for structured recording and three-dimensional visualization atlas; ; ; ; ; Among them represents the three-dimensional space coordinate point corresponding to the th pixel in the crack region; , represents the lateral difference component of the th spatial difference vector in the three-dimensional space; represents the longitudinal difference component of the th spatial difference vector in the three-dimensional space; represents the height difference component of the th spatial difference vector in the three-dimensional space; represents the th image coordinate extracted from the crack mask in the image frame; represents the component in the height direction of the three-dimensional space coordinate; represents the arc length parameter along the crack path; Among them represents the actual length of the entire crack in three-dimensional space; represents the number of all back-projected points in the crack area image coordinate set.
Citation Information
Patent Citations
Concrete bridge crack detection method based on computer vision
CN119130986A
Generalized neural radiation field reconstruction method based on multi-modal information fusion
CN119359934A