Building structure crack intelligent detection method based on machine vision

By combining AR glasses with SLAM and lightweight convolutional neural networks, three-dimensional identification and prediction of cracks in building structures are achieved, solving the problem of difficulty in correlating the spatial distribution of cracks in traditional detection methods, and improving detection efficiency and the accuracy of safety assessments.

CN120823592AInactive Publication Date: 2025-10-21WUHAN CITY VOCATIONAL COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511004463.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional building structure crack detection methods cannot intuitively reflect the spatial distribution and correlation of cracks, which brings inconvenience to subsequent analysis and maintenance decisions.

Method used

AR glasses are used for scanning, and SLAM technology is combined to build a three-dimensional spatial coordinate system. A lightweight convolutional neural network model is used for crack classification. Combined with speech recognition and voiceprint positioning technology, geometric features are extracted through image processing algorithms, and a crack prediction model is built for extended prediction.

Benefits of technology

It achieves efficient and accurate identification and prediction of cracks in building structures, provides intuitive visual feedback and voice broadcast, improves detection efficiency and the accuracy of safety assessment, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a building structure crack intelligent detection method based on machine vision, and relates to the technical field of detection. The method comprises the steps that AR glasses are used for scanning the surface of a building, an RGB image is obtained and denoised, and then a three-dimensional space coordinate system is constructed through the SLAM technology; based on a lightweight convolutional neural network model, constructing a crack classification semantic model, inputting a de-noised image, and outputting a crack semantic tag and a center point coordinate; training an end side voice RNNT model, converting the voice instruction into a retrieval instruction, and determining a crack corresponding to the instruction through a voiceprint positioning technology; according to the crack center point coordinate and a crack corresponding to the instruction, utilizing an image processing algorithm to extract a crack geometric feature; and constructing a crack prediction model, inputting geometric features, outputting a crack propagation prediction result, and marking a high-risk region on an AR glasses interface by color coding. And building structure crack detection and risk assessment are realized through the crack semantic tag and the crack propagation prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detection engineering, and in particular to a machine vision-based intelligent detection method for cracks in building structures. Background Art

[0002] In the field of building structure safety testing, cracks are an important indicator for assessing the health status of building structures. Their existence and development may reflect problems such as changes in internal structural stress, material aging, or foundation settlement. If not discovered and handled in a timely manner, they may cause structural safety hazards.

[0003] Traditional crack detection in building structures is primarily performed with the assistance of optical instruments. A magnifying glass with a scale is used to directly read the crack width, a ruler is used to measure the length, and the crack location and shape are recorded. However, the crack location and shape recorded using optical instrument-assisted detection are difficult to correlate with the building's three-dimensional model, and cannot intuitively reflect the spatial distribution and correlation of the cracks, which brings inconvenience to subsequent analysis and maintenance decisions. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides an intelligent detection method for building structure cracks based on machine vision, which solves the problem that traditional methods cannot intuitively reflect the spatial distribution and correlation of cracks, causing inconvenience to subsequent analysis and maintenance decisions.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for intelligent detection of cracks in building structures based on machine vision, comprising the following steps: Step S1: Use AR glasses to scan the surface of the building structure to obtain an RGB image of the building, perform denoising on the RGB image to obtain a denoised RGB image of the building, use SLAM technology to match the denoised RGB image of the building, and construct a three-dimensional spatial coordinate system of the building structure; Step S2: Based on the lightweight convolutional neural network model and the three-dimensional spatial coordinate system of the building structure, a crack classification semantic model is constructed. The denoised building RGB image is input into the crack classification semantic model, and the semantic label and center point coordinates of the crack are output; Step S3: Train the device-side voice RNNT model to obtain a trained device-side voice RNNT model. Use the trained device-side voice RNNT model to convert the tester's voice command into an executable search command. Simultaneously, use voiceprint positioning technology to locate the voice command and determine the crack corresponding to the executable search command. Step S4: Based on the coordinates of the crack center point and the crack corresponding to the instruction, extract the features of the crack corresponding to the instruction through an image processing algorithm, and calculate the geometric features of the crack; Step S5: Construct a building structure crack prediction model, input the geometric features of the cracks into the building structure crack prediction model, output a crack extension prediction result, and mark high-risk areas with color coding on the AR glasses interface based on the crack extension prediction result.

[0006] Preferably, the step of scanning the surface of the building structure using AR glasses to obtain a building RGB image, and performing denoising on the building RGB image to obtain the denoised building RGB image comprises: RGB camera: AR smart glasses are equipped with dual RGB cameras with a resolution of ≥1080p and a frame rate of ≥30fps, used to acquire images of building surfaces. Inertial measurement unit: used to obtain the device attitude and assist in determining the device's position and attitude in space; Image denoising: Get denoised building RGB image ; Non-local mean filtering is used to remove stains and noise from building surface images, improve image quality, and reduce interference in subsequent processing. The calculation formula is as follows:

[0007] in, Represents the RGB color value of the image at pixel x after non-local mean filtering; is a normalization coefficient used to ensure that the sum of all weights is 1; is the weight function; It is the search window, which represents a neighborhood range centered on pixel point x in the image, and represents the original RGB color value of the image at pixel point y.

[0008] Preferably, SLAM technology is used to match the denoised building RGB images to construct a three-dimensional spatial coordinate system of the building structure, including: Definition of the three-dimensional space coordinate system of the building structure: With the optical center of the camera in the first mapping frame as the origin, establish a right-handed coordinate system: X-axis: right side of the image; Y-axis: bottom side of the image; Z-axis: scene depth.

[0009] Preferably, the construction of the crack classification semantic model based on the lightweight convolutional neural network model and the three-dimensional space coordinate system of the building structure includes: Network architecture: Built on the lightweight convolutional neural network model MobileNetV3, including a feature extraction backbone network and a multi-task output head; Among them, the feature extraction backbone network also includes feature modulation, and the calculation formula of the feature modulation is as follows:

[0010] in, It is the output feature map after attention modulation; the attention enhancement feature is obtained after element-by-element multiplication , is a spatial attention map, which is used to weight the spatial positions in the feature map; For feature splicing, the results of average pooling and maximum pooling are spliced ​​together; It is a channel attention map, which is used to weight the features of different channels; The multi-task output header is specifically: based on Output the results of the three subtasks in parallel, sharing the backbone network parameters but optimizing them independently. Specifically: Branch 1: Crack existence judgment, 1×1 convolution: compresses the channel dimension and generates existence probability:

[0011] in, is the probability of output crack existence, where is the convolution kernel parameter; Branch 2: Crack type classification, Softmax classification: feature Apply the Softmax function to get the probability distribution of each category The values ​​are converted to probability values ​​so that the sum of the probabilities of all categories is 1. The formula is as follows:

[0012] in is the output type probability distribution; For the Category channel features, i category channel feature index; the process of extracting crack direction features through 3x1 convolution kernel and using Softmax function for classification; Branch 3: Crack width level judgment, Softmax probability mapping:

[0013] in, It is the output level probability distribution, and the width level judgment depends on the pixel span of the crack extending vertically in the feature map.

[0014] Preferably, the outputting of the semantic label and center point coordinates of the crack includes: Three-dimensional coordinate transformation: using the camera intrinsic parameter matrix And the inverse projection transformation:

[0015] in, is the homogeneous coordinate of the point in the three-dimensional space in the camera coordinate system, It is the inverse matrix of the camera's intrinsic parameters. Inverse projection is the process of inferring the camera coordinates from the pixel coordinates. is the depth value of the crack center point, which is finally expressed in the SLAM coordinate system as , is the non-homogeneous coordinate vector of the crack center in the camera coordinate system, and is the intermediate result of the subsequent coordinate system transformation; Coordinate system alignment will Convert to world coordinate system:

[0016] in, is the coordinate vector of the crack center point in the world coordinate system, is the camera pose matrix of the current frame; Label output: Package the crack center coordinates and semantic labels into a structure:

[0017] As the output result of building a crack classification semantic model.

[0018] Preferably, converting the detection person's voice command into an executable retrieval command by using the trained terminal-side voice RNNT model includes: The inspector inputs commands through the voice interface of AR glasses, and the system first pre-processes the voice signal: Signal acquisition: The built-in microphone of AR glasses collects audio streams at a sampling rate of 16kHz to ensure coverage of the main frequencies of human voice; Framing and feature extraction: The continuous audio stream is divided into 25ms windows, and the 80-dimensional MFCC features of each frame are extracted to capture the rhythm and phoneme information of the speech. Device-side voice RNNT model: The preprocessed speech features are input into the RNNT model, which then analyzes the command intent through the following modules: Encoder: Its structure is: 8-layer bidirectional LSTM with a hidden state dimension of 512, which is used to capture the temporal dependency of speech; Its function is to convert the speech feature sequence Convert to high-dimensional contextual representation , preserve the time series information of speech; Prediction Network: Its structure is a single-layer LSTM with a vocabulary size of 1000; Its function is: based on the historical label sequence Predict possible labels for the current frame , output probability distribution ; Joint Network: Its structure is: fully connected layer, output dimension 256, context representation of fusion encoder and the label probability of the prediction network , generate joint features , the characteristics Can be converted into executable retrieval instructions.

[0019] Preferably, the method of locating the voice command using voiceprint positioning technology and determining the crack corresponding to the executable search command includes: Determine the crack corresponding to the sound source: Delay Difference Calculation: The four-element linear microphone array on the AR glasses synchronously collects the sound source signal and calculates the delay difference between any two elements:

[0020] in, is the time delay difference between any two array elements, is the distance between microphone array elements i and j, is the azimuth of the sound source; is the sound source pitch angle; ; Azimuth and elevation angle estimation: Estimation by generalized cross-correlation algorithm , combined with triangulation to calculate the sound source :

[0021] Combined with the sound source localization results ( , ) and the field of view of AR glasses, the system selects the cracks to be retrieved: Field of view mapping: The spherical coordinates of the sound source ( , , ) is converted to the local coordinate system of the AR glasses to determine whether it falls within the field of view:

[0022] Maintain a crack database for the current detection scenario and filter target cracks using the following rules:

[0023] in, is the sound source position, For the The location of the crack.

[0024] Preferably, the crack corresponding to the instruction is extracted using an image processing algorithm, and the geometric features of the crack are calculated to include: Width measurement: at the center of the crack Sample 10 points in the vertical direction and calculate the maximum distance:

[0025] in, is the crack width, and are the three-dimensional coordinates of the left and right edge points of the crack in the qth group of sampling points, is the Euclidean norm, which represents the straight-line distance between two points; Length calculation: The crack length should be the straight-line distance between the two end points of the crack, based on the coordinates of the first and last endpoints extracted from the crack skeleton line. and , the calculation formula is:

[0026] in, is the crack length, and are the starting point and end point coordinates of the crack skeleton line, respectively, based on the three-dimensional space coordinate system constructed in step S1, in millimeters; Tilt angle calculation: skeleton line direction vector , angle with the horizontal plane:

[0027] Where α is the crack inclination angle, which is the acute angle between the vector and its projection vector on the horizontal plane; is the vertical component that determines the change in the "height" of the tilt, It is the horizontal projection length, reflecting the degree of "extension" of the vector on the horizontal plane.

[0028] Preferably, the constructing of a building structure crack prediction model includes: Training set: The training set needs to provide the model with labeled historical data to ensure the accuracy and generalization of the prediction. The specific composition is as follows: Sample source: Contains historical crack detection data of different types of building structures; Sample features: Each sample is structured data, including input features ( ) and the corresponding label, that is, the actual length extension of the sample in the same period of history, which is directly obtained from the historical detection records; Sample size and diversity: The sample size is ≥ 10,000, covering different crack types, width levels, and building materials to ensure the model's adaptability to various scenarios; Spatiotemporal data alignment and prediction input: Data query: According to the current crack center point , spatially adjacent points are retrieved in the historical database, and the time window is set to 1 year; Feature Engineering: Extracting Input Feature Matrix :

[0029] in, is the crack width, is the crack length, It's a crack , is the crack depth; Model selection: This model uses the XGBoost regressor, which minimizes the loss function by iteratively adding trees. It has the ability to handle nonlinear relationships and high-dimensional data, supports parallel computing, and is suitable for deployment on end devices. Training objective function:

[0030] Where N is the number of samples, are the parameters of the model, is the true value of the i-th sample, is the model’s predicted value for the i-th sample, It is a regularization term, including L1 and L2 penalties, which are used to prevent overfitting. L1 penalty will cause the coefficients of some features to become 0, thereby playing the role of feature selection. L2 penalty will make the coefficients of the features smaller, but will not become 0, thereby playing the role of reducing the model variance. The purpose of minimizing the loss function is to minimize the predicted value. and the true value The difference between them is minimized by minimizing the sum of squared errors of all samples to make the model's prediction as accurate as possible; Prediction output: Crack expansion distance in the next three months .

[0031] Preferably, based on the crack propagation prediction results, high-risk areas are marked with color codes on the AR glasses interface, including: Perform AR dynamic risk classification warning:

[0032] Visualization implementation: superimpose translucent color blocks on the crack area, add flashing animations on the edges, and trigger voice alarms synchronously.

[0033] Beneficial effects The present invention provides a method for intelligent detection of cracks in building structures based on machine vision. It has the following beneficial effects: (1) This machine vision-based intelligent detection method for building structure cracks uses AR glasses for real-time scanning and SLAM technology to construct a three-dimensional coordinate system. Combined with the dynamic attention mechanism of a lightweight convolutional neural network, it can automatically focus on the crack area, reduce interference from irrelevant information, improve the efficiency and accuracy of crack identification, and avoid the subjectivity and omissions of manual detection.

[0034] (2) This intelligent detection method for building structure cracks based on machine vision directly superimposes the geometric parameters of the cracks on the crack location through AR visual feedback, and cooperates with voice broadcasting, so that the inspectors can obtain key data intuitively and quickly without having to view complex reports, thereby improving the efficiency of information acquisition; color coding marks high-risk areas, so that inspectors can clearly identify areas that need special attention, facilitating timely response measures.

[0035] (3) This machine vision-based intelligent detection method for building structure cracks can predict the development trend of cracks based on existing geometric parameters through a crack expansion prediction model, providing data support for the maintenance and safety assessment of building structures, helping to formulate maintenance plans in advance, avoid safety accidents caused by crack deterioration, and reduce maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flow chart of a method for intelligent detection of cracks in building structures based on machine vision proposed by the present invention.

[0037] Figure 2 This is a hierarchical diagram of the crack classification semantic model in the machine vision-based intelligent detection method for building structure cracks proposed by the present invention.

[0038] Figure 3 This is a hierarchical diagram of the crack propagation prediction results obtained in the machine vision-based intelligent detection method for building structure cracks proposed by the present invention. DETAILED DESCRIPTION

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0040] See also Figure 1 The present invention provides a technical solution: a method for intelligent detection of cracks in building structures based on machine vision. Specifically, the following method for intelligent detection of cracks in building structures based on machine vision is provided. Figure 1 , the method comprises the following steps: Step S1: Use AR glasses to scan the surface of the building structure to obtain an RGB image of the building, perform denoising on the RGB image to obtain a denoised RGB image of the building, use SLAM technology to match the denoised RGB image of the building, and construct a three-dimensional spatial coordinate system of the building structure.

[0041] RGB camera: AR smart glasses are equipped with dual RGB cameras with a resolution ≥ 1080p and a frame rate ≥ 30fps, used to acquire images of building surfaces I.

[0042] Inertial Measurement Unit (IMU): Used to obtain the device's attitude and assist in determining the device's position and attitude in space.

[0043] Image denoising: Get denoised building RGB image ; Non-local mean filtering (NLM) is used to remove stains and noise from building surface images, improve image quality, and reduce interference in subsequent processing. The calculation formula is as follows:

[0044] in, Represents the RGB color value of the image at pixel x after non-local mean filtering; is a normalization coefficient used to ensure that the sum of all weights is 1; is the weight function; is the search window, which represents a neighborhood range centered on pixel point x in the image, and represents the original RGB color value of the image at pixel point y; The weight function represents the similarity weight between pixel points x and y. The calculation formula is:

[0045] in, ( ) is an exponential function, representing the power of the base of the natural logarithm e (approximately equal to 271828). In the calculation of the weight function, it is used to calculate the similarity weight between pixels x and y; It is the square of the Euclidean distance between the RGB color value vectors of pixel points x and y, which measures the difference between the two. h is the smoothing intensity control parameter. The larger the value, the stronger the denoising effect. The purpose of the function is to measure the similarity of the color values ​​of two pixels. The closer the color values, the greater the weight.

[0046] SLAM three-dimensional coordinate system construction: Feature point extraction: The ORB (Oriented FAST and Rotated BRIEF) algorithm is used to detect image corners, and the descriptor dimension is 256 bits. The FAST key point detection and BRIEF descriptor extraction used in the ORB algorithm are two main components, and their purpose is to quickly find and describe key points in the image.

[0047] The feature point extraction step itself usually does not involve complex mathematical formulas, but is based on some classic algorithms in image processing and computer vision.

[0048] FAST keypoint detection: corners are determined by checking the brightness difference between a pixel and its neighborhood. For each pixel in the image, its brightness relationship with the surrounding 16 neighboring pixels is checked. If the pixel is brighter or darker than N pixels (usually 11) in its neighborhood, the point is considered a corner. BRIEF descriptor extraction: The BRIEF algorithm is a method for extracting image feature descriptors. It generates a binary string as a descriptor. For each key point, a pair of pixels (called a point pair) is randomly selected; the grayscale values ​​of the pair of pixels are compared. If the grayscale value of the first pixel is greater than that of the second pixel, 1 is written to the descriptor, otherwise 0 is written; this process is repeated multiple times (for example, 256 times) to generate a 256-bit binary descriptor.

[0049] Relative pose of adjacent frames: By comparing the descriptors of key points in different images, we can find the matching point pairs between the two frames, solve the relative pose of the camera from frame i to frame j (rotation R and translation t), and minimize the reprojection error. By solving the fundamental matrix F through epipolar geometry, the reprojection error is minimized:

[0050] in, , are the coordinates of the matching points in the two frames; is the optimization variable; F is the basic matrix, describing the epipolar geometric relationship between the two frames; is a set of matching point pairs; is the camera intrinsic parameter matrix, is the relative rotation matrix between two frames; is the relative translation vector between the two frames; the geometric relationship between the matching point pairs is found by solving the fundamental matrix F. The purpose is to determine the rotation and translation of the camera when it moves from one frame to the next.

[0051] Global optimization: Jointly optimizing the camera pose and 3D point coordinates using the Bundle Adjustment (BA) method can improve the consistency and accuracy of the entire map and reduce cumulative error, where the map is a 3D representation of the environment built as you move around it. Objective function: variables include the poses of all frames and all 3D point coordinates .

[0052]

[0053] Among them, the first term is the reprojection error, k is the key frame index, M represents the total number of feature points detected in all frames, and m is used to traverse all detected feature points. is the pixel coordinate (observation value) of the mth 3D point observed in the kth frame, Is the projection function, 3D point (Coordinates in the camera coordinate system) are projected to the The pixel plane of the frame, this error measures the difference between the projection of the 3D point in the camera coordinate system and the actual observed pixel coordinates, reflecting the joint error of the pose and the 3D point coordinates; the second term is the pose smoothness constraint, is the regularization parameter, is the Frobenius norm of the matrix; this constraint requires that the rotation matrices of adjacent frames and As close as possible (the smaller the norm, the smaller the rotation difference), avoiding sudden changes in posture; Definition of the three-dimensional space coordinate system of the building structure: With the optical center of the camera in the first mapping frame as the origin, a right-handed coordinate system is established: X-axis: right side of the image (u-axis of the pixel coordinate system); Y-axis: bottom side of the image (v-axis of the pixel coordinate system); Z-axis: scene depth direction (in the same direction as the camera optical axis).

[0054] This step provides basic image data and spatial positioning references for all subsequent steps. The image is acquired and preprocessed through the hardware equipment of the AR smart glasses, and then a three-dimensional coordinate system is constructed through SLAM technology to ensure that subsequent crack detection, measurement, and visualization are based on a unified spatial reference.

[0055] Step S2: Based on the lightweight convolutional neural network model and the three-dimensional spatial coordinate system of the building structure, a crack classification semantic model is constructed. The denoised building RGB image is input into the crack classification semantic model, and the semantic label and center point coordinates of the crack are output.

[0056] Build a semantic model for crack classification: Network architecture: Built on the lightweight convolutional neural network model MobileNetV3, including a feature extraction backbone network and a multi-task output head: Feature extraction backbone network (depthwise separable convolution): Initial feature extraction ( Depthwise Separable Convolution + Batch Normalization):

[0057] in, represents the result of initial feature extraction, It refers to depthwise separable convolution, and BatchNorm refers to batch normalization, which is a commonly used technique in training deep neural networks to speed up the training process and improve the stability of the model. is the denoised building RGB image, which is the image input. Indicates downsampling with a step size of 2, output feature size , is the number of channels of the output feature map; Feature enhancement module (MobileNetV3Block):

[0058] in, To output the feature map, the module extracts and enhances features through dynamic convolution kernels and SE attention mechanism. The final output feature map will be further downsampled to a resolution of 480×270.

[0059] Global feature aggregation (average pooling):

[0060] Among them, the feature map F2 is globally averaged and pooled to obtain To output channel-level features, the formula aims to globally average the spatial dimensions and preserve the global semantics of the crack.

[0061] The dynamic attention mechanism (CBAM) embeds a channel-space attention module after each layer of the backbone network to enhance the response of key areas: Channel Attention:

[0062] in, It is a channel attention map, which is used to weight the features of different channels; MLP is a multi-layer perceptron (MultiLayerPerceptron), which is used to perform nonlinear transformation on features; For feature splicing, the results of average pooling and maximum pooling are spliced ​​together; is the Sigmoid activation function.

[0063] Spatial Attention:

[0064] in, is a spatial attention map, which is used to weight the spatial positions in the feature map; It is a 3x3 convolution layer, which is used to perform convolution operations on the concatenated features; Used to adjust the order of feature channels and enhance spatial local dependency modeling, and Perform global average pooling and global maximum pooling on the feature maps after adjusting the order.

[0065] Characteristic modulation:

[0066] in, It is the output feature map after attention modulation; the attention enhancement feature is obtained after element-by-element multiplication , focusing on discriminative areas such as crack edges and textures.

[0067] Multi-task output head, based on Output the results of the three subtasks in parallel, sharing the backbone network parameters but optimizing them independently: Branch 1: Crack existence judgment (two-classification), 1×1 convolution: compresses the channel dimension and generates existence probability:

[0068] in, is the probability of output crack existence, where is the convolution kernel parameter.

[0069] Branch 2: Crack type classification (three categories) 3×1 longitudinal convolution: focusing on the characteristics of the crack extension direction

[0070] in, Is the output feature, for the input feature map Apply a 3x1 convolution kernel with weights , the bias is Here, the convolution kernel slides along the height direction, assuming that the crack extends longitudinally, thereby extracting the characteristics of the crack direction, and the number of output channels is 3 (structural / non-structural / undefined).

[0071] Softmax classification: features Apply the Softmax function to get the probability distribution of each category The values ​​are converted to probability values ​​so that the sum of the probabilities of all categories is 1. The formula is as follows:

[0072] in is the output type probability distribution; For the Categorical channel features, i is an index of categorical channel features; the process of extracting crack direction features using a 3x1 convolution kernel and classifying them using the Softmax function. This method helps identify and classify different types of cracks in image processing tasks such as crack detection, thereby improving the accuracy and robustness of the model.

[0073] Branch 3: Crack width grade judgment (three categories) Core logic: Decoupling width-sensitive features to achieve grading (micro-cracks / moderate cracks / severe cracks):

[0074] in, is the width feature, and They are the weight and bias of the convolution kernel respectively; the convolution kernel is designed to adapt to the pixel span in the crack width direction (combined with the input resolution Calibrate with point cloud density to establish pixel width The actual physical width The mapping: ,in is the calibration factor).

[0075] Softmax probability mapping:

[0076] in, is the output level probability distribution, where Corresponding to micro cracks, otherwise judged by threshold; Physical meaning: Width level judgment depends on the pixel span of the crack extending vertically in the feature map (such as micro cracks corresponding to pixel span , middle split ,serious , with the specific threshold determined by training data statistics). This module can convert the crack width in the image into actual physical size and determine the severity of the crack based on the probability distribution, thereby achieving automatic detection and assessment of cracks.

[0077] The multi-task loss function jointly optimizes three branches to balance classification accuracy:

[0078] in, is the multi-task loss function, is the binary cross entropy loss (crack existence) is the multi-class cross entropy loss (type / width level), 、 They are the task weight coefficients for crack existence judgment, crack type classification and crack width level judgment (tuned through the validation set, usually to ensure that basic tasks take priority).

[0079] Center point coordinate mapping: Image-level positioning, calculation of crack geometric center:

[0080] in, is the horizontal (column) coordinate of the crack center in the image, is the vertical (row direction) coordinate of the crack center in the image, is the set of crack pixels, Represents the number of elements in the crack pixel set, that is, the total number of pixels contained in the crack; and Respectively represent the sum of the u coordinates and v coordinates of all pixels in the crack pixel set; this formula calculates the geometric center of the crack by the arithmetic mean method. In essence, it is to perform a weighted average of the coordinates of all pixels on the crack to obtain the feature point representing the center position of the crack.

[0081] Three-dimensional coordinate transformation: Using the camera intrinsic parameter matrix And the inverse projection transformation, the formula is as follows:

[0082] in, is the homogeneous coordinate of the point in the three-dimensional space in the camera coordinate system, It is the inverse matrix of the camera's intrinsic parameters. Inverse projection is the process of inferring the camera coordinates from the pixel coordinates. is the depth value of the crack center point (estimated by monocular vision), which is finally expressed in the SLAM coordinate system as , is the non-homogeneous coordinate vector of the crack center in the camera coordinate system, and is the intermediate result of the subsequent coordinate system transformation; Coordinate system alignment will Convert to world coordinate system:

[0083] in, is the coordinate vector of the crack center point in the world coordinate system, is the camera pose matrix of the current frame (output by SLAM in real time).

[0084] Label output: Package the crack center coordinates and semantic labels into a structure:

[0085] As the output result of constructing the crack classification semantic model, structured crack semantic information is provided. Based on the label output, a crack database of the detection scene is constructed, namely the center point coordinates, type, and width of each crack, providing a retrieval basis for the voice command parsing in step S3.

[0086] Step S3: Train the end-side voice RNNT model to obtain a trained end-side voice RNNT model. Use the trained end-side voice RNNT model to convert the detection person's voice instructions into executable retrieval instructions. At the same time, use voiceprint positioning technology to locate the voice instructions and determine the cracks corresponding to the executable retrieval instructions.

[0087] On-device voice RNNT model training set: Speech: Speech samples containing instructions related to crack detection and analysis, expressed in natural language, such as "measure maximum crack width," "analyze crack depth distribution," and "find transverse cracks." These speech samples should be diverse, covering different accents, speaking speeds, and intonations, and recorded in the presence of background noise (especially ambient noise that may be associated with crack detection equipment).

[0088] Text / label section: Standardized text labels or instruction sequences corresponding to the aforementioned voice commands. These labels include not only the command word itself (e.g., "width," "length"), but also the series of sub-commands, parameters, or opcodes required to perform the complex task. For example, "measure the maximum crack width" might be parsed into a series of operations: [locate cracks, identify the largest crack, extract its width features, calculate and report the result].

[0089] The inspector enters a command (such as "measure the maximum crack width") through the voice interface of the AR glasses. The system first preprocesses the voice signal: Signal acquisition: The built-in microphone in AR glasses collects audio streams at a 16kHz sampling rate to ensure coverage of the main frequencies of human voice (300Hz~3400Hz).

[0090] Framing and feature extraction: The continuous audio stream is divided into 25ms windows (with a frame shift of 10ms) and 80-dimensional MFCC features (Mel-frequency cepstral coefficients) are extracted from each frame to capture the rhythm and phoneme information of the speech.

[0091] Device-side voice RNNT model: The preprocessed speech features are input into the RNNT model, which then analyzes the command intent through the following modules: Encoder (Bidirectional LSTM): Structure: 8-layer bidirectional LSTM with a hidden state dimension of 512, used to capture the temporal dependencies of speech (such as the sequential relationship between keywords such as "measurement", "maximum", and "crack width").

[0092] Function: Sequence speech features ( (number of frames) into a high-dimensional context representation , preserving the time series information of speech.

[0093] Prediction network (single-layer LSTM): Its structure is: single-layer LSTM, vocabulary size 1000 (including instruction-related words such as "measurement", "width", "length", "crack", "maximum", etc.), based on historical label sequence (Identified instruction fragment) predict the possible label of the current frame , output probability distribution .

[0094] Joint network (fully connected layer): Its structure is: fully connected layer, output dimension 256, context representation of fusion encoder and the label probability of the prediction network , generate joint features , the characteristics Can be converted into executable retrieval instructions ( It can be used directly as a query vector for ANN retrieval).

[0095] Loss function and training: Training objective: minimize negative log-likelihood loss and optimize model parameters:

[0096] in, is the loss function of the voice command parsing model, is the total number of time steps of the input sequence (such as speech features); It is before Input features of time steps (such as acoustic features of speech frames) It is before The predicted output of time steps (used to optimize the current prediction using historical information, similar to language models); Before it is known Input and front In the case of predictions, the current time step prediction is probability.

[0097] Inference optimization: CTC BeamSearch is used to accelerate decoding, and beam width control (for example, Beam = 5) is used to balance accuracy and latency, ensuring that the end-side inference latency is ≤ 300ms.

[0098] Determine the crack corresponding to the sound source: Time Delay Difference (TDOA) calculation AR glasses equipped with a 4-element linear microphone array (element spacing ) Synchronously collect the sound source signal and calculate the time delay difference between any two array elements:

[0099] in, is the time delay difference between any two array elements, is the distance between microphone array elements i and j, is the azimuth of the sound source (angle with the x-axis); is the sound source pitch angle (angle with the y-axis); (Speed ​​of sound).

[0100] Azimuth and elevation angle estimation: Estimation by generalized cross-correlation (GCCPHAT) algorithm , combined with triangulation to calculate the sound source :

[0101] (Assume that array elements 1, 2, and 3 are in the x-axis, y-axis, and z-axis directions respectively) Combined with the sound source localization results ( , ) and the field of view of AR glasses (FOV = 90° × 60°), the system selects the cracks to be retrieved: Field of view mapping: The spherical coordinates of the sound source ( , , )( The distance to the sound source) is converted into the local coordinate system of the AR glasses (with the optical center of the glasses as the origin) to determine whether it falls within the field of view:

[0102] Maintain a crack database for the current detection scene (including the IDs, locations, and ), filter target cracks using the following rules:

[0103] in, is the sound source position, For the The position of the crack belongs to step S2 , and feed back the matched target crack ID (or location coordinates) to the AR interface. The inspector can confirm it through voice or touch, triggering the subsequent crack measurement and analysis process.

[0104] This step parses the inspector's voice commands (such as "measure the maximum crack width" and "analyze horizontal cracks") by building an end-side voice RNNT model, converting natural language commands into system-executable retrieval commands. At the same time, combined with voiceprint positioning technology (using the microphone array's time delay difference calculation and triangulation positioning method), the specific crack area associated with the voice command is determined, ensuring that subsequent steps (such as the parameter calculation in step S4) can accurately point to the cracks of interest to the inspector, reducing manual operation costs and improving detection efficiency.

[0105] Step S4: Based on the coordinates of the crack center point and the crack corresponding to the instruction, feature extraction is performed on the crack corresponding to the instruction through an image processing algorithm to calculate the geometric features of the crack.

[0106] Crack geometry feature extraction: Width measurement: at the center of the crack Sample 10 points along the vertical direction (gradient direction) and calculate the maximum distance:

[0107] in, is the crack width, and are the three-dimensional coordinates of the left and right edge points of the crack in the qth group of sampling points, is the Euclidean norm (distance formula), which represents the straight-line distance between two points.

[0108] Length calculation: The crack length should be the straight-line distance between the two end points of the crack, based on the coordinates of the first and last endpoints extracted from the crack skeleton line. and (in three-dimensional space), the calculation formula is:

[0109] in, is the crack length, and are the starting and ending coordinates of the crack skeleton line (extracted by the morphological thinning algorithm), based on the three-dimensional space coordinate system constructed in step S1, in millimeters (mm).

[0110] Tilt angle calculation: skeleton line direction vector , angle with the horizontal plane:

[0111] Where α is the crack inclination angle, which is the acute angle between the vector and its projection vector on the horizontal plane; is the vertical component that determines the change in the "height" of the tilt, It is the horizontal projection length, reflecting the degree of "extension" of the vector on the horizontal plane.

[0112] AR visualization feedback: Coordinate system mapping: The center point of the crack in the world coordinate system is mapped to the coordinate system. Convert to AR device coordinate system (device optical center is the origin):

[0113] in, is the coordinate of the crack on the AR device, is the center point of the crack in the world coordinate system; This is the pose transformation matrix from the device coordinate system to the world coordinate system, acquired in real time by the IMU (Inertial Measurement Unit). Through this coordinate mapping, the AR system can "translate" the crack location in the real world to the device's perspective, achieving precise virtual-reality fusion visualization and providing users with intuitive spatial positioning feedback.

[0114] AR Feedback: Ruler rendering, using OpenGLES to draw a red line segment (RGB: 255, 0, 0), starting from , direction along the main direction of the crack, length is , the width is marked as (Pixel width is proportional to actual size); Transparency control: Adjust according to crack level (micro cracks , middle split ,serious ).

[0115] Voice broadcast: Output parameters through TTS engine (iFlytek offline engine): .

[0116] The RNNT model for voice recognition is used to parse the command intent, and the sound source localization module is used to determine the crack location corresponding to the sound source. Finally, the goal of "which crack to search for" is output, thus achieving a closed loop from "voice interaction" to "precise detection."

[0117] Step S5: Construct a building structure crack prediction model, input the geometric features of the cracks into the building structure crack prediction model, output a crack extension prediction result, and mark high-risk areas with color coding on the AR glasses interface based on the crack extension prediction result.

[0118] Constructing a building structure crack prediction model: Training set: The training set needs to provide the model with labeled historical data to ensure the accuracy and generalization of the prediction. The specific composition is as follows: Sample source: Contains historical crack detection data of different types of building structures (such as concrete frames, shear walls, steel structures, etc.); Sample features: Each sample is structured data, including input features ( ) and the corresponding label, that is, the actual length expansion of the sample in the same historical period (such as the past 3 months), which is directly obtained from the historical detection records.

[0119] Sample size and diversity: The sample size is ≥10,000, covering different crack types (structural / non-structural), width levels (micro-cracks / moderate cracks / severe cracks), and building materials (concrete, steel bars, masonry), ensuring the model's adaptability to various scenarios.

[0120] Data query: According to the current crack center point , spatially adjacent points (radius ≤ 50 cm) are retrieved in the historical database, and the time window is set to 1 year.

[0121] Feature Engineering: Extracting Input Feature Matrix :

[0122] in, is the crack width, is the crack length, It's a crack , is the crack depth.

[0123] Model selection: This model uses the XGBoost regressor to minimize the loss function by iteratively adding trees. It has the ability to handle nonlinear relationships and high-dimensional data, supports parallel computing, and is suitable for deployment on end devices.

[0124] Training objective function:

[0125] Where N is the number of samples, are the parameters of the model, is the true value of the i-th sample, is the model’s predicted value for the i-th sample, It is a regularization term, including L1 and L2 penalties, which are used to prevent overfitting. L1 penalty (Lasso) will cause the coefficients of some features to become 0, thereby playing a role in feature selection. L2 penalty (Ridge) will make the coefficients of the features smaller, but will not become 0, thereby playing a role in reducing the model variance. The purpose of minimizing the loss function is to minimize the predicted value. and the true value The difference between them is minimized to make the model's prediction as accurate as possible by minimizing the sum of squared errors of all samples.

[0126] Prediction output: Crack expansion distance in the next three months , confidence intervals (95% CI) were calculated based on error propagation.

[0127] Perform AR dynamic risk classification warning:

[0128] Visualization is achieved by superimposing translucent blocks of color on the crack area, adding a flashing animation on the edge (frequency 2Hz, green→yellow→red gradient), and synchronously triggering a voice alarm: "High-risk cracks, please review immediately."

[0129] The present invention uses the high-resolution camera of AR glasses to scan the building surface in real time, obtains RGB images and performs denoising processing to eliminate interference such as stains and shadows, and then constructs a three-dimensional spatial coordinate system of the building structure through SLAM technology to provide a spatial reference for subsequent detection; a crack classification semantic model is constructed based on a lightweight convolutional neural network model, the denoised RGB images are processed, and the semantic labels and center point coordinates of the cracks are output to achieve accurate identification and classification of cracks; the end-side voice RNNT model is trained to parse the voice commands of the inspector and convert them into executable retrieval commands; at the same time, the specific crack area corresponding to the command is determined by combining voiceprint positioning technology to achieve accurate response of human-computer interaction; based on the retrieval command and the corresponding crack, the geometric features of the crack are extracted through image processing algorithms, and these parameters are superimposed and displayed at the crack position through AR technology, and voice broadcast is performed at the same time to provide intuitive information for the inspector; a building structure crack prediction model is constructed, the crack geometric parameters are input to obtain extended prediction results, and the high-risk areas are marked with color coding on the AR interface according to the results to achieve early warning of crack risks.

[0130] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions. The sentence "including an element defined by..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element."

[0131] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent detection of cracks in building structures based on machine vision, characterized in that: The method comprises the following steps: Step S1: Use AR glasses to scan the surface of the building structure to obtain an RGB image of the building, perform denoising on the RGB image to obtain a denoised RGB image of the building, use SLAM technology to match the denoised RGB image of the building, and construct a three-dimensional spatial coordinate system of the building structure; Step S2: Based on the lightweight convolutional neural network model and the three-dimensional spatial coordinate system of the building structure, a crack classification semantic model is constructed. The denoised building RGB image is input into the crack classification semantic model, and the semantic label and center point coordinates of the crack are output; Step S3: Train the device-side voice RNNT model to obtain a trained device-side voice RNNT model. Use the trained device-side voice RNNT model to convert the tester's voice command into an executable search command. Simultaneously, use voiceprint positioning technology to locate the voice command and determine the crack corresponding to the executable search command. Step S4: Based on the coordinates of the crack center point and the crack corresponding to the instruction, extract the features of the crack corresponding to the instruction through an image processing algorithm, and calculate the geometric features of the crack; Step S5: Construct a building structure crack prediction model, input the geometric features of the cracks into the building structure crack prediction model, output a crack extension prediction result, and mark high-risk areas with color coding on the AR glasses interface based on the crack extension prediction result.

2. The method for intelligent detection of cracks in building structures based on machine vision according to claim 1, characterized in that: The AR glasses are used to scan the surface of the building structure to obtain an RGB image of the building, and the denoising process is performed on the RGB image of the building to obtain a denoised RGB image of the building, including: RGB camera: AR smart glasses are equipped with dual RGB cameras with a resolution of ≥1080p and a frame rate of ≥30fps, used to acquire images of building surfaces. Inertial measurement unit: used to obtain the device attitude and assist in determining the device's position and attitude in space; Image denoising: Get denoised building RGB image ; Non-local mean filtering is used to remove stains and noise from building surface images, improve image quality, and reduce interference in subsequent processing. The calculation formula is as follows: ; in, Represents the RGB color value of the image at pixel x after non-local mean filtering; is a normalization coefficient used to ensure that the sum of all weights is 1; is the weight function; It is the search window, which represents a neighborhood range centered on pixel point x in the image, and represents the original RGB color value of the image at pixel point y.

3. The method for intelligent detection of building structure cracks based on machine vision according to claim 2, characterized in that: SLAM technology is used to match the denoised building RGB images and construct the three-dimensional spatial coordinate system of the building structure, including: Definition of the three-dimensional space coordinate system of the building structure: With the optical center of the camera in the first mapping frame as the origin, establish a right-handed coordinate system: X-axis: right side of the image; Y-axis: bottom side of the image; Z-axis: scene depth.

4. The method for intelligent detection of building structure cracks based on machine vision according to claim 3, characterized in that: The crack classification semantic model is constructed based on the lightweight convolutional neural network model and the three-dimensional spatial coordinate system of the building structure, including: Network architecture: Built on the lightweight convolutional neural network model MobileNetV3, including a feature extraction backbone network and a multi-task output head; Among them, the feature extraction backbone network also includes feature modulation, and the calculation formula of the feature modulation is as follows: ; in, It is the output feature map after attention modulation; the attention enhancement feature is obtained after element-by-element multiplication , is a spatial attention map, which is used to weight the spatial positions in the feature map; For feature splicing, the results of average pooling and maximum pooling are spliced ​​together; It is a channel attention map, which is used to weight the features of different channels; The multi-task output header is specifically: based on Output the results of the three subtasks in parallel, sharing the backbone network parameters but optimizing them independently. Specifically: Branch 1: Crack existence judgment, 1×1 convolution: compresses the channel dimension and generates existence probability: ; in, is the probability of output crack existence, where is the convolution kernel parameter; Branch 2: Crack type classification, Softmax classification: feature Apply the Softmax function to get the probability distribution of each category The values ​​are converted to probability values ​​so that the sum of the probabilities of all categories is 1. The formula is as follows: ; in is the output type probability distribution; For the Category channel features, i category channel feature index; the process of extracting crack direction features through 3x1 convolution kernel and using Softmax function for classification; Branch 3: Crack width level judgment, Softmax probability mapping: ; in, It is the output level probability distribution, and the width level judgment depends on the pixel span of the crack extending vertically in the feature map.

5. The method for intelligent detection of building structure cracks based on machine vision according to claim 4, characterized in that: The output obtains the semantic label and center point coordinates of the crack, including: Three-dimensional coordinate transformation: using the camera intrinsic parameter matrix And the inverse projection transformation: ; in, is the homogeneous coordinate of the point in the three-dimensional space in the camera coordinate system, It is the inverse matrix of the camera's intrinsic parameters. Inverse projection is the process of inferring the camera coordinates from the pixel coordinates. is the depth value of the crack center point, which is finally expressed in the SLAM coordinate system as , is the non-homogeneous coordinate vector of the crack center in the camera coordinate system, and is the intermediate result of the subsequent coordinate system transformation; Coordinate system alignment will Convert to world coordinate system: ; in, is the coordinate vector of the crack center point in the world coordinate system, is the camera pose matrix of the current frame; Label output: Package the crack center coordinates and semantic labels into a structure: ; As the output result of building a crack classification semantic model.

6. The method for intelligent detection of cracks in building structures based on machine vision according to claim 5, characterized in that: The method of converting the voice commands of the detection personnel into executable retrieval commands through the trained end-side voice RNNT model includes: The inspector inputs commands through the voice interface of AR glasses, and the system first pre-processes the voice signal: Signal acquisition: The built-in microphone of AR glasses collects audio streams at a sampling rate of 16kHz to ensure coverage of the main frequencies of human voice; Framing and feature extraction: The continuous audio stream is divided into 25ms windows, and the 80-dimensional MFCC features of each frame are extracted to capture the rhythm and phoneme information of the speech. Device-side voice RNNT model: The preprocessed speech features are input into the RNNT model, which then analyzes the command intent through the following modules: Encoder: Its structure is: 8-layer bidirectional LSTM with a hidden state dimension of 512, which is used to capture the temporal dependency of speech; Its function is to convert the speech feature sequence Convert to high-dimensional contextual representation , preserve the time series information of speech; Prediction Network: Its structure is a single-layer LSTM with a vocabulary size of 1000; Its function is: based on the historical label sequence Predict possible labels for the current frame , output probability distribution ; Joint Network: Its structure is: fully connected layer, output dimension 256, context representation of fusion encoder and the label probability of the prediction network , generate joint features , the characteristics Can be converted into executable retrieval instructions.

7. The method for intelligent detection of cracks in building structures based on machine vision according to claim 6, characterized in that: The method of using voiceprint positioning technology to locate the voice command and determine the crack corresponding to the executable search command includes: Determine the crack corresponding to the sound source: Delay Difference Calculation: The four-element linear microphone array on the AR glasses synchronously collects the sound source signal and calculates the delay difference between any two elements: ; in, is the time delay difference between any two array elements, is the distance between microphone array elements i and j, is the azimuth of the sound source; is the sound source pitch angle; ; Azimuth and elevation angle estimation: Estimation by generalized cross-correlation algorithm , combined with triangulation to calculate the sound source : ; Combined with the sound source localization results ( , ) and the field of view of AR glasses, the system selects the cracks to be retrieved: Field of view mapping: The spherical coordinates of the sound source ( , , ) is converted to the local coordinate system of the AR glasses to determine whether it falls within the field of view: ; in, is the azimuth, is the pitch angle; Maintain a crack database for the current detection scenario and filter target cracks using the following rules: ; in, is the sound source position, For the The location of the crack.

8. The method for intelligent detection of building structure cracks based on machine vision according to claim 7, characterized in that: The crack corresponding to the instruction is extracted through image processing algorithms, and the geometric features of the crack are calculated, including: Width measurement: at the center of the crack Sample 10 points in the vertical direction and calculate the maximum distance: ; in, is the crack width, and are the three-dimensional coordinates of the left and right edge points of the crack in the qth group of sampling points, is the Euclidean norm, which represents the straight-line distance between two points; Length calculation: The crack length should be the straight-line distance between the two end points of the crack, based on the coordinates of the first and last endpoints extracted from the crack skeleton line. and , the calculation formula is: ; in, is the crack length, and are the starting point and end point coordinates of the crack skeleton line, respectively, based on the three-dimensional space coordinate system constructed in step S1, in millimeters; Tilt angle calculation: skeleton line direction vector , angle with the horizontal plane: ; Where α is the crack inclination angle, which is the acute angle between the vector and its projection vector on the horizontal plane; It is the vertical component that determines the change in the "height" of the tilt. is the horizontal projection length.

9. The method for intelligent detection of building structure cracks based on machine vision according to claim 8, characterized in that: The method of constructing a building structure crack prediction model includes: Training set: The training set needs to provide the model with labeled historical data to ensure the accuracy and generalization of the prediction. The specific composition is as follows: Sample source: Contains historical crack detection data of different types of building structures; Sample features: Each sample is structured data, including input features ( ) and the corresponding label, that is, the actual length extension of the sample in the same period of history, which is directly obtained from the historical detection records; Sample size and diversity: The sample size is ≥10,000, covering different crack types, width levels, and building materials to ensure the model's adaptability to various scenarios; Spatiotemporal data alignment and prediction input: Data query: According to the current crack center point , spatially adjacent points are retrieved in the historical database, and the time window is set to 1 year; Feature Engineering: Extracting Input Feature Matrix : ; in, is the crack width, is the crack length, It's a crack , is the crack depth; Model selection: This model uses the XGBoost regressor, which minimizes the loss function by iteratively adding trees. It has the ability to handle nonlinear relationships and high-dimensional data, supports parallel computing, and is suitable for deployment on end devices. Training objective function: ; Where N is the number of samples, are the parameters of the model, is the true value of the i-th sample, is the model’s predicted value for the i-th sample, It is a regularization term, including L1 and L2 penalties, which are used to prevent overfitting. L1 penalty will cause the coefficients of some features to become 0, thereby playing the role of feature selection. L2 penalty will make the coefficients of the features smaller, but will not become 0, thereby playing the role of reducing the model variance. The purpose of minimizing the loss function is to minimize the predicted value. and the true value The difference between them is minimized by minimizing the sum of squared errors of all samples to make the model's prediction as accurate as possible; Prediction output: Crack expansion distance in the next three months .

10. The method for intelligent detection of building structure cracks based on machine vision according to claim 9, characterized in that: Based on the crack propagation prediction results, high-risk areas are color-coded on the AR glasses interface, including: Perform AR dynamic risk classification warning: ; Visualization implementation: superimpose translucent color blocks on the crack area, add flashing animations on the edges, and trigger voice alarms synchronously.

Citation Information

Cited By

  • Structural crack identification and positioning method based on multi-modal deep learning

    CN121706007A