Complex noise scene concrete defect fine identification method
By combining preliminary class activation map generation, adaptive segmentation threshold processing, dense condition random field optimization and Ivy algorithm hyperparameter optimization, the boundary blur and regional incompleteness of concrete defect detection in complex noise scenarios are solved, and high-precision defect recognition and automated detection are achieved.
Patent Information
- Application Number
- CN202510553335.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art is difficult to achieve high-precision automated detection of concrete structure defects in complex noise scenarios, especially in terms of boundary blur and regional incompleteness, which cannot meet the safety and durability requirements of water-walked buildings.
The combination of preliminary class activation graph generation, adaptive segmentation threshold processing, dense condition random field optimization and Ivy algorithm automation hyperparameter optimization is adopted to improve defect boundary clarification and regional consistency through deep learning technology.
In complex noise environments, the accuracy and boundary alignment of defect detection are improved, the dependence on high-quality labeled data is reduced, the stability and adaptability of the algorithm are enhanced, and it is suitable for the detection of water-resistant building structures such as bridges and dams.
Smart Images

Figure CN120495195A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for finely identifying concrete defects in complex noise scenes, and belongs to the technical field of concrete defect identification in image processing. Background Art
[0002] As water-related structures age, their concrete structures often develop various types of damage, including cracks, holes, spalling, and aggregate exposure, under the coupled effects of complex environmental factors such as water pressure, temperature loads, erosion, and seepage stress. These defects typically originate on the surface and gradually expand into the structure under the action of external forces or environmental loads, resulting in a decrease in overall stiffness and load-bearing capacity, which may eventually lead to structural failure and seriously threaten the safety and durability of water-related structures. Therefore, the efficient identification and accurate detection of concrete structural defects are important technical requirements for ensuring the safe operation of water-related structures.
[0003] Traditional concrete defect detection methods rely on manual inspections or local instrument testing. These methods typically require point-by-point testing after exposing the structure by pumping out water or drying it. This is not only time-consuming and costly, but also carries limitations such as high risk and limited coverage. Furthermore, test results are often subject to subjective factors, making it difficult to accurately inspect large structures in complex environments.
[0004] In recent years, deep learning-based image segmentation techniques have demonstrated advantages in automated defect detection. Weakly supervised segmentation methods, in particular, have become a research hotspot because they rely solely on image-level annotation data, reducing the cost of pixel-level annotation. However, existing weakly supervised segmentation methods still suffer from issues such as blurred boundaries, incomplete detection areas, and severe background interference in complex scenes, making them difficult to meet the high requirements for defect detection accuracy in practical engineering projects. Furthermore, existing methods are unable to adequately handle defect area boundaries and internal consistency in complex noisy environments, further limiting their practical application. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method for fine identification of concrete defects in complex noise scenes. By combining the generation of preliminary class activation maps, segmentation result optimization and the automated parameter optimization strategy of the Ivy algorithm, the defect boundaries are clarified and the regional consistency is improved in complex noise environments.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A method for finely identifying concrete defects in complex noise scenarios includes the following steps:
[0008] Step 1: Extract features from the concrete structure image, generate a preliminary class activation map, and complete the preliminary positioning of the defect area in the concrete structure image;
[0009] Step 2: Set the adaptive segmentation threshold, perform binarization on the preliminary class activation map, divide the pixels in the preliminary class activation map into defect areas or background areas, and generate preliminary segmentation results;
[0010] Step 3: The dense conditional random field is improved using the iterative message passing algorithm of mean field variational inference. The boundary of the preliminary segmentation result is optimized by the improved dense conditional random field to generate a defect segmentation mask to realize the defect recognition of concrete structure images.
[0011] In step 4, before using the improved dense conditional random field to optimize the boundaries of the preliminary segmentation results, the Ivy algorithm is used to automatically optimize the hyperparameters of the improved dense conditional random field, and the optimal hyperparameters found are applied to the improved dense conditional random field.
[0012] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:
[0013] 1. This invention innovatively combines preliminary class activation map (CAM) generation, segmentation result optimization, and the automated hyperparameter optimization strategy of the Ivy Algorithm (IVYA), effectively overcoming existing issues such as fuzzy defect detection boundaries and incomplete segmentation results. It demonstrates strong adaptability and excellent stability in practical engineering applications, providing reliable technical support for automated defect detection and safety assessment of concrete structures in water-related buildings, and has broad application prospects.
[0014] 2. The present invention not only improves the accuracy of defect detection, but also greatly improves the boundary alignment and defect area consistency in complex noise environments, ensuring more accurate and reliable defect positioning and identification. By using CAM generation, excessive reliance on pixel-level data sets is avoided, thereby reducing the need for fine-grained annotated data. The introduction of CAM effectively improves the perception of defect areas, so that the model can still capture key defect features in the absence of high-quality annotated data. At the same time, IVYA's automated hyperparameter optimization strategy improves the adaptability of the model under different conditions and enhances the stability and robustness of the algorithm. By optimizing the segmentation boundary through dense conditional random fields (DenseCRF), the boundary accuracy is further improved, ensuring accurate segmentation of defect areas and reducing false detections and missed detections.
[0015] 3. This invention focuses on automated identification and precise boundary modeling of defects such as defects, erosion, and voids in concrete structures, and has significant engineering application value. It can be widely used to inspect key areas of various water-related structures, including bridges, dams, and tunnels, effectively improving the accuracy of engineering safety inspections and the efficiency of operations and maintenance management. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of a method for finely identifying concrete defects in complex noise scenarios according to the present invention;
[0017] Figure 2 It is a line chart of Kullback-Leibler divergence of different optimization algorithms;
[0018] Figure 3 The radar chart of segmentation indicators of the present invention and different fully supervised segmentation methods, where (a) is the mean intersection over union (MIoU), (b) is the recall rate (Re), (c) is the precision rate (Pr), and (d) is the F1 score (F1);
[0019] Figure 4 This is a visual comparison of the segmentation effects of the present invention and different fully supervised segmentation methods;
[0020] Figure 5 is the segmentation index histogram of the present invention and different weakly supervised segmentation methods;
[0021] Figure 6 This is a visual comparison of the segmentation effects of the present invention and different weakly supervised segmentation methods. DETAILED DESCRIPTION
[0022] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be interpreted as limiting the present invention.
[0023] like Figure 1 As shown in the figure, the present invention proposes a method for fine recognition of concrete defects in complex noise scenes, which is implemented based on Python 3.8 and Pytorch 1.10.1 deep learning framework and includes the following steps:
[0024] Step 1: By inputting a high-resolution image of the concrete structure into an arbitrary feature extraction network, a preliminary class activation map (CAM) is generated to complete the preliminary positioning of the defect area and provide a rough defect location distribution.
[0025] By inputting each concrete structure image with a resolution of 2048×2048 into any pre-trained feature extraction network (such as ResNet-50, which contains 50 layers in total). After the network forward propagation, the last layer of convolutional feature map with a size of 7×7×2048 is extracted from each image. Subsequently, a 7×7 class activation map (CAM) is generated by global weighting by combining the feature map of this layer with the weight of the corresponding category of the classification layer. Then, by upsampling to the original image size (2048×2048), the preliminary defect response class activation map corresponding to each image is obtained to complete the preliminary positioning of the defect area. Among them, CAM is a classic discriminative positioning technology, which is generated by multiplying the feature map of the last convolution layer of the feature network with the target specific weight of the classification layer. The generation of the activation map related to the cth class can be expressed as:
[0026]
[0027] Where, is the class-specific activation map associated with a class of defects c, is the classification layer weight corresponding to class c, It is the feature map of the last convolutional layer.
[0028] Step 2: Based on the preliminary class activation map, the preliminary class activation map is binarized by setting a threshold, and the pixels are divided into the defect area and the background area. The preliminary segmentation result is generated to extract the basic area contour of the defect and describe the general shape of the defect.
[0029] The threshold used for binarization is an adaptive threshold. This method effectively distinguishes multi-category defect areas from background areas and generates preliminary segmentation results by combining the class activation value maximization operation. The specific calculation process is as follows:
[0030]
[0031] Where S(i) is the final multi-category defect segmentation result of pixel i, and the value is the category c to which the pixel belongs. is the set of all possible categories, M c (i) is the activation map value of category c to which pixel i belongs, T c is the adaptive segmentation threshold for category c.
[0032] Adaptive threshold T for each class c c It is dynamically calculated based on the global characteristics of the class activation map, and its formula is:
[0033] T c =μ c +α·σ c
[0034] Where μc is the global mean of the activation map of category c; α is the adjustment coefficient, which controls the sensitivity of the threshold and is usually in the range of [0.5, 2.0]; σ c is the standard deviation of the activation map of category c. The global mean and standard deviation of the activation map of category c are calculated by the following formulas:
[0035]
[0036] Where M and N are the row and column sizes of the image.
[0037] Step 3: Use dense conditional random fields (DenseCRF) to optimize the initial segmentation results. By modeling the color and spatial relationship between pixels, the segmentation results are dynamically adjusted to make the defect boundaries clearer while suppressing the interference of background noise.
[0038] DenseCRF is used to refine the boundaries of the binarized results and model pixel-level label consistency to enhance the spatial coherence and edge accuracy of the defect area. By modeling the color and spatial relationship between pixels, the defect area boundary is refined and the segmentation accuracy is improved. The dense conditional random field determines the maximum posterior probability P(X|T) through inference. The posterior probability follows the Gibbs distribution. The specific formula is as follows:
[0039]
[0040] Where P(X|T) is the posterior probability; X represents the pixel-level label distribution (class configuration), which usually corresponds to the class assignment result of each pixel in the image, that is, the semantic segmentation mask; T is the input image, which is the original image data to be processed and segmented; Z(T) is the normalization factor (partition function), which is used to normalize all possible label configurations to ensure that the sum of the posterior probability distribution is 1; E(x|T) is the energy function, which is used to characterize the match between the class configuration and the image content and the smoothness between labels.
[0041] The allocation function Z(T) and the energy function E(x|T) are calculated by the following formulas (for convenience, T is omitted in the following part of the present invention):
[0042]
[0043] Where exp is an exponential function used to convert the distance between feature vectors into a similarity measure; x is a specific instance of the category configuration, that is, a specific category assignment is made to each pixel in the image; E(x) is the energy function, ψ u (x i ) is a unary potential function that measures the probability of assigning pixel position i to class label xi The cost caused by this is the degree of consistency between the category label and the observation data at the corresponding position i in the observation image I, ψ p (x i ,x j ) is a binary potential energy function, representing the category labels x of adjacent pixels i and j i and x j The relationship between them.
[0044] The unary potential energy function is a key energy term in the conditional random field model. It is obtained by taking the logarithm of the probability that each pixel belongs to a certain category. The calculation of the unary potential energy associated with category c can be expressed as:
[0045] ψ u (x i =c) = -lnP(x i =c)
[0046] Where, ψ u (x i =c) is the unary potential energy of pixel i belonging to category c, which represents the cost of pixel i belonging to this category; P(x i =c) is the probability that pixel i belongs to category c, defined as follows:
[0047]
[0048] In the formula, gt_prob is the confidence of the category corresponding to the true label, which is used to control the degree of trust the model has in the true label in supervised learning; label(x i ) is the ground truth label of pixel i, x i represents the predicted class label for pixel i, and n_labels is the total number of classes. This definition essentially constructs a "soft label" distribution, assigning a higher probability to the true class and an equal residual probability to the remaining classes, thereby establishing a reasonable confidence level for each pixel label during training. This probabilistic form is often used in the definition of unary potential functions, providing prior support based on supervisory information for subsequent conditional random field models.
[0049] The binary potential energy function is the core of the conditional random field model. It describes the relationship between pixels, encourages similar pixels to be assigned to the same category, and assigns pixels with large differences to different categories. The corresponding calculation can be expressed as:
[0050]
[0051] Where μ p (x i ,xj ) is the binary potential energy function, μ(x i ,x j ) is the label compatibility function, when x i ≠x j The value is 1 when k(f i ,f j ) is a linear combination of Gaussian kernel functions and is assigned weight w (m) , m=1,...,M, M is the total number of kernel functions, indicating how many different Gaussian kernel functions are used to calculate the similarity between pixels; k (m) (f i ,f j ) is the mth Gaussian kernel function, which is used to measure the similarity between pixel i and pixel j, f i and f j It is a multidimensional feature representation vector of pixel i and pixel j, which is used to describe the attribute information of each pixel in the image in the feature space (such as color, texture, etc.). i -f j ) T is the transpose of the difference between the feature vectors of pixel i and pixel j, representing the difference between them, Λ (m) It is the precision matrix corresponding to the mth Gaussian kernel function. It is a symmetric positive definite matrix used to adjust the distance metric between eigenvectors, control the width of the Gaussian kernel function, and affect the similarity calculation between pixels.
[0052] The parameters of the DenseCRF are automatically optimized by combining the Ivy Algorithm (IVYA), and the final high-precision segmentation mask is generated through mean-field variational inference for defect identification and subsequent applications. Due to the large number of edge connections in the dense conditional random field, it is difficult to directly solve the posterior probability P(X). This paper adopts an iterative message passing algorithm with mean-field variational inference to facilitate the calculation of the approximate distribution probability Q(X) instead of the exact distribution probability P(X). The specific calculation formula is as follows:
[0053] Q(X)≈P(X)
[0054]
[0055] Where Q(X) is the approximate distribution probability, which represents the approximate posterior distribution probability calculated by variational inference; P(X) is the true posterior distribution probability; Q i (x i ) indicates that the i-th pixel belongs to the label random variable x i The approximate category distribution probability of each pixel has its own approximate distribution, which is obtained by independent calculation; Z i represents the normalization constant, ψ u(x i ) represents the unary potential energy function of a single pixel; It can be expanded as follows:
[0056]
[0057] In the formula, μ(x i ,c)=[x i ≠c]; [·] is an Iverson bracket, which has a value of 1 if the condition inside the bracket is met, and 0 otherwise; is the set of all categories; M is the number of Gaussian kernels; w (m) is the mth weight; It can be expanded into the following form:
[0058]
[0059] Where k (m) (f i ,f j ) is the mth Gaussian kernel function, Indicates the approximate probability that pixel j belongs to category c in the last mean field update. The first output is
[0060] For k(f i ,f j ) Gaussian kernel function linear combination. In the multi-class image segmentation task, the present invention uses a contrast-sensitive linear combination that combines two Gaussian kernel functions. The specific calculation is as follows:
[0061]
[0062] In the formula, k(f i ,f j ) is a linear combination of Gaussian kernel functions, w (1) and w (2) are weight coefficients that control the influence of appearance kernel and smoothness kernel in the total similarity calculation, p i and p j is the two-dimensional spatial position variable of pixel i and pixel j in the image plane, usually representing the position of the pixel in the image (for example, the row and column number or two-dimensional coordinate of the pixel), |p i -p j | is the spatial distance between pixel i and pixel j, which is used to measure their spatial proximity. A smaller distance means that the two pixels are closer and have higher similarity. i and F jis the color variable of pixel i and pixel j, which usually contains the color information of the pixel (for example, RGB value or value of other color space), and is used to describe the visual characteristics of the pixel; |F i -F j | is the color difference between pixel i and pixel j, which measures their similarity in color space. Pixels with smaller color differences have higher similarity. θ α ,θ β ,θ γ are parameters that control the width of the Gaussian kernel. They control the influence of spatial distance, the influence of color difference, and the influence of spatial distance in the smoothness kernel respectively.
[0063] Step 4: The parameters of DenseCRF are automatically optimized in combination with the Ivy algorithm (IVYA), and the final high-precision segmentation mask is generated through mean field variational inference to ensure the high-precision performance of the defect segmentation results in terms of boundary clarity and regional integrity, meeting the needs of automated defect detection of concrete defects in complex noisy environments.
[0064] Since mean field variational inference involves multiple hyperparameters (gt_prob, w1, w2, θ α ,θ β ,θ γ The selection of these hyperparameters relies on experience and manual adjustment, making them difficult to accurately determine manually. To find the optimal hyperparameter combination, the Ivy algorithm (IVYA) was used for optimization search. This algorithm gradually optimizes the solution by simulating the growth and expansion of plants.
[0065] The steps of the Ivy algorithm are as follows:
[0066] (1) Initialize the population: In the initialization phase, the initial sample positions in the population are randomly generated. The position of each sample is determined by the following formula:
[0067]
[0068] Where, is the position of the nth sample, i.e. the initial candidate solution of the parameters; I min , I max represents the minimum and maximum values of the search space; rand(1,D) represents a vector of dimension D, each component of which independently obeys the uniform distribution in the interval [0,1] and is used to introduce random perturbations in the initialization process.
[0069] (2) Growth phase (local selection, the sample will decide whether to update its position based on the fitness value): During the growth phase, each sample will be updated based on the positions of its neighbors and the best neighbor. The update formula is as follows:
[0070]
[0071] Among them, I nn Indicates the optimal neighbor position of the nth sample, usually the neighbor with the best fitness; ΔGv n represents the growth rate of the nth sample, which controls the sample growth rate; N(1,D) is a standard normal distribution random vector used to simulate random changes in the growth process.
[0072] (3) Diffusion stage (global selection): In the diffusion stage, the sample will update its position towards the global optimal solution. The update formula is as follows:
[0073]
[0074] Where, I best represents the position of the global optimal sample; N(1,D) represents a standard normal distribution random vector, which is used to control random fluctuations in the diffusion process.
[0075] (4) Termination condition: The termination condition of the Ivy algorithm is to reach the maximum number of iterations Itermax.
[0076] The present invention quantitatively evaluates the performance and final segmentation results of different optimization methods using Kullback-Leibler and five classic evaluation indicators. The relevant formulas are as follows:
[0077]
[0078] Where P(i) is the probability value of distribution P at i; Q(i) is the probability value of distribution Q at i; TP, FP, FN, and TN represent the classification results under multiple defect categories, respectively, where TP is the true positive (the number of pixels correctly predicted to belong to a certain defect category), FP is the false positive (the number of pixels incorrectly predicted to belong to a certain defect category), FN is the false negative (the number of pixels not correctly predicted to belong to a certain defect category), and TN is the true negative (the number of pixels correctly predicted not to belong to the defect category).
[0079] The effectiveness of IVYA in DenseCRF parameter optimization was verified by comparing it with seven other optimization algorithms: the Ivy algorithm (IVYA) used in the present invention was compared with the Sparrow Search Algorithm (SSA), the Particle Swarm Optimization Algorithm (MPSO) based on adaptive strategy, the Wavelet-based Differential Evolution Algorithm (WMSDE), the Northern Falcon Optimization Algorithm (NGO), the Teamwork Optimization Algorithm (TOA), the Tasmanian Devil Optimization Algorithm (TDO), and the Osprey Optimization Algorithm (OOA). KL divergence (Kullback-Leibler divergence) was used as the optimization objective, with a population size of 90 and an iteration number of 100 to evaluate the convergence speed and final optimization accuracy. A lower KL divergence indicates better parameter optimization, such as Figure 2 As shown in the figure, the IVYA optimization algorithm proposed in this paper performs excellently in DenseCRF parameter optimization. IVYA reduces the KL divergence to 0.450 within 15 iterations and further reduces it to 0.441 after 50 iterations, achieving the lowest value among all methods. Compared with manual tuning, ICRF provides more accurate boundary positioning, reduces misclassification and background noise, and enhances detection stability. In summary, IVYA performs best in terms of KL divergence and convergence speed, becoming an efficient and stable tool for DenseCRF parameter optimization, especially suitable for defect detection in large concrete structures.
[0080] The ICRF method is further compared with AT (Adaptive Threshold) and five advanced fully supervised segmentation methods (FCN, PSPNet, DeepLabV3+, U-Net, CT-CrackSeg). All methods are trained and tested on the same hyperparameters and datasets. The experimental results are evaluated by global evaluation indicators ( Figure 3 (a)-(d)) and local visualization ( Figure 4 ) for presentation. Comparison shows that the ICRF method of the present invention excels in defect boundary clarity and regional integrity, performing comparable to fully supervised methods, and exhibits greater robustness and adaptability in complex environments. Combined with the experimental results of the aforementioned IVYA comparison method, these data further verify the high accuracy and practical engineering application value of ICRF in defect localization and segmentation tasks.
[0081] The ICRF method of the present invention is compared with five advanced weakly supervised segmentation methods (PSA, SEAM, SIPE, L2G and MCTformer). Among them, PSA, SEAM, SIPE and L2G are based on CNN for feature extraction, while MCTformer relies on Transformer to model global information. All methods are trained on the same dataset and only use image-level labels for supervision. The experimental results are evaluated by global evaluation indicators ( Figure 5) and local visualization ( Figure 6 ). Comparison shows that the proposed method excels in defect boundary clarity and regional integrity, outperforming several existing weakly supervised methods. It can improve defect detection stability and robustness, while also enhancing the adaptive adjustment capability during the detection process, making it particularly suitable for automated detection of concrete defects.
[0082] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the aforementioned method for fine identification of concrete defects in complex noise scenarios are implemented.
[0083] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the aforementioned method for fine identification of concrete defects in complex noise scenes.
[0084] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0085] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0086] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0088] The above embodiments are only for illustrating the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for fine identification of concrete defects in complex noise scenes, characterized by: The steps include: Step 1: Extract features from the concrete structure image, generate a preliminary class activation map, and complete the preliminary positioning of the defect area in the concrete structure image; Step 2: Set the adaptive segmentation threshold, perform binarization on the preliminary class activation map, divide the pixels in the preliminary class activation map into defect areas or background areas, and generate preliminary segmentation results; Step 3: The dense conditional random field is improved using the iterative message passing algorithm of mean field variational inference. The boundary of the preliminary segmentation result is optimized by the improved dense conditional random field to generate a defect segmentation mask to realize the defect recognition of concrete structure images. In step 4, before using the improved dense conditional random field to optimize the boundaries of the preliminary segmentation results, the Ivy algorithm is used to automatically optimize the hyperparameters of the improved dense conditional random field, and the optimal hyperparameters found are applied to the improved dense conditional random field.
2. The method for fine identification of concrete defects in complex noise scenes according to claim 1 is characterized in that: In step 1, a pre-trained feature extraction network is used to extract features from the concrete structure image, a feature map generated by the last layer of the feature extraction network is obtained, and the feature map generated by the last layer is multiplied by the classification layer weight corresponding to each category to obtain a preliminary class activation map; The preliminary class activation map is represented as follows: Among them, M c is the class activation map associated with defect class c, θ c is the classification layer weight corresponding to the defect category c, and f is the feature map generated by the last layer of the feature extraction network.
3. The method for fine identification of concrete defects in complex noise scenes according to claim 2 is characterized in that: In step 2, the adaptive segmentation threshold T of the defect category c is dynamically calculated based on the global characteristics of the preliminary class activation map. c : T c =μ c +a·s c Among them, μ c Activate the map M for the category c The global mean of , α is the adjustment coefficient, σ c Activate the map M for the category c The standard deviation of M and N are the class activation maps M and N, respectively. c The row and column size, M c (i) is the activation map value of pixel i corresponding to defect category c; Using T c The formula for binarization is as follows: Where S(i) is the multi-category defect segmentation result of pixel i, and the value is the defect category to which pixel i belongs; is the set of all defect categories; the activation map value is greater than or equal to T c The pixels with negative pixels are classified into the defect area, otherwise they are classified into the background area.
4. The method for fine identification of concrete defects in complex noise scenes according to claim 3 is characterized in that: In step 3, the dense conditional random field determines the category configuration that maximizes the posterior probability P(X) through inference. The improved dense conditional random field uses the approximate distribution probability Q(X) calculated by mean field variational inference to replace P(X). The formula is as follows: Q(X)≈P(X) Among them, Q i (x i ) indicates that pixel i belongs to the label random variable x i The approximate category distribution probability, Z i represents the normalization constant, ψ u (x i ) represents the unary potential energy function of pixel i, ψ u (x i )and Expand as follows: ψ u (x i )=-lnP(x i =c) Among them, gt_prob is the confidence of the category corresponding to the true label, label(x i ) is the true category label of pixel i, n_labels is the total number of defect categories, μ(x i ,c)=[x i ≠c], [] is the Iverson bracket, when the condition in the bracket is met, its value is 1, otherwise it is 0; M is the number of Gaussian kernels, M = 2; w (m) is the mth weight coefficient; Expands to the following form: Among them, k (m) (f i ,f j ) is the mth Gaussian kernel function, when m=1, k (1) (f i ,f j ) is the appearance kernel function, and When m=2, k (2) (f i ,f j ) is the smoothness kernel function, and p i and p j are the two-dimensional spatial position variables of pixel i and pixel j in the image plane, F i and F j are the color variables of pixel i and pixel j respectively, θ α ,θ β ,θ γ Both are parameters that control the width of the Gaussian kernel; Indicates the approximate probability that pixel j belongs to defect category c at the last mean field variation iteration. for The iterative message passing algorithm of mean field variational inference ends with the maximum number of iterations set in advance, and the last iteration is used to obtain Calculate Q(X).
5. The method for fine identification of concrete defects in complex noise scenes according to claim 4 is characterized in that: In step 4, the following six hyperparameters are optimized using the Ivy algorithm: gt_prob, w (1) ,w (2) ,θ α ,θ β ,θ γ The specific process is as follows: (1) In the initialization phase, the initial sample positions in the population are randomly generated, and the position of each sample is initialized by the following formula: in, is the initial position of the nth sample, and one value of the six hyperparameters represents one sample; I min , I max are the minimum and maximum values of the search space, respectively, and N is the number of samples; rand(1,D) represents a vector of dimension D, each component of which independently obeys a uniform distribution in the interval [0,1] and is used to introduce random perturbations in the initialization process; (2) In the growth phase, the first iteration number T1 is set. During the first to T1th iterations, the position of each sample is updated according to the positions of its neighbors and the optimal neighbors. The update formula is as follows: in, is the updated position of the nth sample, I n is the position of the nth sample before updating; N(1,D) is a standard normal distribution random vector; I nn is the optimal neighbor position of the nth sample, ΔGv n is the growth rate of the nth sample, f(I n ) is the fitness value of the nth sample, f(I best ) is the fitness value corresponding to the sample with the best fitness in the current iteration; (3) In the diffusion stage, the second iteration number T2 is set. During the T1+1th to T1+T2th iterations, the position update formula of each sample in each iteration is as follows: Among them, I best is the position of the sample with the best fitness in the current iteration; (4) When the T1+T2th iteration is reached, the Ivy algorithm terminates and the hyperparameter value corresponding to the position of the optimal sample in the last selection is output as the final hyperparameter value.
6. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the method for finely identifying concrete defects in complex noise scenes according to any one of claims 1 to 5 are implemented.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for finely identifying concrete defects in complex noise scenes according to any one of claims 1 to 5 are implemented.