Data-centered airport pavement crack segmentation method and device
Through self-training method and pseudo-label screening technology, combined with comparison learning and data enhancement, the dependence on labeled data in airport road fracture segmentation is solved, the robustness and generalization ability of the model are improved, and the efficient and low-cost crack segmentation effect is achieved.
Patent Information
- Application Number
- CN202510042726.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art relies on a large amount of labeled data in the crack segmentation of airport road surfaces, which is expensive to label and poor generalization performance, making it difficult to effectively utilize labelless data.
The self-training method is used to combine contrast learning and data augmentation technology to improve the robustness and generalization ability of the model through pseudo-label screening and orthogonal constraint loss function, and reduce dependence on labeled data.
It significantly improves the performance of the crack segmentation model, reduces the annotation cost, and improves the precise segmentation ability of the model in practical application scenarios, especially in different materials and complex backgrounds.
Smart Images

Figure CN119963841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of airport pavement safety monitoring and image processing technology, and in particular to a data-centric airport pavement crack segmentation method and device. Background Art
[0002] Crack detection is critical to ensuring the safety and functionality of airport pavement facilities [1] Crack image segmentation is a technique that can accurately distinguish crack pixels from non-crack pixels. [2] , providing data support and quantitative tools for airport operations and maintenance [3] The crack segmentation method based on traditional image processing faces many challenges. It not only requires a lot of manual intervention for feature recognition and extraction, but also has poor generalization performance. [4] .
[0003] In recent years, with the rapid development of deep learning algorithms, many crack segmentation models based on deep learning have been proposed, including encoder-decoder based networks and attention mechanism based networks. In order to further improve the detection accuracy and efficiency, researchers have proposed a series of improvement measures, such as: a) using a deeper pre-trained model as an encoder; b) building a new decoder containing more information; c) designing a new feature fusion module to improve the ability to extract multi-scale information; d) improving the connection mechanism to enhance the ability to capture global context information; e) proposing a new loss function to improve the model training accuracy; f) integrating robust pre-processing and post-processing methods.
[0004] Large-scale, high-quality crack annotation datasets determine the upper limit of crack segmentation model performance [5] Relying on the airport pavement high-speed inspection robot, it can efficiently obtain airport pavement crack data, and has accumulated unlabeled data covering multiple airports. [6] However, due to the small and irregular shape characteristics, the annotation cost of crack images is high compared to other conventional targets, requiring professional training and multiple rounds of inspection and screening. The annotation cost eventually becomes a bottleneck that limits the effectiveness of the crack segmentation model, making it difficult to effectively utilize massive amounts of unlabeled data. Therefore, how to use the original data to provide the model with richer learning materials and improve the segmentation effect without increasing the annotation cost is a problem with practical application value.
[0005] In order to reduce the reliance on a large amount of labeled data, relevant researchers have proposed methods based on unsupervised learning. However, it is difficult to achieve efficient crack segmentation using models pre-trained with unsupervised learning. In contrast, semi-supervised learning (SSL) balances the advantages and disadvantages of fully supervised learning and unsupervised learning by using a small number of labeled samples and a large number of unlabeled samples. The segmentation model based on SSL not only significantly reduces the labeling cost, but also maintains the accuracy of the model by fully tapping the potential of unlabeled data. [7] .
[0006] References
[0007] [1] Li Haifeng, Jing Pan, Han Hongyang. Airport pavement crack detection algorithm based on deformable convolution and feature fusion[J]. Journal of Nanjing University of Aeronautics & Astronautics, 2021, 53(06): 981-988.
[0008] [2]Li H, Zong J, Huang R, et al. AggCrack: An aggregated attention model for robotic crack detection in challenging airport runway environment [C] / / 2022IEEE 18th International Conference on Automation Science and Engineering (CASE). IEEE, 2022: 1747-1752.
[0009] [3]Dong CZ,Catbas F N.Areview of computer vision–based structuralhealth monitoring at local and global levels[J]. Structural Health Monitoring, 2021,20(2):692-743.
[0010] [4]Guo JM, Markoni H, Lee J D. BARNet: Boundary aware refinement network for crack detection[J]. IEEE Transactions on Intelligent TransportationSystems, 2021, 23(7):7343-7358.
[0011] [5]Lin Y,Xu ZZ,Chen D,Ai ZJ,Qiu Y,Yuan Y Z.Wood crack detectionbased on data-driven semantic segmentation network[J]. IEEE / CAAJournal ofAutomatica Sinica,2023,10(6):1510-1512.
[0012] [6] Li Haifeng, Fan Tianxiao, Huang Rui, et al. Airport runway crack segmentation based on combined self-attention and axial-attention[J]. Journal of Zhengzhou University (Science Edition), 2023, 55(04): 30-38.
[0013] [7] Li G, Wan J, He S, et al. Semi-supervised semantic segmentation using adversarial learning for pavement crack detection [J]. IEEE Access, 2020, 8: 51446-51459. Summary of the invention
[0014] The present invention provides a data-centric airport pavement crack segmentation method and device. The present invention is based on a self-training method, draws on the idea of contrastive learning, and combines data enhancement technology to evaluate the stability of pseudo-labels generated by different enhancement methods, and gradually screen out more reliable pseudo-labels; secondly, a batch-based data enhancement method is adopted to alleviate the problem of unbalanced data distribution by increasing the proportion of crack pixels in a single image in batch sampling, while enhancing the robustness of the model; finally, an orthogonal constraint is introduced in the loss function to prompt the model to learn more independent features, thereby improving the generalization ability in cross-domain tasks. The method and device can not only achieve high-precision detection and segmentation of airport pavement cracks, but also have strong adaptability and extensibility, providing reliable technical support for the intelligent detection and maintenance of airport pavement cracks, and helping to improve the efficiency and accuracy of airport pavement maintenance work. See the following description for details:
[0015] In a first aspect, a data-centric airport pavement crack segmentation method is provided, the method comprising:
[0016] The BatchMix mechanism is used to reorganize high-density crack areas and generate enhanced samples; the standard cross entropy loss function is used as the objective function to optimize the teacher model;
[0017] We use progressive consistency to filter pseudo labels, apply data augmentation to unlabeled data, and evaluate the stability of pseudo labels under different data augmentation conditions. We combine the output results of the teacher model at different training stages to comprehensively judge the reliability of pseudo labels.
[0018] According to the reliability of pseudo labels, the unlabeled data are divided into credible data and uncredible data from high to low. The screened credible data are combined with the original labeled data to generate a new training data set, and the BatchMix enhancement mechanism is applied to generate more high-density crack samples.
[0019] An orthogonal constraint loss function consisting of a cross entropy loss and an orthogonal constraint loss is introduced to train the student model and iteratively update the parameters of the student model.
[0020] The trained student model is deployed to the airport pavement inspection system, and works with a variety of acquisition devices to output crack segmentation results.
[0021] The method of using progressive consistency to select pseudo labels and applying data enhancement to unlabeled data is specifically as follows:
[0022] Calculate the average similarity of the pseudo labels of all enhanced versions and unenhanced data to obtain the image-level enhancement consistency score S aug :
[0023]
[0024] The enhanced consistency score is used to reflect the stability of pseudo labels under different data augmentations;
[0025] During the first stage of training of the teacher model, K training checkpoints are saved and each checkpoint is used to train the same unlabeled image I i ∈D u Make predictions and generate pseudo label sets {M i,1 ,M i,2 ,...,M i,k}, calculate the mIoU value between these predictions to measure the consistency of the pseudo labels, and the multi-checkpoint consistency score S ckpt The calculation is as follows:
[0026]
[0027] The final pseudo-label reliability score is obtained by enhancing the consistency score S aug and multi-checkpoint consistency score S ckpt The comprehensive reliability score S is obtained by combining p for:
[0028] S p =S aug +S ckpt
[0029] According to the comprehensive reliability score S p , for the unlabeled image set D u The samples in are sorted, and the top R most reliable unlabeled images and their pseudo labels are selected as high-quality pseudo-label samples to participate in the subsequent training stage.
[0030] The BatchMix enhancement mechanism is used to generate more high-density crack samples:
[0031] a) Input and initialization, assuming that the batch sampling data is where x i represents the image, y i Indicates the corresponding label, the image size is s i , the grid size is s p , define the crack pixel label value as c;
[0032] b) Synchronous segmentation, x i and i Synchronous segmentation is divided into pieces of size s p Grid blocks of the grid are obtained to obtain a set of blocks:
[0033]
[0034] Among them, m is the number of grids after block division, x i,j andi,j The images x i and label y i The jth grid of
[0035] c) Crack pixel statistics and sorting, for each label grid y i,j , calculate the number of crack pixels as follows:
[0036]
[0037] Among them, (k,l) is the label grid y i,j Pixels in ;
[0038] d) Grid selection and stitching: select n grid blocks from the sorted queue, stitch them into a complete image, and insert them into the sampled batch data.
[0039] Among them, the orthogonal constraint loss function is:
[0040] The calculation process of the orthogonal constraint loss function is to initialize the loss value L orth Start with zero; for each parameter matrix W in the model parameter set, if W is a bias term, skip the processing of the matrix; otherwise, convert it to a shape of (C in ,C out ) of the two-dimensional matrix W f , where C in and C out Respectively represent the number of output channels and the number of input channels;
[0041] Calculate the matrix W f The autocorrelation matrix R is obtained by subtracting the identity matrix I from R, and the error matrix ΔR is obtained; the sum of the absolute values of the elements of ΔR is calculated, multiplied by the regularization factor λ, and accumulated to L orth ;
[0042] Returns the orthogonality constraint loss value L orth , used to constrain the characteristic orthogonality of the weight matrix; the calculated orthogonal constraint loss value L orth The constraint factor, which is an orthogonal constraint, is integrated into the overall loss function calculation and used as part of the supervision signal to train the student model.
[0043] In a second aspect, a data-centric airport pavement crack segmentation device is provided, the device comprising: a processor and a memory, the memory storing program instructions, the processor calling the program instructions stored in the memory to enable the device to execute any one of the methods described in the first aspect.
[0044] In a third aspect, a computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes any one of the methods described in the first aspect.
[0045] The beneficial effects of the technical solution provided by the present invention are:
[0046] 1. Aiming at the problem of data scarcity and imbalanced distribution in crack segmentation tasks, this paper proposes an optimized self-training framework, which significantly improves the performance of the segmentation model by making full use of large-scale unlabeled data; the proposed pseudo-label screening mechanism is based on data enhancement consistency and multi-checkpoint consistency evaluation, which can effectively filter low-quality pseudo-labels and improve the reliability of training data;
[0047] 2. The BatchMix data enhancement method adopted in the present invention alleviates the data imbalance problem in the crack segmentation task by splicing the high-density crack areas; the introduction of the orthogonal constraint loss function further enhances the feature diversity and generalization ability of the model;
[0048] 3. Experiments on the APD dataset show that the proposed STCrack method has achieved significant improvements in multiple indicators compared with the baseline method and mainstream segmentation methods, including a 2.25% increase in mIoU and a 2.71% increase in F1 score. In addition, STCrack performs well in the detection of crack edge details and sparse crack areas, showing strong generalization ability and robustness.
[0049] 4. Without increasing the labeling cost, the present invention effectively utilizes unlabeled data, improves the model's ability to accurately segment cracks in actual application scenarios, and especially maintains excellent segmentation performance under different airport pavement materials and complex backgrounds, providing an efficient and low-cost solution for the intelligent detection and maintenance of airport pavement cracks. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of a data-centric airport pavement crack segmentation method;
[0051] Figure 2 This is the BatchMix flow chart;
[0052] Figure 3 A visualization of the improvement effect. DETAILED DESCRIPTION
[0053] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.
[0054] Self-training is a classic SSL method. It predicts pseudo labels of unlabeled images and combines them with real labeled images to increase the amount of training data. It has significant advantages in improving the efficiency and practicality of crack segmentation model development. However, there are some challenges in directly applying self-training methods to crack segmentation tasks. First, due to the significant imbalance in crack image data, that is, the crack area usually occupies a smaller part of the image, which makes it difficult for the model to capture and correctly classify crack features in the initial training stage. Therefore, the method of directly using fixed threshold screening to generate pseudo labels can often only produce a very limited number of high-confidence pseudo labels. Secondly, pavement crack images of different airports have significant differences in shooting conditions, surface materials, crack types, and background noise. This cross-domain image difference interferes with the process of generating pseudo labels by the model. This interference may lead to a decrease in the accuracy of pseudo labels, further limiting the effectiveness of self-training methods in cross-domain situations. These factors make the direct application of self-training in crack segmentation tasks challenging, and further strategy adjustments and method improvements are needed to improve the quality and quantity of pseudo label generation.
[0055] Example 1
[0056] Based on the traditional self-training method, the STCrack framework proposes an improved solution that combines data enhancement and pseudo-label screening. The technical route specifically includes the following three core links:
[0057] 1. Teacher Model Training
[0058] Data preparation: Use a small-scale labeled dataset as training data and train the teacher model through standard supervised learning methods.
[0059] The BatchMix data augmentation method is applied. First, the original training data is gridded, and high-density crack areas are selected according to the distribution density of crack pixels. Then, the high-density crack areas are reorganized using the BatchMix mechanism to generate enhanced samples to alleviate the data imbalance problem.
[0060] The standard cross-entropy loss function is used as the objective function to optimize the teacher model.
[0061] 2. Pseudo-label generation and screening
[0062] Use the trained teacher model to perform inference on the unlabeled data and generate pseudo labels.
[0063] The pseudo-labels are screened using a progressive consistency screening method. First, multiple data augmentations (e.g., rotation, flipping, etc.) are applied to the unlabeled data to evaluate the stability of the pseudo-labels under different data augmentations. Second, the reliability of the pseudo-labels is comprehensively judged by combining the output results of the teacher model at different training stages.
[0064] Finally, according to the reliability of the pseudo-labels, the unlabeled data are divided into credible data and uncredible data from high to low, and credible data is used preferentially for subsequent training.
[0065] 3. Student Model Training
[0066] Training data construction: Combine the filtered high-quality pseudo-labeled data with the original labeled data to generate a new training data set.
[0067] Continue to apply the BatchMix enhancement mechanism to generate more high-density crack samples. Introduce the orthogonal constraint loss function to improve the complementarity of the model in feature learning and avoid learning redundant features. The loss function consists of two parts: cross entropy loss and orthogonal constraint loss.
[0068] Use the new dataset and improved loss function to train the student model, iteratively update the parameters of the student model, and improve the segmentation performance.
[0069] Example 2
[0070] The embodiment of the present invention provides a data-centric airport pavement crack segmentation method, see Figures 1 to 3 , the method comprises the following steps:
[0071] 1. Overview of the Airport Pavement Crack Segmentation Framework
[0072] STCrack framework Figure 1 As shown, the dotted box indicates the specific improvements corresponding to each stage of the general self-training framework.
[0073] For a given airport pavement annotation data (x,y)∈(X,Y), where x∈R H×W×3 is a sample image in the airport pavement crack image set X, y∈R H×W×C is a sample label in the corresponding label set Y. H, W, and C represent the height, width, and number of categories of the image, respectively. In the crack segmentation task, the pixel categories are divided into cracks and background.
[0074] STCrack consists of two stages. First, the teacher model is trained on an existing labeled dataset using a cross-entropy loss function, enabling it to learn valuable crack features from limited labeled data and generate high-quality prediction results.
[0075] The loss function used in teacher model training is shown in formula (1).
[0076]
[0077] Where N is the number of pixels in the image, x i is a pixel in the original image, y i is the one-hot encoding of the corresponding class label, p T represents the Softmax prediction of the teacher model T containing the class probabilities, and A() represents the data augmentation operation.
[0078] Then, the teacher model T is applied to a large-scale unlabeled dataset x' and generates pseudo labels y'. The pseudo labels characterize the model's understanding of the unlabeled data and become an indication of potential crack areas, thus enabling the unlabeled data to play a role in the training process, as shown in formula (2).
[0079] y'~argmax C p T (x') (2)
[0080] Here, y' is the same one-hot encoded pseudo-label as y, which can save a lot of storage space and training time. T (x') is the pixel-level prediction of the teacher model T for the unlabeled data x', and C is the number of categories of the pixel.
[0081] Finally, the generated pseudo labels are filtered according to the progressive consistency screening method and merged with the corresponding crack images into the original annotated dataset to construct a larger hybrid dataset to further train the student model S. Different from using a single cross entropy loss function to train the teacher model T, STCrack introduces orthogonal constraints in the training process of the student model S, as shown in formula (3).
[0082]
[0083] Among them, M is the number of unlabeled data, p s represents the Softmax prediction containing category information output by the student model S. σ is the orthogonal constraint factor, which is specifically expressed by formula (4). Compared with the teacher model T, the student model S can gradually improve the segmentation ability of unlabeled crack data through the guidance of pseudo labels.
[0084]
[0085] Among them, λ is the regularization factor, W is the parameter matrix, and I is the identity matrix. The student model S is optimized with the cross entropy loss of integrated soft orthogonal regularization as the goal. Through this teacher-student self-training framework, the model can fully utilize the potential of unlabeled data without significantly increasing the annotation workload, and improve the accuracy and stability of the crack segmentation task.
[0086] The above is an overall overview of the STCrack framework. The improvements made in the embodiments of the present invention for the crack segmentation scenario will be described in more detail in the following sections.
[0087] 2. Progressive consistency screening method
[0088] In order to solve the reliability evaluation problem in pseudo-label screening, a new scheme is proposed to perform pseudo-label screening at the image level based on a single training model. i ∈D u , apply different data enhancement methods to generate N enhanced versions {I i,1 ,I i,2 ,...,I i,N}, where the enhancement methods include: Gaussian blur, Gaussian noise and edge filtering, that is, N = 3, which is intended to simulate the diversity of data distribution. Then, the teacher model is used to perform segmentation prediction on each enhanced crack image to generate a pseudo label set {M i,1 ,M i,2 ,...,M i,N By calculating the pseudo label M of the unenhanced image data i,0 The consistency between the pseudo-labels in the pseudo-label set is used to quantify the stability of the pseudo-labels under different data augmentation situations. The cosine similarity is used to measure the similarity between the pseudo-labels. The average similarity of the pseudo-labels of all enhanced versions and the unenhanced data is calculated to obtain the image-level enhancement consistency score S aug , the specific calculation method is as follows:
[0089]
[0090] The enhanced consistency score is used to reflect the stability of pseudo-labels under different data enhancements. The higher the consistency, the more stable the pseudo-labels are under data changes and the higher the reliability.
[0091] In order to further improve the robustness of reliability assessment, it is necessary to combine multi-checkpoint consistency scoring. During the first stage of teacher model training, K training checkpoints are saved, and each checkpoint is used to evaluate the same unlabeled image I i ∈D u Make predictions and generate pseudo label sets {M i,1 ,M i,2,...,M i,k}. The consistency of pseudo labels is measured by calculating the mIoU value between these predictions. Multi-checkpoint consistency score S ckpt The calculation method is as follows:
[0092]
[0093] Here, mIoU represents the intersection over union ratio between pseudo labels predicted at different checkpoints, which can reflect the pseudo label consistency of the model under different training states. The final pseudo label reliability score is obtained by enhancing the consistency score S aug and multi-checkpoint consistency score S ckpt Define the comprehensive reliability score S p for:
[0094] S p =S aug +S ckpt (7)
[0095] This comprehensive scoring can more comprehensively measure the reliability of unlabeled images and avoid the bias that may be caused by a single evaluation standard. p , for the unlabeled image set D u The samples in the dataset are sorted and the top R most reliable unlabeled images and their pseudo labels are selected as high-quality pseudo-label samples to participate in the subsequent training phase. This strategy ensures that the model learns the features of reliable samples first and gradually expands to more complex samples, which helps improve the generalization ability and robustness of the model.
[0096] 3. Batch-based data enhancement method
[0097] The embodiment of the present invention designs a data enhancement method based on intra-Batch splicing, called BatchMix, to alleviate the problem of unbalanced crack distribution, such as Figure 2 As shown. By synchronously segmenting the images and labels in the batch sampling data, the high-density areas containing cracks are prioritized and spliced to increase the proportion of crack pixels, thereby enhancing the robustness and segmentation performance of the model. The specific method is as follows:
[0098] a) Input and initialization. Assume that the batch sampling data is where x i represents the image, y i Indicates the corresponding label, the image size is s i , the grid size is s p , define the crack pixel label value as c.
[0099] b) Synchronous segmentation. i and iSynchronous segmentation is divided into pieces of size s p Grid blocks of the grid are obtained to obtain a set of blocks:
[0100]
[0101] Among them, m is the number of grids after block division, x i,j and i,j The images x i and label y i The j-th grid of .
[0102] c) Crack pixel statistics and sorting. For each label grid y i,j , calculate the number of crack pixels (that is, the number of pixels with label value c), the calculation formula is as follows:
[0103]
[0104] Among them, (k,l) is the label grid y i,j The pixels in .
[0105] Since the label value of crack pixels is 1 and the label value of background pixels is 0, the label grid y can be directly calculated i,j The density of crack pixels in the grid is measured by the sum of the pixels in . Each grid block and its crack pixel number are stored in a queue and sorted in descending order according to the crack pixel number.
[0106] d) Grid selection and splicing. Select n grid blocks from the sorted queue, splice them into a complete image, and insert them into the sampled batch data. In order to ensure that a complete image can be spliced each time, n=s i / s p .
[0107] Enhanced batch data D b+1 The grid blocks with high density of crack pixels are retained, so more crack pixels are included than normal crack image data. BatchMix increases the proportion of crack samples in batch sampling data by giving priority to grid areas containing more crack pixels, effectively alleviating the common data imbalance problem in crack segmentation tasks, and enhancing the model's sensitivity to small crack areas.
[0108] 4. Loss Function with Regularization Constraints
[0109] Considering that the teacher model and the student model use the same network structure and are similarly initialized, they tend to make similar correct or incorrect predictions on unlabeled images, resulting in the student model being unable to learn additional information from them except for entropy minimization. In order to further improve the diversity of pseudo-labels and avoid the teacher model and the student model from generating overly consistent prediction results on unlabeled images due to the similarity of structure and initialization, the embodiment of the present invention introduces orthogonal constraints when training the student model. The orthogonal constraint loss is calculated as follows:
[0110] 1. Calculation process of orthogonal constraint loss function to initialize loss value L orth Start from zero;
[0111] 2. Then, for each parameter matrix W in the model parameter set, if W is a bias term, skip processing the matrix; otherwise, convert it to a matrix with a shape of (C in ,C out ) of the two-dimensional matrix W f , where C in and C out Respectively represent the number of output channels and the number of input channels;
[0112] 3. On this basis, calculate the matrix W f The autocorrelation matrix R is obtained, and the error matrix ΔR is obtained by subtracting the identity matrix I from R;
[0113] 4. Next, calculate the sum of the absolute values of the elements of ΔR, multiply it by the regularization factor λ, and add it to L orth ;
[0114] 5. Finally, return the orthogonal constraint loss value L orth , which is used to constrain the feature orthogonality of the weight matrix, thereby improving the expressiveness and generalization performance of the model.
[0115] The calculated orthogonal constraint loss value L orth The constraint factor, which is an orthogonal constraint, is integrated into the calculation of the overall loss function of formula (3) and used as part of the supervisory signal to train the student model. In the crack segmentation task of the airport pavement, the model needs to identify small cracks, mesh cracks, and damage of various shapes and textures. The orthogonal constraint can help the student model avoid making overly consistent predictions with the teacher model, forcing the model to explore more effectively in the feature space, so that the model can capture more differentiated features. This design reduces the reliance on a single feature and improves the generalization ability of the model in complex and dynamic scenarios.
[0116] V. Experimental part
[0117] In order to more intuitively demonstrate the improvement effect brought by the improved self-training framework, the airport pavement crack segmentation results are visualized, such as Figure 3 The first row in the figure is the original image of the airport pavement, the second row is the corresponding label, the third row is the segmentation result of the baseline model, and the fourth row is the segmentation result after applying the improved self-training framework.
[0118] 6. Model deployment and application
[0119] The trained student model is deployed to the airport pavement inspection system, which can work with a variety of acquisition devices, such as drones, high-resolution ground cameras, or vehicle-mounted inspection equipment. The inspection equipment will capture airport pavement images at a high frequency and transmit the image data to the backend server or edge device in real time. The student model will perform inference on the backend or edge device, quickly process the input pavement image, and output the segmentation result of the crack.
[0120] During the reasoning process, the student model shows good adaptability to various crack forms (including: small cracks, irregular cracks, etc.) and complex environments (such as: lighting changes, pavement of different materials) through its efficient network architecture and orthogonal constraint optimization capabilities. The segmentation results will be presented in the form of binary images or crack annotations superimposed on the original image, and further integrated into the airport's pavement maintenance management platform. Through this platform, airport operation and maintenance personnel can visualize the location and distribution of cracks, automatically generate maintenance tasks, and even interact with other systems (such as pavement health assessment systems) to provide support for decision-making.
[0121] Example 3
[0122] A data-centric airport pavement crack segmentation device, the device comprising: a processor and a memory, the memory storing program instructions, the processor calling the program instructions stored in the memory to enable the device to execute the following method steps in embodiment 1:
[0123] The BatchMix mechanism is used to reorganize high-density crack areas and generate enhanced samples; the standard cross entropy loss function is used as the objective function to optimize the teacher model;
[0124] We use progressive consistency to filter pseudo labels, apply data augmentation to unlabeled data, and evaluate the stability of pseudo labels under different data augmentation conditions. We combine the output results of the teacher model at different training stages to comprehensively judge the reliability of pseudo labels.
[0125] According to the reliability of pseudo labels, the unlabeled data are divided into credible data and uncredible data from high to low. The screened credible data are combined with the original labeled data to generate a new training data set, and the BatchMix enhancement mechanism is applied to generate more high-density crack samples.
[0126] An orthogonal constraint loss function consisting of a cross entropy loss and an orthogonal constraint loss is introduced to train the student model and iteratively update the parameters of the student model.
[0127] The trained student model is deployed to the airport pavement inspection system, and works with a variety of acquisition devices to output crack segmentation results.
[0128] Among them, using progressive consistency to filter pseudo labels, applying data enhancement to unlabeled data is specifically as follows:
[0129] Calculate the average similarity of the pseudo labels of all enhanced versions and unenhanced data to obtain the image-level enhancement consistency score S aug :
[0130]
[0131] The enhanced consistency score is used to reflect the stability of pseudo labels under different data augmentations;
[0132] During the first stage of training of the teacher model, K training checkpoints are saved and each checkpoint is used to train the same unlabeled image I i ∈D u Make predictions and generate pseudo label sets {M i,1 ,M i,2 ,...,M i,k}, calculate the mIoU value between these predictions to measure the consistency of the pseudo labels, and the multi-checkpoint consistency score S ckpt The calculation is as follows:
[0133]
[0134] The final pseudo-label reliability score is obtained by enhancing the consistency score S aug and multi-checkpoint consistency score S ckpt The comprehensive reliability score S is obtained by combining p for:
[0135] S p =S aug +S ckpt
[0136] According to the comprehensive reliability score S p , for the unlabeled image set D u The samples in are sorted, and the top R most reliable unlabeled images and their pseudo labels are selected as high-quality pseudo-label samples to participate in the subsequent training stage.
[0137] Among them, the BatchMix enhancement mechanism is used to generate more high-density crack samples:
[0138] a) Input and initialization, assuming that the batch sampling data is where x i represents the image, y i Indicates the corresponding label, the image size is s i , the grid size is s p , define the crack pixel label value as c;
[0139] b) Synchronous segmentation, x i and i Synchronous segmentation is divided into pieces of size s p Grid blocks of the grid are obtained to obtain a set of blocks:
[0140] X p ={x i,j |j=1,2,...,m}
[0141] Y p ={y i,j |j=1,2,...,m}
[0142] Among them, m is the number of grids after block division, x i,j and i,j The images x i and label y i The jth grid of
[0143] c) Crack pixel statistics and sorting, for each label grid y i,j , calculate the number of crack pixels as follows:
[0144]
[0145] Among them, (k,l) is the label grid y i,j Pixels in ;
[0146] d) Grid selection and stitching: select n grid blocks from the sorted queue, stitch them into a complete image, and insert them into the sampled batch data.
[0147] Among them, the orthogonal constraint loss function is:
[0148] The calculation process of the orthogonal constraint loss function is to initialize the loss value L orth Start with zero; for each parameter matrix W in the model parameter set, if W is a bias term, skip the processing of the matrix; otherwise, convert it to a shape of (C in ,C out ) of the two-dimensional matrix W f , where C in and C out Respectively represent the number of output channels and the number of input channels;
[0149] Calculate the matrix Wf The autocorrelation matrix R is obtained by subtracting the identity matrix I from R, and the error matrix ΔR is obtained; the sum of the absolute values of the elements of ΔR is calculated, multiplied by the regularization factor λ, and accumulated to L orth ;
[0150] Returns the orthogonality constraint loss value L orth , used to constrain the characteristic orthogonality of the weight matrix; the calculated orthogonal constraint loss value L orth The constraint factor, which is an orthogonal constraint, is integrated into the overall loss function calculation and used as part of the supervision signal to train the student model.
[0151] It should be pointed out here that the device description in the above embodiment corresponds to the method description in the embodiment, and the embodiment of the present invention will not be described in detail here.
[0152] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as computers, single-chip microcomputers, and microcontrollers. In specific implementation, the embodiments of the present invention do not limit the execution subjects and are selected according to the needs of actual applications.
[0153] The data signal is transmitted between the memory and the processor via a bus, which is not described in detail in the embodiment of the present invention.
[0154] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, the storage medium includes a stored program, and when the program is running, the device where the storage medium is located is controlled to execute the method steps in the above embodiment.
[0155] The computer-readable storage medium includes but is not limited to a flash memory, a hard disk, a solid-state drive, and the like.
[0156] It should be pointed out here that the description of the readable storage medium in the above embodiment corresponds to the description of the method in the embodiment, and the embodiment of the present invention will not be described in detail here.
[0157] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present invention are generated.
[0158] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions may be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be accessed by the computer or a data storage device such as a server or a data center that includes one or more available media. The available medium may be a magnetic medium or a semiconductor medium, etc.
[0159] Unless otherwise specified, the models of the components in the embodiments of the present invention are not limited, and any device that can perform the above functions may be used.
[0160] Those skilled in the art will appreciate that the accompanying drawing is only a schematic diagram of a preferred embodiment, and the serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0161] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A data-centric airport pavement crack segmentation method, characterized in that: The method comprises: The BatchMix mechanism is used to reorganize high-density crack areas and generate enhanced samples; the standard cross entropy loss function is used as the objective function to optimize the teacher model; We use progressive consistency to filter pseudo labels, apply data augmentation to unlabeled data, and evaluate the stability of pseudo labels under different data augmentation conditions. We combine the output results of the teacher model at different training stages to comprehensively judge the reliability of pseudo labels. According to the reliability of pseudo labels, the unlabeled data are divided into credible data and uncredible data from high to low. The screened credible data are combined with the original labeled data to generate a new training data set, and the BatchMix enhancement mechanism is applied to generate more high-density crack samples. An orthogonal constraint loss function consisting of a cross entropy loss and an orthogonal constraint loss is introduced to train the student model and iteratively update the parameters of the student model. The trained student model is deployed to the airport pavement inspection system, and works with a variety of acquisition devices to output crack segmentation results.
2. The data-centric airport pavement crack segmentation method according to claim 1, characterized in that: The method of using progressive consistency to filter pseudo labels and applying data enhancement to unlabeled data is specifically as follows: Calculate the average similarity of the pseudo labels of all enhanced versions and unenhanced data to obtain the image-level enhancement consistency score S aug : The enhanced consistency score is used to reflect the stability of pseudo labels under different data augmentations; During the first stage of training of the teacher model, K training checkpoints are saved and each checkpoint is used to train the same unlabeled image I i ∈D u Make predictions and generate pseudo label sets {M i,1 ,M i,2 ,...,M i,k }, calculate the mIoU value between these predictions to measure the consistency of the pseudo labels, and the multi-checkpoint consistency score S ckpt The calculation is as follows: The final pseudo-label reliability score is obtained by enhancing the consistency score S aug and multi-checkpoint consistency score S ckpt The comprehensive reliability score S is obtained by combining p for: S p =S aug +S ckpt According to the comprehensive reliability score S p , for the unlabeled image set D u The samples in are sorted, and the top R most reliable unlabeled images and their pseudo labels are selected as high-quality pseudo-label samples to participate in the subsequent training stage.
3. The data-centric airport pavement crack segmentation method according to claim 1, characterized in that: The BatchMix enhancement mechanism is applied to generate more high-density crack samples: a) Input and initialization, assuming that the batch sampling data is where x i Represents the image, y i Indicates the corresponding label, the image size is s i , the grid size is s p , define the crack pixel label value as c; b) Synchronous segmentation, x i and i Synchronous segmentation is divided into pieces of size s p Grid blocks of the grid are obtained to obtain a set of blocks: Among them, m is the number of grids after block division, x i,j and i,j The images x i and label y i The jth grid of c) Crack pixel statistics and sorting, for each label grid y i,j , calculate the number of crack pixels as follows: Among them, (k,l) is the label grid y i,j Pixels in ; d) Grid selection and stitching: select n grid blocks from the sorted queue, stitch them into a complete image, and insert them into the sampled batch data.
4. The data-centric airport pavement crack segmentation method according to claim 1, characterized in that: The orthogonal constraint loss function is: The calculation process of the orthogonal constraint loss function is to initialize the loss value L orth Start with zero; for each parameter matrix W in the model parameter set, if W is a bias term, skip the processing of the matrix; otherwise, convert it to a shape of (C in ,C out ) of the two-dimensional matrix W f , where C in and C out Respectively represent the number of output channels and the number of input channels; Calculate the matrix W f The autocorrelation matrix R is obtained by subtracting the identity matrix I from R; the sum of the absolute values of the elements of ΔR is calculated, multiplied by the regularization factor λ, and added to L orth ; Returns the orthogonality constraint loss value L orth , used to constrain the characteristic orthogonality of the weight matrix; the calculated orthogonal constraint loss value L orth The constraint factor, which is an orthogonal constraint, is integrated into the overall loss function calculation and used as part of the supervision signal to train the student model.
5. A data-centric airport pavement crack segmentation device, characterized in that: The device comprises: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method according to any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is enabled to perform the method according to any one of claims 1 to 4.