A method and device for marking pavement structure layer defects under three-dimensional multi-view imaging

By intercepting and stitching views in the visual sequence multi-view of the pavement structure layer, a feature association labeling model is constructed, which solves the problems of low efficiency and low accuracy of the pavement structure layer disease labeling in the prior art, and achieves efficient and accurate automated labeling.

CN119399168BActive Publication Date: 2025-05-13CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411521939.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-05-13
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

In the prior art, the disease labeling efficiency of pavement structure layer is low and the accuracy is low, resulting in untimely and inaccurate disease monitoring.

Method used

By separating the horizontal plane view and vertical section view from the visual sequence multi-view of the pavement structure layer, stitching it into a two-dimensional image across the view, labeling is used using the visual image annotation software, and a horizontal plane-longitudinal section feature association annotation model is constructed for automatic annotation.

Benefits of technology

It improves the efficiency and accuracy of disease marking of pavement structure layers, reduces the time of manual marking and the possibility of missed marking, and realizes automatic marking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399168B_ABST
    Figure CN119399168B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for marking defects in pavement structure layers under three-dimensional multi-view imaging, and relates to the technical field of computer image recognition. The method comprises: splicing a horizontal plane view and a longitudinal section view to obtain a two-dimensional image across views; constructing a horizontal plane-longitudinal section feature association labeling model including multiple multi-level feature fusion networks and a spatial pyramid pooling layer; and obtaining a trained feature association labeling model for the model using a training set. In this way, defects are automatically labeled by using the trained horizontal plane-longitudinal section feature association labeling model, which reduces the time for manual labeling and improves the efficiency of defect labeling; using the two-dimensional image across views, considering the similar defect feature association information between adjacent horizontal plane views and longitudinal section views, optimizing the labeling performance of the model, and automatically labeling defects by using the trained horizontal plane-longitudinal section feature association labeling model, which reduces mislabeling and missed labeling and improves the accuracy of defect labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image recognition, and in particular to a method and device for marking road structure layer defects under three-dimensional multi-view imaging. Background Art

[0002] With the rapid development of transportation infrastructure, highway assets are gradually transitioning from large-scale infrastructure to large-scale maintenance. The safety monitoring and management of the pavement structure layer space is particularly important. The factors that affect the safety of the pavement structure layer are mainly water-rich diseases (such as leakage, water gushing, etc.), which may cause serious structural damage, failure of roadbed and pavement service performance, and even cause personnel safety accidents. Therefore, accurate and timely monitoring of pavement structure layer diseases is of great significance to the long-term safety and stable service of the pavement.

[0003] At present, 3D radar technology is usually used as an effective and non-destructive means of monitoring the pavement structure layer. By converting radar signals into high-resolution images of multiple views, the health of the structure layer can be effectively reflected. By analyzing the 3D radar sequence multi-view imaging, the imaging characteristics of different diseases can be used to indirectly identify the diseases in the underground pavement structure layer. After identifying the diseases in the underground pavement structure layer, it is necessary to mark the diseases in the underground pavement structure layer to provide a data basis for the subsequent 3D radar automatic detection. However, the diseases in the underground pavement structure layer are usually marked manually. When the amount of disease data in the underground pavement structure layer is large, manual marking takes a long time, which makes the efficiency of disease marking low. Manual marking may result in missing labels and errors, which makes the marking accuracy low. Summary of the invention

[0004] The purpose of the embodiments of the present invention is to provide a method and device for marking pavement structure layer defects under three-dimensional multi-view imaging, so as to solve the problem of low efficiency and low accuracy in marking pavement structure layer defects.

[0005] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:

[0006] A first aspect of the present invention provides a method for marking pavement structure layer defects under three-dimensional multi-view imaging, the method comprising:

[0007] In the visualized sequence multi-view of the pavement structure layer, a horizontal plane view and a longitudinal section view are cut out;

[0008] The horizontal plane view and the longitudinal section view are stitched together to obtain a two-dimensional image across the views;

[0009] Using visual image annotation software, the defects in the two-dimensional images across views are annotated to obtain corresponding annotation files, and the annotation files and the two-dimensional images across views are combined into data pairs, and the data pairs are divided into a training set and a validation set;

[0010] Construct a horizontal plane-longitudinal section feature association annotation model, which is used to annotate pavement structure layer defects. The horizontal plane-longitudinal section feature association annotation model includes a backbone network and a head network. The backbone network is a network that includes multiple multi-level feature fusion networks and a spatial pyramid pooling layer.

[0011] The horizontal plane-longitudinal section feature association annotation model is trained by using the training set to obtain a trained horizontal plane-longitudinal section feature association annotation model;

[0012] The validation set is input into the trained horizontal plane-longitudinal section feature association annotation model, and the trained horizontal plane-longitudinal section feature association annotation model is used to annotate the defects in the validation set, and the defect annotation results are output.

[0013] A second aspect of the present invention provides a pavement structure layer disease labeling device under three-dimensional multi-view imaging, the device comprising:

[0014] A clipping unit, used for clipping out a horizontal plane view and a longitudinal section view from a sequence of multiple views of a visualized pavement structure layer;

[0015] A stitching unit, used for stitching the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the views;

[0016] A first labeling unit is used to label the defects in the two-dimensional image across views by using visual image labeling software to obtain a corresponding labeling file, and to form a data pair with the labeling file and the two-dimensional image across views, and to divide the data pair into a training set and a verification set;

[0017] A construction unit is used to construct a horizontal plane-longitudinal section feature association annotation model, which is used to annotate pavement structure layer defects. The horizontal plane-longitudinal section feature association annotation model includes a backbone network and a head network. The backbone network is a network including multiple multi-level feature fusion networks and a spatial pyramid pooling layer.

[0018] A training unit, used for training the horizontal plane-longitudinal section feature association annotation model using the training set to obtain a trained horizontal plane-longitudinal section feature association annotation model;

[0019] The second labeling unit is used to input the verification set into the trained horizontal plane-longitudinal section feature association labeling model, use the trained horizontal plane-longitudinal section feature association labeling model to label the defects in the verification set, and output the defect labeling result.

[0020] Compared with the prior art, the method and device for marking pavement structure layer defects under three-dimensional multi-view imaging provided by the present invention cut out the horizontal plane view and the longitudinal section view in the visualized sequential multi-view of the pavement structure layer; splice the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the view; use the visual image annotation software to mark the defects in the two-dimensional image across the view to obtain the corresponding annotation file, and form a data pair with the annotation file and the two-dimensional image across the view, and divide the data pair into a training set and a verification set; construct a horizontal plane-longitudinal section feature association annotation model, and the horizontal plane-longitudinal section feature association annotation model is used to identify the defects in the two-dimensional image across the view. The surface feature association labeling model is used to label the pavement structure layer defects. The horizontal plane-longitudinal section feature association labeling model includes a backbone network and a head network. The backbone network is a network containing multiple multi-level feature fusion networks and spatial pyramid pooling layers. The horizontal plane-longitudinal section feature association labeling model is trained using a training set to obtain a trained horizontal plane-longitudinal section feature association labeling model. The verification set is input into the trained horizontal plane-longitudinal section feature association labeling model, and the defects in the verification set are labeled using the trained horizontal plane-longitudinal section feature association labeling model, and the defect labeling results are output. In this way, the pavement structure layer defects can be automatically labeled through the trained horizontal plane-longitudinal section feature association labeling model, reducing the time and workload of manual labeling and improving the efficiency of pavement structure layer defect labeling. By utilizing the two-dimensional images across views, the similar defect feature association information between adjacent horizontal plane views and longitudinal section views is considered to optimize the labeling performance of the horizontal plane-longitudinal section feature association labeling model. In addition, the pavement structure layer defects can be automatically labeled through data pairs and the trained horizontal plane-longitudinal section feature association labeling model, reducing the possibility of mislabeling and missing labels, and improving the accuracy of pavement structure layer defect labeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] By reading the detailed description below with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0022] Figure 1 The flowchart of the pavement structure layer disease labeling method under three-dimensional multi-view imaging is schematically shown;

[0023] Figure 2 schematically illustrates a schematic diagram of a two-dimensional image across a view;

[0024] Figure 3 The structural diagram of the backbone network is schematically shown;

[0025] Figure 4 A schematic diagram schematically shows the disease annotation results;

[0026] Figure 5 The structure diagram of the pavement structure layer disease labeling device under three-dimensional multi-view imaging is schematically shown. DETAILED DESCRIPTION

[0027] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0028] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present invention should have the common meanings understood by those skilled in the art to which the present invention belongs.

[0029] The method in the embodiment of the present invention is described in detail below.

[0030] Figure 1 The flowchart of the method for marking pavement structure layer defects under three-dimensional multi-view imaging in an embodiment of the present invention is schematically shown. Figure 1 As shown, the method may include:

[0031] S101. In the visualized sequence multi-view of the pavement structure layer, extract a horizontal plane view and a longitudinal section view.

[0032] Before extracting the horizontal plane view and the longitudinal section view from the visualized sequence multi-view, the method further includes:

[0033] Step A1: Acquire a three-dimensional image of the pavement structure layer.

[0034] The 3D radar system emits high-frequency electromagnetic wave pulses to penetrate the highway pavement and deep into the underground medium. By comprehensively analyzing the measurement data of multiple pulses or different angles, it can construct a 3D image of the pavement structure layer.

[0035] Step A2: Convert the 3D image into multiple views of the sequence for visualization.

[0036] Among them, the visualized sequence multi-view consists of a horizontal plane view, a longitudinal section view and a cross-section view.

[0037] In order to process radar signals, the 3D images are converted into multiple views for visualization.

[0038] The visualized sequence multi-views include multiple horizontal pixel points of the horizontal plane view and multiple vertical pixel points of the longitudinal section view.

[0039] The multiple horizontal pixel points at least include pixel points at the four corners of the horizontal plane view, namely, the upper left pixel point, the lower left pixel point, the upper right pixel point, and the lower right pixel point of the horizontal plane view. The multiple vertical pixel points at least include pixel points at the four corners of the longitudinal section view, namely, the upper left pixel point, the lower left pixel point, the upper right pixel point, and the lower right pixel point of the longitudinal section view.

[0040] The characteristic information of the disease is mainly reflected in the horizontal view and the longitudinal section view. It is necessary to extract the horizontal view and the longitudinal section view from the visualized sequence multi-view of the pavement structure layer. The horizontal view and the longitudinal section view are views in the same time and space.

[0041] Specifically, in the visualized sequence multi-view of the pavement structure layer, the horizontal plane view and the longitudinal section view are cut out, including:

[0042] Step B1: According to a plurality of horizontal pixel points, a horizontal plane view is cut out from the visualized sequence multi-view.

[0043] According to the upper left pixel point, the lower left pixel point, the upper right pixel point and the lower right pixel point of the horizontal plane view, the horizontal plane view is cut out from the visualized sequence multi-view.

[0044] Step B2: According to a plurality of vertical pixel points, a longitudinal section view is cut out from the visualized sequence multi-view.

[0045] The width of the intercepted horizontal plane view is the same as the width of the horizontal plane view and the height of the horizontal plane view is different from the height of the horizontal plane view.

[0046] The longitudinal section view is cut out in the visualized sequence multi-view according to the upper left pixel point, the lower left pixel point, the upper right pixel point and the lower right pixel point of the longitudinal section view.

[0047] After the horizontal plane view and the longitudinal section view are captured in the visualized sequence multi-view, it can also be verified whether the upper left pixel point, lower left pixel point, upper right pixel point and lower right pixel point of the horizontal plane view are within the pixel range of the horizontal plane view, and whether the upper left pixel point, lower left pixel point, upper right pixel point and lower right pixel point of the longitudinal section view are within the pixel range of the longitudinal section view. If so, the horizontal plane view and the longitudinal section view are spliced. If not, the horizontal plane view and the longitudinal section view are captured again in the visualized sequence multi-view.

[0048] S102: splice the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the views.

[0049] A cross-view 2D image is a 2D image containing two views.

[0050] In order to construct a two-dimensional image across the view, the longitudinal view and the horizontal view are spliced ​​on the side with the same width, wherein the longitudinal view is located at the top and the horizontal view is located at the bottom, thereby obtaining a two-dimensional image across the view.

[0051] Figure 2 A schematic diagram of a two-dimensional image across views is schematically shown. The two-dimensional image across views is a spliced ​​image, wherein the upper portion is a longitudinal section view, and the lower portion is a horizontal plane view.

[0052] S103, using visual image annotation software, annotating the defects in the two-dimensional image across views to obtain a corresponding annotation file, and forming a data pair with the annotation file and the two-dimensional image across views, and dividing the data pair into a training set and a validation set.

[0053] The visual image annotation software includes an open source graphic image annotation tool (labelimg).

[0054] The disease in the two-dimensional image across the view is annotated using the visual image annotation software based on the characteristic information and geographic location information of the disease. Specifically, the annotation box is determined along the minimum circumscribed rectangle of the disease. The coordinates of the upper left pixel point of the annotation box in the image are P1 (x1, y1), the coordinates of the lower right pixel point are P2 (x2, y2), and the category c. The annotation content is exported to a plain text (txt) data format. The two-dimensional image in the txt data format contains detailed information of the annotation box. The detailed information of the annotation box includes the width of the annotation box, the height of the annotation box, the category of the disease, the relative position of the center point of the annotation box in the width of the two-dimensional image, and the relative position of the center point of the annotation box in the height of the two-dimensional image. The relative position is between 0 and 1. The width of the annotation box and the height of the annotation box can be expressed as a ratio of the width and height of the two-dimensional image. The categories of diseases include loose, water-rich, and hollow.

[0055] After the network automatically annotates defects in two-dimensional images across views, the annotated results can be standardized. The standardization includes two aspects: one is format alignment, and the other is value normalization.

[0056] Add P1(x1, y1) and P2(x2, y2) to the width w of the annotation box, the height h of the annotation box, the category c of the disease, the conversion relationship of the center point of the annotation box, and the normalization process between 0 and 1. The specific process of normalization is: the center point x and y of the annotation box are divided by the width and height of the two-dimensional image across the view, and the width and height of the annotation box are divided by the width and height of the two-dimensional image across the view.

[0057] S104. Construct a horizontal plane-longitudinal section feature association annotation model.

[0058] Among them, the horizontal plane-longitudinal section feature association annotation model is used to annotate the pavement structure layer defects. The horizontal plane-longitudinal section feature association annotation model includes a backbone network and a head network. The backbone network is a network containing multiple multi-level feature fusion networks and spatial pyramid pooling layers.

[0059] The backbone network can extract features from the input image.

[0060] Specifically, Figure 3 The structural diagram of the backbone network is schematically shown, and the backbone network includes a first convolutional layer, a first multi-level feature fusion network, a second convolutional layer, a second multi-level feature fusion network, a first dimensionality increase / decrease feature extraction module, a third convolutional layer, a third multi-level feature fusion network, a second dimensionality increase / decrease feature extraction module, a fourth convolutional layer, a fourth multi-level feature fusion network, a third dimensionality increase / decrease feature extraction module, a fifth convolutional layer, a fifth multi-level feature fusion network, a fourth dimensionality increase / decrease feature extraction module and a spatial pyramid pooling layer, which are connected in sequence.

[0061] The first dimensionality-lifting feature extraction module includes three sixth convolutional layers connected in sequence, the second dimensionality-lifting feature extraction module includes six seventh convolutional layers connected in sequence, the third dimensionality-lifting feature extraction module includes nine eighth convolutional layers connected in sequence, and the fourth dimensionality-lifting feature extraction module includes three ninth convolutional layers connected in sequence.

[0062] The number of filters in each convolutional layer in the backbone network starts from 64 and gradually increases to 1024, while the size and step size of the convolution kernel are adjusted as needed to change the size of the feature map. Each up-down dimension feature extraction module (i.e., C3 module) extracts image features by repeatedly applying convolutional layers. The spatial pyramid pooling layer can capture features of different scales and provide rich contextual information for subsequent image annotation.

[0063] The first convolutional layer has 64 filters, each of which is 6×6 in size and has a stride of 2. It will halve the width and height of the input image in the training set. The size of the output feature map of the first convolutional layer is 1 / 2 of the input image size. The second convolutional layer has 128 filters, each of which is 3×3 in size and has a stride of 2. It halves the size of the feature map of the input features of the second convolutional layer, i.e., the output of the first multi-level feature fusion network. The size of the output feature map of the second convolutional layer is 1 / 4 of the input image size. The third convolutional layer has 256 filters, and the size of the output feature map of the third convolutional layer is 1 / 8 of the input image. The fourth convolutional layer has 512 filters, with a stride of 2, and the size of the output feature map of the fourth convolutional layer is 1 / 16 of the input image. The fifth convolutional layer has 1024 filters, with a stride of 2, and the size of the output feature map of the fifth convolutional layer is 1 / 32 of the input image. Each dimension-raising and lowering feature extraction module does not change the size of the input feature map. The spatial pyramid pooling layer has 5 pooling windows of different sizes to capture features of different scales.

[0064] Specifically, the first multi-level feature fusion network, the second multi-level feature fusion network, the third multi-level feature fusion network, the fourth multi-level feature fusion network, and the fifth multi-level feature fusion network all include an image block segmentation layer, an image embedding layer, multiple multi-level feature fusion blocks, and multiple image downsampling layers. Each multi-level feature fusion block includes a first normalization layer, a window multi-head self-attention layer, a second normalization layer, a first multi-layer perception mechanism layer, a third normalization layer, a sliding window multi-head self-attention layer, a fourth normalization layer, and a second multi-layer perception mechanism layer connected in sequence.

[0065] The image block segmentation layer is used to convert the smallest unit of the image into an image block (patch). The size of the image block in the image block segmentation layer is 4×4, and the output is a feature matrix of H / 4×W / 4×48. The image embedding layer is used to convert the number of channels of the input feature map of this layer into the number of channels C. Each image downsampling layer is used for downsampling. In the tensor matrix corresponding to the input image of this layer, a sample is taken every other point, and then the sampled tensor matrix is ​​spliced, so that the number of channels is doubled and the size is doubled. Each multi-level feature fusion network is used to perform multi-layer processing on the input features of this layer to extract rich feature information. It can fully extract the local and global features in the input image of this layer while maintaining efficient calculation. The window multi-head self-attention layer and the sliding window multi-head self-attention layer realize the effective interaction of local and global information through sliding window operation and window shift. Each multi-layer perception mechanism layer can perform nonlinear transformation, which enhances the abstraction and expression ability of features.

[0066] Specifically, the head network includes a tenth convolutional layer, a first upsampling layer, a first connection layer, a fifth up-dimensionality feature extraction module, an eleventh convolutional layer, a second upsampling layer, a second connection layer, a sixth up-dimensionality feature extraction module, a twelfth convolutional layer, a third connection layer, a seventh up-dimensionality feature extraction module, a thirteenth convolutional layer, a fourth connection layer, an eighth up-dimensionality feature extraction module and a detection layer, which are connected in sequence.

[0067] The third connection layer is also connected to the fifth up-down dimensionality feature extraction module, and the fourth connection layer is also connected to the sixth up-down dimensionality feature extraction module.

[0068] The fourth multi-level feature fusion network of the backbone network is connected to the first connection layer of the head network, and the third multi-level feature fusion network of the backbone network is connected to the second connection layer of the head network. The spatial pyramid pooling layer of the backbone network is connected to the tenth convolutional layer of the head network.

[0069] The head network can detect diseases, and each convolution layer in the head network can reduce the number of channels of the input feature map of each convolution layer. Each upsampling layer enlarges the size of the feature map input to each upsampling layer. The first connection layer can fuse the feature map output by the first upsampling layer with the feature map output by the fourth multi-level feature fusion network. The second connection layer can fuse the feature map output by the second upsampling layer with the feature map output by the third multi-level feature fusion network, and can achieve splicing of different feature maps in the channel dimension, thereby achieving multi-scale fusion of features. The detection layer combines feature maps of different scales and outputs the final detection result, including the category and location of the disease.

[0070] The tenth convolutional layer has 512 filters, each of which is 1×1 in size and has a step size of 1. The tenth convolutional layer is used to reduce the number of channels of the feature map. The first upsampling layer uses the nearest neighbor interpolation method to double the size of the input feature map. The first connection layer doubles the number of channels of the input feature map of this layer. The fifth up-dimensional feature extraction module contains 3 convolutional layers. The eleventh convolutional layer has 256 filters, each of which is 1×1 in size and has a step size of 1. The second upsampling layer also doubles the size of the feature map input to this layer. The twelfth convolutional layer has 256 filters, each of which is 3×3 in size and has a step size of 2. The thirteenth convolutional layer has 512 filters, each of which is 3×3 in size and has a step size of 2, and is used to downsample the feature map. The detection layer combines the feature map output by the fifth up-dimensional feature extraction module, the feature map output by the sixth up-dimensional feature extraction module, and the feature map output by the eighth up-dimensional feature extraction module, and outputs the final annotation result, i.e., the detection result.

[0071] S105. Train the horizontal plane-longitudinal section feature association annotation model using the training set to obtain a trained horizontal plane-longitudinal section feature association annotation model.

[0072] Specifically, the horizontal plane-longitudinal section feature association annotation model is trained using the training set to obtain a trained horizontal plane-longitudinal section feature association annotation model, including:

[0073] Step C1: Iteratively train the horizontal plane-longitudinal section feature association annotation model using the training set to obtain the preset disease annotation results of the current round and the model parameters of the current round.

[0074] Specifically, the horizontal plane-longitudinal section feature association annotation model is iteratively trained using the training set to obtain the preset disease annotation results and model parameters of the current round, including:

[0075] Step C11: Input the training set into the backbone network so that the backbone network outputs multiple intermediate features.

[0076] Among them, the multiple intermediate features include a first intermediate feature, a second intermediate feature and a third intermediate feature.

[0077] Specifically, the training set is input into the backbone network so that the backbone network outputs multiple intermediate features, including:

[0078] Step C111: input the training set into the first convolution layer, the first multi-level feature fusion network, the second convolution layer, the second multi-level feature fusion network, the first up-down dimension feature extraction module, the third convolution layer and the third multi-level feature fusion network in sequence, so that the third multi-level feature fusion network outputs the first intermediate feature.

[0079] Step C112: input the first intermediate feature into the second up-down dimension feature extraction module, the fourth convolutional layer and the fourth multi-level feature fusion network in sequence, so that the fourth multi-level feature fusion network outputs the second intermediate feature.

[0080] Step C113: input the second intermediate feature into the third dimensionality reduction feature extraction module, the fifth convolutional layer, the fifth multi-level feature fusion network, the fourth dimensionality reduction feature extraction module and the spatial pyramid pooling layer in sequence, so that the spatial pyramid pooling layer outputs the third intermediate feature.

[0081] Step C12: Input multiple intermediate features into the head network so that the head network outputs the preset disease labeling results of the current round and the model parameters of the current round.

[0082] The preset disease marking result of the current round is in the same data format as the two-dimensional image after marking in step S103, including the width of the marking box, the height of the marking box, the category of the disease, the relative position of the center point of the marking box to the width of the two-dimensional image, and the relative position of the center point of the marking box to the height of the two-dimensional image.

[0083] Step C2: Obtain target disease labeling results.

[0084] Among them, the target disease labeling result is the result of labeling the same disease as the preset disease labeling result of the current round.

[0085] The target disease labeling result is the result of accurate labeling.

[0086] Step C3: Compare the preset disease labeling result of the current round with the target disease labeling result to determine the accuracy of the preset disease labeling result of the current round.

[0087] Specifically, the preset disease labeling results of the current round are automatically matched and calibrated with the target disease labeling results. During the calibration process, it is important to check whether the labeling box accurately identifies the characteristics of the disease, and record the precision, that is, the accuracy, of the preset disease labeling results of the current round.

[0088] Step C4: Determine whether the accuracy of the preset disease marking result of the current round reaches the preset accuracy value.

[0089] The preset accuracy value may be 95%.

[0090] It is determined whether the accuracy of the preset disease marking result of the current round reaches the preset accuracy value. If so, step C5 is executed; if not, step C6 is executed.

[0091] Step C5: Stop the iteration to obtain the trained horizontal plane-longitudinal section feature association annotation model.

[0092] Determine whether the accuracy of the preset disease labeling results of the current round reaches the preset accuracy value. If so, stop the iteration, indicating that the current horizontal plane-longitudinal section feature association labeling model has the ability of fully automatic labeling, and a trained horizontal plane-longitudinal section feature association labeling model can be obtained.

[0093] Step C6: Input the target disease labeling result into the horizontal plane-longitudinal section feature association labeling model, use the target disease labeling result to iteratively train the horizontal plane-longitudinal section feature association labeling model to adjust the model parameters of the current round, obtain the preset disease labeling result of the next round and the model parameters of the next round, use the preset disease labeling result of the next round as the preset disease labeling result of the current round, use the model parameters of the next round as the model parameters of the current round, and return to the step of comparing the preset disease labeling result of the current round with the target disease labeling result, until the accuracy of the preset disease labeling result of the next round reaches the preset accuracy value, stop the iteration, and obtain the trained horizontal plane-longitudinal section feature association labeling model.

[0094] Determine whether the accuracy of the preset disease labeling results of the current round reaches the preset precision value. If not, input the target disease labeling results into the horizontal plane-longitudinal section feature association labeling model. Use the target disease labeling results to iteratively train the horizontal plane-longitudinal section feature association labeling model to adjust the model parameters of the current round and obtain the preset disease labeling results and model parameters of the next round. Use the preset disease labeling results of the next round as the preset disease labeling results of the current round, use the model parameters of the next round as the model parameters of the current round, and return to step C3 to continue executing until the accuracy of the preset disease labeling results of the next round reaches the preset precision value of 95%, then stop the iteration to obtain the trained horizontal plane-longitudinal section feature association labeling model.

[0095] The preset disease marking result of the next round is in the same data format as the two-dimensional image after marking in step S103, including the width of the marking box, the height of the marking box, the category of the disease, the relative position of the center point of the marking box to the width of the two-dimensional image, and the relative position of the center point of the marking box to the height of the two-dimensional image.

[0096] The trained horizontal plane-longitudinal section feature association annotation model can compare the disease features extracted from the input images in adjacent training sets, analyze the similarities and differences between them, and thus obtain richer feature information. The trained horizontal plane-longitudinal section feature association annotation model can accurately locate and identify the disease objects in the image. The horizontal plane-longitudinal section feature association annotation model uses cross-scale aggregation technology to fuse features from different layers to enhance the ability to identify diseases.

[0097] The training process of the horizontal plane-longitudinal section feature association annotation model involves the process of optimizing the loss function, which not only considers the performance of the horizontal plane-longitudinal section feature association annotation model on two-dimensional images, but also considers the feature association between adjacent views (i.e., horizontal plane views and longitudinal section views) in three-dimensional data, thereby making the horizontal plane-longitudinal section feature association annotation model more robust when processing complex data.

[0098] S106, inputting the verification set into the trained horizontal plane-longitudinal section feature association labeling model, labeling the defects in the verification set using the trained horizontal plane-longitudinal section feature association labeling model, and outputting the defect labeling result.

[0099] The defect marking result has the same data format as the two-dimensional image after marking in step S103, including the width of the marking box, the height of the marking box, the type of the defect, the relative position of the center point of the marking box to the width of the two-dimensional image, and the relative position of the center point of the marking box to the height of the two-dimensional image.

[0100] Figure 4A schematic diagram of the defect marking result is schematically shown, where the red frame is the marking frame for the pavement structure layer, and loose type defects are marked through the marking frame.

[0101] In the comparative experiment of the present invention, two different pavement structure layer disease labeling methods are evaluated: manual labeling method and the present invention method. The experimental results show that compared with manual labeling, the present invention method significantly improves the labeling efficiency, almost twice that of manual labeling, and also performs well in accuracy, reaching more than 90%, which is slightly higher than the accuracy of manual labeling.

[0102] Based on the above Figure 1 It can be seen from the implementation method that the embodiment of the present invention cuts out the horizontal plane view and the longitudinal section view from the visualized sequence multi-view of the pavement structure layer; splices the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the view; uses the visual image annotation software to annotate the defects in the two-dimensional image across the view to obtain a corresponding annotation file, and the annotation file and the two-dimensional image across the view form a data pair, and the data pair is divided into a training set and a verification set; constructs a horizontal plane-longitudinal section feature association annotation model, the horizontal plane-longitudinal section feature association annotation model is used to annotate the pavement structure layer defects, the horizontal plane-longitudinal section feature association annotation model includes a backbone network and a head network, and the backbone network is a network including multiple multi-level feature fusion networks and a spatial pyramid pooling layer; uses the training set to train the horizontal plane-longitudinal section feature association annotation model to obtain a trained horizontal plane-longitudinal section feature association annotation model; inputs the verification set into the trained horizontal plane-longitudinal section feature association annotation model, uses the trained horizontal plane-longitudinal section feature association annotation model to annotate the defects in the verification set, and outputs the defect annotation result. In this way, the pavement structure layer defects can be automatically labeled through the trained horizontal plane-longitudinal section feature association labeling model, reducing the time and workload of manual labeling and improving the efficiency of pavement structure layer defect labeling. By utilizing the two-dimensional images across views, the similar defect feature association information between adjacent horizontal plane views and longitudinal section views is considered to optimize the labeling performance of the horizontal plane-longitudinal section feature association labeling model. In addition, the pavement structure layer defects can be automatically labeled through data pairs and the trained horizontal plane-longitudinal section feature association labeling model, reducing the possibility of mislabeling and missing labels, and improving the accuracy of pavement structure layer defect labeling.

[0103] Based on the same inventive concept, as an implementation of the above-mentioned pavement structure layer disease labeling method under three-dimensional multi-view imaging, an embodiment of the present invention further provides a pavement structure layer disease labeling device under three-dimensional multi-view imaging. Figure 5 is a structural diagram of the device in the embodiment of the present invention, see Figure 5 As shown, the device may include:

[0104] The interception unit 501 is used to intercept the horizontal plane view and the longitudinal section view from the visualized sequence multi-view of the pavement structure layer;

[0105] A stitching unit 502 is used to stitch the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the views;

[0106] The first labeling unit 503 is used to label the defects in the two-dimensional image across the view using visual image labeling software to obtain a corresponding labeling file, and to form a data pair with the labeling file and the two-dimensional image across the view, and to divide the data pair into a training set and a validation set;

[0107] A construction unit 504 is used to construct a horizontal plane-longitudinal section feature association annotation model, which is used to annotate pavement structure layer defects. The horizontal plane-longitudinal section feature association annotation model includes a backbone network and a head network. The backbone network is a network including multiple multi-level feature fusion networks and a spatial pyramid pooling layer.

[0108] The training unit 505 is used to train the horizontal plane-longitudinal section feature association annotation model using the training set to obtain a trained horizontal plane-longitudinal section feature association annotation model;

[0109] The second labeling unit 506 is used to input the verification set into the trained horizontal plane-longitudinal section feature association labeling model, use the trained horizontal plane-longitudinal section feature association labeling model to label the defects in the verification set, and output the defect labeling result.

[0110] The device may also include: an acquisition unit, for acquiring a three-dimensional image of the pavement structure layer before cutting out the horizontal plane view and the longitudinal section view from the visualized sequence multi-view;

[0111] The conversion unit is used to convert the three-dimensional image into a visualized sequence of multiple views, wherein the visualized sequence of multiple views consists of a horizontal plane view, a longitudinal section view and a cross-section view.

[0112] The interception unit 501 is specifically used to intercept a horizontal plane view in the visualized sequence multi-view according to multiple horizontal pixel points; intercept a longitudinal section view in the visualized sequence multi-view according to multiple vertical pixel points, the width of the horizontal plane view is the same as the width of the horizontal plane view and the height of the horizontal plane view is different from the height of the horizontal plane view; the visualized sequence multi-view includes multiple horizontal pixel points of the horizontal plane view and multiple vertical pixel points of the longitudinal section view.

[0113] Construction unit 504, the backbone network includes a first convolutional layer, a first multi-level feature fusion network, a second convolutional layer, a second multi-level feature fusion network, a first dimensionality-lifting feature extraction module, a third convolutional layer, a third multi-level feature fusion network, a second dimensionality-lifting feature extraction module, a fourth convolutional layer, a fourth multi-level feature fusion network, a third dimensionality-lifting feature extraction module, a fifth convolutional layer, a fifth multi-level feature fusion network, a fourth dimensionality-lifting feature extraction module and a spatial pyramid pooling layer, which are connected in sequence; the first dimensionality-lifting feature extraction module includes three sixth convolutional layers connected in sequence, the second dimensionality-lifting feature extraction module includes six seventh convolutional layers connected in sequence, the third dimensionality-lifting feature extraction module includes nine eighth convolutional layers connected in sequence, and the fourth dimensionality-lifting feature extraction module includes three ninth convolutional layers connected in sequence.

[0114] Construction unit 504, the first multi-level feature fusion network, the second multi-level feature fusion network, the third multi-level feature fusion network, the fourth multi-level feature fusion network, and the fifth multi-level feature fusion network all include an image block segmentation layer, an image embedding layer, multiple multi-level feature fusion blocks, and multiple image downsampling layers; each multi-level feature fusion block includes a first normalization layer, a window multi-head self-attention layer, a second normalization layer, a first multi-layer perception mechanism layer, a third normalization layer, a sliding window multi-head self-attention layer, a fourth normalization layer, and a second multi-layer perception mechanism layer connected in sequence.

[0115] Construction unit 504, the head network includes a tenth convolutional layer, a first upsampling layer, a first connection layer, a fifth dimensionality increase / decrease feature extraction module, an eleventh convolutional layer, a second upsampling layer, a second connection layer, a sixth dimensionality increase / decrease feature extraction module, a twelfth convolutional layer, a third connection layer, a seventh dimensionality increase / decrease feature extraction module, a thirteenth convolutional layer, a fourth connection layer, an eighth dimensionality increase / decrease feature extraction module and a detection layer, which are connected in sequence.

[0116] The training unit 505 is specifically used to iteratively train the horizontal plane-longitudinal section feature association labeling model using the training set to obtain the preset disease labeling result of the current round and the model parameters of the current round; obtain the target disease labeling result, which is the result of labeling the same disease as the preset disease labeling result of the current round; compare the preset disease labeling result of the current round with the target disease labeling result to determine the accuracy of the preset disease labeling result of the current round; determine whether the accuracy of the preset disease labeling result of the current round reaches the preset accuracy value; if so, stop the iteration to obtain the trained horizontal plane-longitudinal section feature association labeling model; if not, label the target disease The annotation results are input into the horizontal plane-longitudinal section feature association annotation model, and the horizontal plane-longitudinal section feature association annotation model is iteratively trained using the target disease annotation results to adjust the model parameters of the current round, obtain the preset disease annotation results of the next round and the model parameters of the next round, use the preset disease annotation results of the next round as the preset disease annotation results of the current round, use the model parameters of the next round as the model parameters of the current round, and return to the step of comparing the preset disease annotation results of the current round with the target disease annotation results, until the accuracy of the preset disease annotation results of the next round reaches the preset accuracy value, stop the iteration, and obtain the trained horizontal plane-longitudinal section feature association annotation model.

[0117] Compared with manual labeling, the automatic labeling of the present invention significantly improves the labeling efficiency, which is almost twice that of manual labeling. At the same time, it also performs well in accuracy, reaching more than 90%, which is slightly higher than the accuracy of manual labeling.

[0118] The embodiment of the present invention includes a capture unit, which is used to capture a horizontal plane view and a longitudinal section view from a visualized sequence of multiple views of a pavement structure layer; a splicing unit, which is used to splice the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the view; a first annotation unit, which is used to use a visual image annotation software to annotate the defects in the two-dimensional image across the view to obtain a corresponding annotation file, and to form a data pair with the annotation file and the two-dimensional image across the view, and to divide the data pair into a training set and a verification set; a construction unit, which is used to construct a horizontal plane-longitudinal section feature association annotation model, and a horizontal plane-longitudinal section feature association annotation model. Used to mark the pavement structure layer defects, the horizontal plane-longitudinal section feature association labeling model includes a backbone network and a head network, the backbone network is a network containing multiple multi-level feature fusion networks and spatial pyramid pooling layers; the training unit is used to train the horizontal plane-longitudinal section feature association labeling model using the training set to obtain the trained horizontal plane-longitudinal section feature association labeling model; the second labeling unit is used to input the verification set into the trained horizontal plane-longitudinal section feature association labeling model, use the trained horizontal plane-longitudinal section feature association labeling model to label the defects in the verification set, and output the defect labeling results. In this way, the pavement structure layer defects can be automatically labeled through the feature association labeling model horizontal plane-longitudinal section feature association labeling model trained in the training unit, which reduces the time and workload of manual labeling and improves the efficiency of pavement structure layer defect labeling. By utilizing the two-dimensional images of the cross-views in the splicing unit, the similar disease feature association information between adjacent horizontal plane views and longitudinal section views is considered to optimize the labeling performance of the feature association labeling model horizontal plane-longitudinal section feature association labeling model. In addition, the pavement structure layer defects are automatically labeled through data pairs and the trained feature association labeling model horizontal plane-longitudinal section feature association labeling model, which reduces the possibility of mislabeling and missing labels and can improve the accuracy of pavement structure layer defect labeling.

[0119] It should be pointed out here that the above description of the embodiment of the pavement structure layer disease labeling device under three-dimensional multi-view imaging is similar to the description of the embodiment of the pavement structure layer disease labeling method under three-dimensional multi-view imaging, and has similar beneficial effects as the embodiment of the pavement structure layer disease labeling method under three-dimensional multi-view imaging. For technical details not disclosed in the embodiment of the pavement structure layer disease labeling device under three-dimensional multi-view imaging of the embodiment of the present invention, please refer to the description of the embodiment of the pavement structure layer disease labeling method under three-dimensional multi-view imaging of the present invention for understanding.

[0120] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for marking pavement structure layer defects under three-dimensional multi-view imaging, characterized in that: The pavement structure layer disease labeling method under three-dimensional multi-view imaging includes: In the visualized sequence multi-view of the pavement structure layer, a horizontal plane view and a longitudinal section view are cut out; splicing the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the view; Using visual image annotation software, annotating the defects in the two-dimensional image across the view to obtain a corresponding annotation file, and forming a data pair with the annotation file and the two-dimensional image across the view, and dividing the data pair into a training set and a validation set; Constructing a horizontal plane-longitudinal section feature association annotation model, wherein the horizontal plane-longitudinal section feature association annotation model is used to annotate pavement structure layer defects, wherein the horizontal plane-longitudinal section feature association annotation model comprises a backbone network and a head network, wherein the backbone network is a network comprising a plurality of multi-level feature fusion networks and a spatial pyramid pooling layer; Using the training set to train the horizontal plane-longitudinal section feature association annotation model to obtain a trained horizontal plane-longitudinal section feature association annotation model; The verification set is input into the trained horizontal plane-longitudinal section feature association annotation model, the trained horizontal plane-longitudinal section feature association annotation model is used to annotate the defects in the verification set, and the defect annotation result is output.

2. The pavement structure layer disease labeling method under three-dimensional multi-view imaging according to claim 1 is characterized in that: Before extracting the horizontal plane view and the longitudinal section view from the visualized sequence multi-view, the method further includes: Obtain three-dimensional images of pavement structure layers; The three-dimensional image is converted into a visualized sequential multi-view, wherein the visualized sequential multi-view consists of the horizontal plane view, the longitudinal section view and the cross-section view.

3. The method for marking pavement structure layer defects under three-dimensional multi-view imaging according to claim 1, characterized in that: The visualized sequential multi-views include a plurality of horizontal pixel points of the horizontal plane view and a plurality of vertical pixel points of the longitudinal section view; The method of extracting a horizontal plane view and a longitudinal section view from the visualized sequence multi-view of the pavement structure layer includes: According to the plurality of horizontal pixel points, extracting the horizontal plane view from the visualized sequence of multiple views; The longitudinal section view is cut out in the visualized sequential multi-view according to the multiple vertical pixel points, the width of the horizontal plane view is the same as the width of the horizontal plane view and the height of the horizontal plane view is different from the height of the horizontal plane view.

4. The method for marking pavement structure layer defects under three-dimensional multi-view imaging according to claim 1, characterized in that: The backbone network includes a first convolutional layer, a first multi-level feature fusion network, a second convolutional layer, a second multi-level feature fusion network, a first up-down dimensionality feature extraction module, a third convolutional layer, a third multi-level feature fusion network, a second up-down dimensionality feature extraction module, a fourth convolutional layer, a fourth multi-level feature fusion network, a third up-down dimensionality feature extraction module, a fifth convolutional layer, a fifth multi-level feature fusion network, a fourth up-down dimensionality feature extraction module and the spatial pyramid pooling layer, which are connected in sequence; the first up-down dimensionality feature extraction module includes three sixth convolutional layers connected in sequence, the second up-down dimensionality feature extraction module includes six seventh convolutional layers connected in sequence, the third up-down dimensionality feature extraction module includes nine eighth convolutional layers connected in sequence, and the fourth up-down dimensionality feature extraction module includes three ninth convolutional layers connected in sequence.

5. The method for marking pavement structure layer defects under three-dimensional multi-view imaging according to claim 4, characterized in that: The first multi-level feature fusion network, the second multi-level feature fusion network, the third multi-level feature fusion network, the fourth multi-level feature fusion network, and the fifth multi-level feature fusion network all include an image block segmentation layer, an image embedding layer, multiple multi-level feature fusion blocks, and multiple image downsampling layers; each multi-level feature fusion block includes a first normalization layer, a window multi-head self-attention layer, a second normalization layer, a first multi-layer perception mechanism layer, a third normalization layer, a sliding window multi-head self-attention layer, a fourth normalization layer, and a second multi-layer perception mechanism layer connected in sequence.

6. The method for marking pavement structure layer defects under three-dimensional multi-view imaging according to claim 1, characterized in that: The head network includes a tenth convolutional layer, a first upsampling layer, a first connection layer, a fifth up-dimensionality feature extraction module, an eleventh convolutional layer, a second upsampling layer, a second connection layer, a sixth up-dimensionality feature extraction module, a twelfth convolutional layer, a third connection layer, a seventh up-dimensionality feature extraction module, a thirteenth convolutional layer, a fourth connection layer, an eighth up-dimensionality feature extraction module and a detection layer, which are connected in sequence.

7. The method for marking pavement structure layer defects under three-dimensional multi-view imaging according to claim 1, characterized in that: The step of training the horizontal plane-longitudinal section feature association annotation model using the training set to obtain a trained horizontal plane-longitudinal section feature association annotation model includes: Iteratively training the horizontal plane-longitudinal section feature association annotation model using the training set to obtain a preset disease annotation result of the current round and a model parameter of the current round; Obtaining a target disease labeling result, wherein the target disease labeling result is a result of labeling the same disease as the preset disease labeling result of the current round; Comparing the preset disease labeling result of the current round with the target disease labeling result to determine the accuracy of the preset disease labeling result of the current round; Determine whether the accuracy of the preset disease marking result of the current round reaches a preset accuracy value; If yes, stop the iteration to obtain the trained horizontal plane-longitudinal section feature association annotation model; If not, the target disease labeling result is input into the horizontal plane-longitudinal section feature association labeling model, and the horizontal plane-longitudinal section feature association labeling model is iteratively trained using the target disease labeling result to adjust the model parameters of the current round, obtain the preset disease labeling result of the next round and the model parameters of the next round, use the preset disease labeling result of the next round as the preset disease labeling result of the current round, use the model parameters of the next round as the model parameters of the current round, and return to the step of comparing the preset disease labeling result of the current round with the target disease labeling result, until the accuracy of the preset disease labeling result of the next round reaches the preset accuracy value, then stop the iteration to obtain the trained horizontal plane-longitudinal section feature association labeling model.

8. The method for marking pavement structure layer defects under three-dimensional multi-view imaging according to claim 4, characterized in that: The iterative training of the horizontal plane-longitudinal section feature association annotation model using the training set to obtain the preset disease annotation results of the current round and the model parameters of the current round includes: Inputting the training set into the backbone network so that the backbone network outputs a plurality of intermediate features; The multiple intermediate features are input into the head network so that the head network outputs the preset disease labeling result of the current round and the model parameters of the current round.

9. The method for marking pavement structure layer defects under three-dimensional multi-view imaging according to claim 8, characterized in that: The multiple intermediate features include a first intermediate feature, a second intermediate feature, and a third intermediate feature, and the inputting the training set into the backbone network so that the backbone network outputs the multiple intermediate features includes: Inputting the training set into the first convolution layer, the first multi-level feature fusion network, the second convolution layer, the second multi-level feature fusion network, the first dimension-up and dimension-down feature extraction module, the third convolution layer and the third multi-level feature fusion network in sequence, so that the third multi-level feature fusion network outputs the first intermediate feature; Inputting the first intermediate feature into the second up-down dimension feature extraction module, the fourth convolutional layer and the fourth multi-level feature fusion network in sequence, so that the fourth multi-level feature fusion network outputs the second intermediate feature; The second intermediate feature is sequentially input into the third dimensionality increase / decrease feature extraction module, the fifth convolution layer, the fifth multi-level feature fusion network, the fourth dimensionality increase / decrease feature extraction module and the spatial pyramid pooling layer, so that the spatial pyramid pooling layer outputs the third intermediate feature.

10. A pavement structure layer disease labeling device under three-dimensional multi-view imaging, characterized in that: The pavement structure layer disease marking device under three-dimensional multi-view imaging includes: A clipping unit, used for clipping out a horizontal plane view and a longitudinal section view from a sequence of multiple views of a visualized pavement structure layer; A splicing unit, used for splicing the horizontal plane view and the longitudinal section view to obtain a two-dimensional image across the view; A first labeling unit is used to label the defects in the two-dimensional image across the view using visual image labeling software to obtain a corresponding labeling file, and to form a data pair with the labeling file and the two-dimensional image across the view, and to divide the data pair into a training set and a validation set; A construction unit is used to construct a horizontal plane-longitudinal section feature association annotation model, wherein the horizontal plane-longitudinal section feature association annotation model is used to annotate pavement structure layer defects, wherein the horizontal plane-longitudinal section feature association annotation model includes a backbone network and a head network, wherein the backbone network is a network including multiple multi-level feature fusion networks and a spatial pyramid pooling layer; A training unit, used to train the horizontal plane-longitudinal section feature association annotation model using the training set to obtain a trained horizontal plane-longitudinal section feature association annotation model; The second labeling unit is used to input the verification set into the trained horizontal plane-longitudinal section feature association labeling model, use the trained horizontal plane-longitudinal section feature association labeling model to label the defects in the verification set, and output the defect labeling result.

Citation Information

Patent Citations

  • CT image lung lobe image segmentation system based on attention mechanism

    CN113936011A

  • Construction method and use method of three-dimensional data set of internal diseases of road

    CN117079268A