Road disease segmentation method and device based on multi-scale cross-layer attention fusion, equipment and medium
The road defect segmentation method using multi-scale cross-layer attention fusion solves the problem of unclear segmentation of small cracks in complex backgrounds, achieves higher accuracy in road defect detection, and improves the topological integrity and edge detail clarity of the segmentation results.
Patent Information
- Application Number
- CN202511053404.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing road crack detection algorithms struggle to accurately segment small cracks in complex environments, and the segmentation results lack topological integrity and have blurred edge details.
A road defect segmentation method using multi-scale cross-layer attention fusion is adopted. By introducing a multi-scale strip pyramid module to suppress background noise, and using a cross-layer directional attention module and a crack boundary refinement module to optimize edge details, the segmentation accuracy is enhanced.
It effectively suppresses background noise, preserves the features of deep micro-cracks, alleviates the loss of topological information, improves the segmentation accuracy of fine cracks, and enhances the global topology modeling capability.
Smart Images

Figure BDA0005524021950000061 
Figure BDA0005524021950000093 
Figure BDA0005524021950000111
Abstract
Description
Technical Field
[0001] This application relates to the field of road defect detection, and in particular to a road defect segmentation method, device, equipment and medium based on multi-scale cross-layer attention fusion. Background Technology
[0002] Accurate road defect detection is crucial for maintaining road quality and ensuring traffic safety. Existing road crack detection algorithms suffer from two major problems:
[0003] (1) Fine cracks are not clearly characterized in complex backgrounds and are easily confused with noise.
[0004] (2) The topological integrity of the segmentation results is insufficient and the edge details are blurred. Summary of the Invention
[0005] The purpose of this application is to provide a road defect segmentation method, device, equipment, and medium based on multi-scale cross-layer attention fusion, which can improve the segmentation accuracy of road defects.
[0006] To achieve the above objectives, this application provides the following solution:
[0007] Firstly, this application provides a road defect segmentation method based on multi-scale cross-layer attention fusion, including:
[0008] Obtain images of road defects to be segmented;
[0009] The road defect image to be segmented is input into the trained road crack segmentation model for road defect segmentation. The road crack segmentation model includes an encoder and a decoder. The encoder includes multiple encoder sub-modules, and the decoder includes multiple decoder sub-modules and multiple crack boundary refinement sub-modules. The encoder and decoder sub-modules are connected by a cross-layer directional attention fusion module. The crack boundary refinement sub-module is used to refine the boundary of the feature map output by the decoder sub-module.
[0010] Secondly, this application provides a road defect segmentation device based on multi-scale cross-layer attention fusion, comprising:
[0011] The acquisition module is used to acquire images of road defects to be segmented.
[0012] The road defect segmentation module is used to input the road defect image to be segmented into the trained road crack segmentation model for road defect segmentation. The road crack segmentation model includes an encoder and a decoder. The encoder includes multiple encoder sub-modules, and the decoder includes multiple decoder sub-modules and multiple crack boundary refinement sub-modules. The encoder and decoder sub-modules are connected by a cross-layer directional attention fusion module. The crack boundary refinement sub-module is used to refine the boundary of the feature map output by the decoder sub-modules.
[0013] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the road defect segmentation method based on multi-scale cross-layer attention fusion as described above.
[0014] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the road defect segmentation method based on multi-scale cross-layer attention fusion as described above.
[0015] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0016] This application provides a road defect segmentation method, device, equipment, and medium based on multi-scale cross-layer attention fusion. The method segments road defects using a road crack segmentation model, which incorporates a multi-scale strip pyramid module to effectively suppress background noise and preserve deep, minute crack features. A cross-layer directional attention module dynamically aggregates contextual features from each stage of the encoder and decoder, mitigating the loss of overall crack topology information caused by downsampling and enhancing global topology modeling capabilities. Finally, a crack boundary refinement sub-module optimizes multi-scale edge details, thereby improving the segmentation accuracy of fine cracks. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is an application environment diagram of a road defect segmentation method based on multi-scale cross-layer attention fusion in one embodiment of this application;
[0019] Figure 2A flowchart illustrating a road defect segmentation method based on multi-scale cross-layer attention fusion, provided as an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of the network structure of a road crack segmentation model provided in one embodiment of this application;
[0021] Figure 4 This is a schematic diagram of the structure of a multi-scale strip pyramid module provided in an embodiment of this application;
[0022] Figure 5 This is a schematic diagram of the structure of a cross-layer directional attention fusion module provided in an embodiment of this application;
[0023] Figure 6 A schematic diagram of the structure of a first directional attention mechanism unit provided in an embodiment of this application;
[0024] Figure 7 This is a schematic diagram of the structure of the first feature recombination unit provided in an embodiment of this application;
[0025] Figure 8 A schematic diagram of the structure of a crack boundary refinement module provided in an embodiment of this application.
[0026] Figure 9 A schematic diagram of the functional modules of a road defect segmentation device based on multi-scale cross-layer attention fusion is provided for another embodiment of this application;
[0027] Figure 10 This application provides a schematic diagram of the structure of a computer device according to one embodiment. Detailed Implementation
[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] The road defect segmentation method based on multi-scale cross-layer attention fusion provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on other servers. Terminal 102 can send the road defect image to be segmented to server 104. Server 104 receives the road defect image, acquires the road defect image, and inputs it into a trained road crack segmentation model for road defect segmentation. Server 104 can then feed back the obtained road defect segmentation result to terminal 102. Furthermore, in some embodiments, the road defect segmentation method based on multi-scale cross-layer attention fusion can also be implemented independently by server 104 or terminal 102. For example, terminal 102 can directly perform road defect segmentation on the road defect image, or server 104 can acquire the road defect image from the data storage system and perform road defect segmentation.
[0031] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.
[0032] In one exemplary embodiment, such as Figure 2 As shown, a road defect segmentation method based on multi-scale cross-layer attention fusion is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 202. Wherein:
[0033] Step 201: Obtain the road defect images to be segmented.
[0034] Step 202: Input the road defect image to be segmented into the trained road crack segmentation model to perform road defect segmentation.
[0035] The road crack segmentation model includes an encoder and a decoder. The encoder includes multiple encoder sub-modules, and the decoder includes multiple decoder sub-modules and multiple crack boundary refinement sub-modules. The encoder and decoder sub-modules are connected by a cross-layer directional attention fusion module (Clda). The crack boundary refinement sub-module (Br) is used to refine the boundary of the feature map output by the decoder sub-modules.
[0036] Step 202 specifically includes steps 301-303:
[0037] Step 301: The road defect image to be segmented is downsampled layer by layer by the encoder to obtain feature maps of different sizes.
[0038] For the i-th encoding submodule, the input is F i-1 The output is F i F i-1 and F i These are the feature maps output by the (i-1)th and ith encoding submodules, respectively; when i equals 1, F0 is the road defect image to be segmented.
[0039] When i = 1, it is the top-level encoding submodule, which includes the first convolutional module and the second convolutional module connected in sequence; when i ≥ 2, the i-th encoding submodule includes the first convolutional module, the second convolutional module, the max pooling module and the multi-scale strip pyramid module connected in sequence.
[0040] The processing steps of the top-level encoding submodule include:
[0041] 1. Input the road defect image to be segmented into the first convolutional module for feature extraction to obtain the feature map F. 11 .
[0042] 2. Transfer feature map F 11 The input is fed into the second convolutional module for feature extraction to obtain the feature map F. 12 .
[0043] When i ≥ 2, the processing procedure for the i-th encoding submodule specifically includes:
[0044] 1. Transfer the feature map F i-1 The input is fed into the first convolutional module for feature extraction to obtain the feature map F. i1 .
[0045] 2. Transfer feature map F i1 The input is fed into the second convolutional module for feature extraction to obtain the feature map F. i2.
[0046] 3. Transfer feature map F i2 The input is fed into the max pooling module to obtain the feature map F. i3 .
[0047] 4. Transfer feature map F i3 The feature map F is obtained by inputting the multi-scale strip pyramid module for feature extraction. i .
[0048] In one exemplary embodiment, such as Figure 3 As shown, the encoder contains five encoding sub-modules, with i equal to 5. The top-level encoding sub-module consists of a first convolutional module and a second convolutional module connected in sequence. Both the first and second convolutional modules include a convolutional layer (Conv), a batch normalization (Bn) operation, and a corrected linear layer (ReLU). The remaining four encoding sub-modules have the same structure, each consisting of a first convolutional module, a second convolutional module, a max pooling module, and a multiscale strip pyramid module (Mstp) connected in sequence.
[0049] like Figure 4 As shown, Mstp includes a first branch structure unit, a second branch structure unit, an average pooling unit, and a channel attention unit; the first branch structure unit includes a hollow pyramid sub-unit and a first convolution sub-unit, and the second branch structure unit includes strip convolution sub-units and second convolution sub-units in multiple directions.
[0050] The dual-branch feature extraction module is composed of a first branch structure unit and a second branch structure unit. The hollow pyramid sub-unit in the first branch structure unit includes three 3x3 depthwise convolutions with dilation rates of 3, 6 and 12, respectively. By constructing the hollow pyramid sub-unit, the receptive field is increased while keeping the feature map size unchanged, thereby improving the feature extraction capability.
[0051] In the second branch structure unit, four-directional strip convolution sub-units are used to capture global context information from different directions.
[0052] Where, X∈R H×W×C Let H be the input feature map, and H, W, and C be the height, width, and number of channels of the input tensor, respectively. W∈R 2K+1 The size of the strip convolution filter is D = (D h D w ) represents the direction of filter W, and D h D w These represent the filter's direction in height and direction in width, respectively; Z D ∈RH×W×C The result of strip convolution can be defined as:
[0053]
[0054] Where [i, j] are the position coordinates of the currently calculated pixel, X is the input feature map, * indicates convolution operation, D is the direction vector of the strip convolution sub-unit, and the convolutions of the horizontal, vertical, left diagonal and right diagonal are (0,1), (1,0), (1,1) and (-1,1) respectively; l is the position index of sliding along direction D, k is the parameter that controls the length of the strip convolution sub-unit, and w is the filter parameter. For filter W, the parameter k is set to 4.
[0055] In the encoder section, Mstp adaptively extracts the local linear structure and global directional distribution features of cracks through strip convolution sub-units and multi-scale strip convolution, thereby improving the road crack segmentation model's ability to perceive complex disease morphologies.
[0056] As the encoder deepens, the number of feature map channels increases, and the contribution of different channels to crack semantics varies significantly. Some channels focus on crack texture details, while most are dominated by complex background information. To enhance the weight of crack semantics in channel features and suppress interference from irrelevant regions, inspired by the Squeeze-and-Excitation (SE) channel attention mechanism, this application employs channel attention units to learn weights for different channels, emphasizing the importance of crack features and thus compensating for the tendency to lose subtle crack features. Specifically, global average pooling is performed on each channel feature to capture the global distribution between channels. A nonlinear activation function is used to learn the nonlinear relationship between channels to generate weights for each channel. The generated weights are then multiplied channel-by-channel with the original feature map to enhance the feature response of important channels.
[0057] Step 302: Dynamically aggregate the context features of each encoding sub-module of the encoder through the cross-layer directional attention fusion module.
[0058] The number of cross-layer directional attention fusion modules is one less than the number of encoding submodules. When there are 5 encoding submodules, there are four cross-layer directional attention fusion modules.
[0059] like Figure 5 As shown, the cross-layer directional attention fusion module includes a first directional attention mechanism unit, a second directional attention mechanism unit, a first feature reorganization unit, a second feature reorganization unit, and a convolution unit.
[0060] The processing procedure for the i-th cross-layer directional attention fusion module is as follows:
[0061] 1) Transfer feature map Fi The input is fed into the first directional attention mechanism unit for feature enhancement, resulting in feature map F. i '.
[0062] 2) Transfer feature map F i+1 The input is fed into the second-direction attention mechanism unit for feature enhancement, resulting in feature map F'. i+1 ; the feature map F' i+1 The input is fed into a convolutional unit for feature extraction, resulting in a feature map F”. i+1 .
[0063] 3) Transfer feature map F i 'and feature map F' i+1 The input is fed into the first feature recombination unit for feature recombination, resulting in the recombined feature map Q. i .
[0064] 4) Transfer the feature map Q i and feature map E i The input is fed into the second feature recombination unit for feature recombination, resulting in feature map T. i Among them, feature map E i This is the feature map output by the i-th decoding submodule.
[0065] The first directional attention mechanism unit and the second directional attention mechanism unit have the same structure. For different levels of features of the encoder, the first directional attention mechanism unit and the second directional attention mechanism unit introduce a collaborative mechanism of directional sensitive feature extraction and attention to achieve refined feature enhancement of directional perception.
[0066] In an exemplary embodiment, for encoder features, this application achieves direction-aware feature decoupling and fusion by constructing a three-branch structure. The first direction attention mechanism unit includes a first branch feature extraction subunit, a second branch feature extraction subunit, a third branch feature extraction subunit, and a fusion subunit.
[0067] like Figure 6 As shown, the first branch feature extraction subunit includes a multi-dimensional strip convolutional layer, a product fusion layer, and a first convolutional layer in the horizontal direction. In the first branch feature extraction subunit, XCoorDonv (XCoordinateConvolution, horizontal coordinate convolution) is used to introduce spatial coordinates, improving the positional sensitivity of the feature map in the horizontal dimension. Secondly, a multi-dimensional strip convolutional layer in the horizontal direction is used to capture local details and global dependencies in the horizontal direction with multiple receptive fields. The encoder features are multiplied and fused with the multi-dimensional horizontal features, and the fused horizontal features are passed through the first convolutional layer to obtain the horizontal attention Q. x K x .
[0068] The second branch feature extraction subunit comprises a multi-scale strip pyramid module and a second convolutional layer. In this subunit, features are extracted using the multi-scale strip pyramid module, and then unified in dimensionality by the second convolutional layer to obtain multi-dimensional features V. Cross-directional attention operations are then performed on these multi-dimensional features to enhance features in different directions.
[0069] The third branch feature extraction subunit includes a multi-dimensional strip convolutional layer, a product fusion layer, and a third convolutional layer in the vertical horizontal direction. In this subunit, Y CoordConv (Y CoordinateConvolution, vertical coordinate convolution) and the multi-dimensional strip convolutional layer in the vertical horizontal direction are used to extract multi-scale features in the vertical direction. The processing flow is consistent with the horizontal branch. The fused vertical features are then passed through the third convolutional layer to obtain the vertical dimension attention Q. y K y .
[0070] The specific calculation formula is as follows:
[0071]
[0072]
[0073]
[0074] Where X represents the encoder feature, conv 1*i For multi-scale horizontal convolution, conv i*1 For multi-scale vertical convolution, coordonv x ,conrdonv y These are coordinate convolutions in the X and Y directions, respectively. Mstp is a multi-scale strip pyramid layer. attention For output features.
[0075] The encoder expands the receptive field through layer-by-layer downsampling, but continuous spatial compression leads to the loss of pixel-level positional information. The decoder gradually restores spatial resolution through upsampling, but this introduces noise information. Direct feature concatenation easily amplifies noisy features at each level and generates feature redundancy. Therefore, this application proposes a dynamic feature reassembly mechanism, namely... Figure 5 The first feature recombination unit and the second feature recombination unit in the structure are identical.
[0076] like Figure 7 As shown, the first feature recombination unit will reconstruct the feature map F i 'and feature map F' i+1Channel shuffle is used to generate fused features, promoting information interaction between different channel groups. The fused features are then used to generate an adaptive spatial attention weight matrix, accurately quantifying the saliency of different pixel regions. Channel-dimensional pooling operations (i.e., channel max pooling and channel average pooling) are then employed to obtain the semantic relevance between different channels, ultimately yielding the recombined feature map Q. i This enables the recombination of features in both the channel and spatial dimensions.
[0077] In skip connections, Clda mitigates the problem of loss of overall topological information in cracks caused by downsampling by introducing orientation-sensitive spatial attention weights and dynamically aggregating contextual features of each stage of the encoder submodule.
[0078] Step 303: The feature map after aggregating context features is upsampled layer by layer and pixel-level feature optimization is performed by the decoder to obtain the road disease segmentation result.
[0079] The number of encoding submodules and decoding submodules is the same; the input of the bottom-level decoding submodule is the output of the bottom-level encoding submodule, and the input of other layer decoding submodules is the output of the corresponding cross-layer directional attention fusion module. The output of the i-th decoding submodule is the feature map E. i Step 303 specifically includes steps 401-403:
[0080] Step 401: Extract features from the input feature map using the i-th decoding submodule to obtain feature map E. i .
[0081] Step 402, refine each feature map E using the crack boundary refinement submodule. i Boundary refinement processing is performed to obtain multiple boundary refinement feature maps of different sizes.
[0082] Step 403: Input the feature map after stitching together the boundary refinement feature maps of each size into the crack boundary refinement submodule for boundary refinement processing again to obtain the road defect segmentation result.
[0083] The crack boundary refinement module (Br) refines the boundaries and calculates the loss on feature maps of different scales output by the decoder, preventing the gradual loss of micro-texture details and overall topological features. Specifically, Br is designed as a residual structure, refining the boundaries by calculating the residual features between the label and the coarse feature map. The i-th decoder submodule outputs E. i ∈R H*W*CIntermediate features are obtained through 1x1 convolution and sigmoid activation. Then, a residual structure consisting of 3x3 convolution, batch normalization, and ReLU activation is used. Finally, the output features are directly compared with the labels for loss calculation. By refining the boundaries in each decoder submodule, the ability of the road crack segmentation model to extract the overall topology of small cracks is improved. The processing of the crack boundary refinement module is illustrated using any decoder submodule as an example. Figure 8 As shown, the input feature E(N×C×H×W) of the crack boundary refinement module is N, where N is the batch size, C is the number of channels, H is the height of the feature map, and W is the width of the feature map. The output feature is E', and the specific calculation formula is as follows:
[0084] E1=δ s B n (f conv1 (E)) (6);
[0085] E' = f shout-cut (E1)+E1 (7);
[0086] Among them, f shot-cut Represents the residual structure, δ s Represents the sigmoid activation function, f conv1 Represents a 1x1 convolution operation, B n E1 represents batch normalization, and E1 is an intermediate feature.
[0087] The loss function is a combination of binary cross-entropy loss function (BCE) and Dice loss, used as the loss function to train the network. The definitions of BCE loss and Dice loss are shown in Equations (8) and (9):
[0088]
[0089] Among them, L BCE For binary cross-entropy loss, L Dice The loss function is the Dice coefficient, where n is the training batch size, N is the number of pixels, and p... i It is the predicted segmentation result, t i It's a real label.
[0090] In the multi-scale output stage, different weights are used for outputs at different scales, and the total loss function is defined as shown in formula (10):
[0091]
[0092] Where L is the total loss, ω zrepresents the weight coefficients of the z-th Br output layer.
[0093] This application introduces a multi-scale strip pyramid module to effectively suppress background noise while preserving deep, minute crack features. A cross-layer directional attention module designs a directional attention mechanism between the encoder and decoder to enhance global topology modeling capabilities. A boundary refinement module optimizes edge details at multiple scales through a residual learning strategy. Validation on multiple datasets demonstrates that the road defect segmentation method based on multi-scale cross-layer attention fusion provided in this application outperforms mainstream algorithms in segmentation accuracy, exhibiting strong robustness, especially in low-contrast and minute crack scenarios. This method provides a new technical path for intelligent road maintenance, and its modular design can be extended to related tasks such as asphalt aging and pothole detection.
[0094] Based on the same inventive concept, this application also provides a road defect segmentation device for implementing the aforementioned multi-scale cross-layer attention fusion-based method. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the road defect segmentation device based on multi-scale cross-layer attention fusion provided below can be found in the limitations of the road defect segmentation method based on multi-scale cross-layer attention fusion described above, and will not be repeated here.
[0095] In one exemplary embodiment, such as Figure 9 As shown, a road defect segmentation device based on multi-scale cross-layer attention fusion is provided, comprising:
[0096] Module 91 is used to acquire images of road defects to be segmented.
[0097] The road defect segmentation module 92 is used to input the road defect image to be segmented into the trained road crack segmentation model for road defect segmentation; wherein, the road crack segmentation model includes an encoder and a decoder, the encoder includes multiple encoder sub-modules, the decoder includes multiple decoder sub-modules and multiple crack boundary refinement sub-modules, and the encoding sub-modules and decoding sub-modules are connected by a cross-layer directional attention fusion module; the crack boundary refinement sub-module is used to refine the boundary of the feature map output by the decoder sub-modules.
[0098] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 10As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores images of road defects to be segmented. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a road defect segmentation method based on multi-scale cross-layer attention fusion.
[0099] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0100] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0101] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0102] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0103] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0104] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0105] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0106] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A road defect segmentation method based on multi-scale cross-layer attention fusion, characterized in that, include: Obtain images of road defects to be segmented; The road defect image to be segmented is input into the trained road crack segmentation model for road defect segmentation. The road crack segmentation model includes an encoder and a decoder. The encoder includes multiple encoder sub-modules, and the decoder includes multiple decoder sub-modules and multiple crack boundary refinement sub-modules. The encoder and decoder sub-modules are connected by a cross-layer directional attention fusion module. The crack boundary refinement sub-module is used to refine the boundary of the feature map output by the decoder sub-module.
2. The road defect segmentation method based on multi-scale cross-layer attention fusion according to claim 1, characterized in that, The road defect images to be segmented are input into the trained road crack segmentation model for road defect segmentation, specifically including: The road defect image to be segmented is downsampled layer by layer by the encoder to obtain feature maps of different sizes; The contextual features of each encoding sub-module of the encoder are dynamically aggregated through a cross-layer directional attention fusion module; The road defect segmentation results are obtained by upsampling the feature map after aggregating context features layer by layer and performing pixel-level feature optimization through the decoder.
3. The road defect segmentation method based on multi-scale cross-layer attention fusion according to claim 2, characterized in that, For the i-th encoding submodule, the input is F i-1 The output is F i F i-1 and F i These are the feature maps output by the (i-1)th and ith encoding submodules, respectively. When i equals 1, F0 is the road defect image to be segmented; When i = 1, it is the top-level encoding submodule, which includes a first convolutional module and a second convolutional module connected in sequence; when i ≥ 2, the i-th encoding submodule includes a first convolutional module, a second convolutional module, a max pooling module, and a multi-scale strip pyramid module connected in sequence. The processing steps of the top-level encoding submodule include: The road defect image to be segmented is input into the first convolutional module for feature extraction to obtain the feature map F. 11 ; Feature map F 11 The input is fed into the second convolutional module for feature extraction to obtain the feature map F. 12 ; When i ≥ 2, the processing procedure for the i-th encoding submodule specifically includes: Feature map F i-1 The input is fed into the first convolutional module for feature extraction to obtain the feature map F. i1 ; Feature map F i1 The input is fed into the second convolutional module for feature extraction to obtain the feature map F. i2 ; Feature map F i2 The input is fed into the max pooling module to obtain the feature map F. i3 ; Feature map F i3 The feature map F is obtained by inputting the multi-scale strip pyramid module for feature extraction. i .
4. The road defect segmentation method based on multi-scale cross-layer attention fusion according to claim 3, characterized in that, The multi-scale strip pyramid module includes a first branch structure unit, a second branch structure unit, an average pooling unit, and a channel attention unit; the first branch structure unit includes a hollow pyramid sub-unit and a first convolution sub-unit, and the second branch structure unit includes strip convolution sub-units and second convolution sub-units in multiple directions.
5. The road defect segmentation method based on multi-scale cross-layer attention fusion according to claim 3, characterized in that, The cross-layer directional attention fusion module has one less element than the coding submodule. The cross-layer directional attention fusion module includes a first directional attention mechanism unit, a second directional attention mechanism unit, a first feature reorganization unit, a second feature reorganization unit, and a convolutional unit; The processing procedure for the i-th cross-layer directional attention fusion module is as follows: Feature map F i The input is fed into the first directional attention mechanism unit for feature enhancement, resulting in feature map F. i '; Feature map F i+1 The input is fed into the second-direction attention mechanism unit for feature enhancement, resulting in feature map F′. i+1 ; the feature map F′ i+1 The input is fed into a convolutional unit for feature extraction, resulting in a feature map F″. i+1 ; Feature map F i 'and feature map F' i+1 The input is fed into the first feature recombination unit for feature recombination, resulting in the recombined feature map Q. i ; feature map Q i and feature map E i The input is fed into the second feature recombination unit for feature recombination, resulting in feature map T. i Among them, feature map E i This is the feature map output by the i-th decoding submodule.
6. The road defect segmentation method based on multi-scale cross-layer attention fusion according to claim 5, characterized in that, The first directional attention mechanism unit and the second directional attention mechanism unit have the same structure. The first directional attention mechanism unit includes a first branch feature extraction subunit, a second branch feature extraction subunit, a third branch feature extraction subunit, and a fusion subunit. The first branch feature extraction subunit includes a multi-dimensional strip convolutional layer, a product fusion layer, and a first convolutional layer in the horizontal direction. The second branch feature extraction subunit includes a multi-scale strip pyramid module and a second convolutional layer. The third branch feature extraction subunit includes a multi-dimensional strip convolutional layer, a product fusion layer, and a third convolutional layer in the vertical direction.
7. The road defect segmentation method based on multi-scale cross-layer attention fusion according to claim 3, characterized in that, The number of encoding submodules and decoding submodules is the same; the input of the bottom-level decoding submodule is the output of the bottom-level encoding submodule, and the input of other layer decoding submodules is the output of the corresponding cross-layer directional attention fusion module. The output of the i-th decoding submodule is the feature map E. i ; Specifically, the road defect segmentation result is obtained by upsampling the feature map after aggregating context features layer by layer and performing pixel-level feature optimization through a decoder, which includes: The i-th decoding submodule extracts features from the input feature map to obtain feature map E. i ; The crack boundary refinement submodule is used to refine each feature map E. i Boundary refinement processing is performed to obtain multiple boundary refinement feature maps of different sizes; The feature map after stitching together the refined feature maps of each size is input into the crack boundary refinement submodule for further boundary refinement processing to obtain the road defect segmentation result.
8. A road defect segmentation device based on multi-scale cross-layer attention fusion, characterized in that, include: The acquisition module is used to acquire images of road defects to be segmented. The road defect segmentation module is used to input the road defect image to be segmented into the trained road crack segmentation model for road defect segmentation. The road crack segmentation model includes an encoder and a decoder. The encoder includes multiple encoder sub-modules, and the decoder includes multiple decoder sub-modules and multiple crack boundary refinement sub-modules. The encoder and decoder sub-modules are connected by a cross-layer directional attention fusion module. The crack boundary refinement sub-module is used to refine the boundary of the feature map output by the decoder sub-modules.
9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the road defect segmentation method based on multi-scale cross-layer attention fusion as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the road defect segmentation method based on multi-scale cross-layer attention fusion as described in any one of claims 1-7.