Road disease detection method based on attention mechanism and feature integration
By employing an attention-based and feature-integrated approach, the specificity issues in road defect detection are addressed, improving detection accuracy and timeliness, and making the method applicable to most equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-08
- Publication Date
- 2026-03-31
AI Technical Summary
Due to specificity issues, existing road defect detection technologies, including traditional target detection techniques, struggle to effectively extract the object features of road defects, resulting in insufficient detection accuracy and timeliness.
This paper adopts an attention mechanism and feature integration method. It preprocesses road images, generates a feature map set using a feature extraction network, focuses features through channel and spatial attention mechanisms, integrates features with pyramid convolution, and finally generates detection results. The detection accuracy is improved by optimizing network parameters.
It improves the accuracy and timeliness of road defect detection, and enables simple and highly portable defect detection that is applicable to most equipment.
Smart Images

Figure CN115578317B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of road defect detection, and specifically to a road defect detection method and apparatus based on attention mechanisms and feature integration. Background Technology
[0002] With the rapid development of my country's economic construction, road construction has flourished, people's living standards have significantly improved, and the number of cars has increased dramatically, facilitating travel and transportation. However, this has also led to a significant increase in traffic volume and severe overloading, making road surface damage unavoidable and causing frequent road defects. This necessitates timely road maintenance. Therefore, improving the accuracy and timeliness of road defect detection to enable early detection and treatment is of paramount importance for extending road lifespan and improving vehicle travel experience and safety.
[0003] Currently, commonly used road inspection methods mainly include manual measurement and statistics, and road surface inspection vehicles. Manual measurement and statistics are costly, inefficient, and dangerous, and also disrupt normal road operations. With the development of computer technology, road defect detection is moving towards automation. The characteristics of objects detected in road defect detection differ significantly from those in traditional target detection, but existing road inspection technologies often use general target detection techniques. Therefore, facing the issue of the specificity of objects detected in road defect detection, traditional target detection technologies extract object features with different focuses, making it difficult to extract the object features of road defects, resulting in a significant performance degradation during defect detection. Summary of the Invention
[0004] The technical problem this invention aims to solve is the inadequacy of traditional target detection techniques due to the specific characteristics of road defects. This invention proposes a road defect detection method and apparatus based on attention mechanisms and feature integration.
[0005] The technical solution adopted by this invention to solve its technical problem is: a road defect detection method based on attention mechanism and feature integration, the method comprising: acquiring road images; preprocessing the road images; extracting features from the preprocessed road images to generate a first feature map set; performing image feature focusing processing on the first feature map set to generate a second feature map set; integrating the second feature map set to generate a third feature map set; and generating a detection result based on the third feature map set and a predetermined detector.
[0006] Furthermore, after generating the detection result based on the third feature map set and the predetermined detector, the method further includes: optimizing network parameters based on the detection result and the real label; and training the predetermined detector based on the optimized network parameters.
[0007] Furthermore, the preprocessing of the road images includes: generating a first road image training set based on the road images; and performing data transformation on the first road image training set to obtain a second road image training set.
[0008] Furthermore, generating the first road image training set based on the road images includes: dividing the road images into a first road image training set and a road image verification set at a predetermined ratio; the second road image training set includes the first road image training set and new road images after data transformation of the first road image training set; the data transformation includes flipping, cropping, brightness and contrast transformation operations.
[0009] Furthermore, the step of extracting features from the preprocessed road images to generate the first feature map set includes: processing the second road image training set using a feature extraction network to obtain a feature map set extracted by each sub-network.
[0010] Furthermore, the step of performing image feature focusing processing on the first feature map set to generate the second feature map set includes: processing the first feature map set through a channel attention mechanism and a spatial attention mechanism to generate the second feature map set; wherein, the channel attention mechanism uses the SE attention mechanism, and the spatial attention mechanism uses the spatial attention mechanism that shares a multilayer perceptron with the SE.
[0011] Furthermore, generating detection results based on the third feature map set and the predetermined detector includes: detecting the category and location of road defects based on the third feature map set and the predetermined detector.
[0012] The present invention also provides a road defect detection device based on attention mechanism and feature integration, comprising: an acquisition module adapted to acquire road images; a preprocessing module adapted to preprocess the road images; a feature extraction module adapted to extract features from the preprocessed road images to generate a first feature map set; a feature focusing processing module adapted to perform image feature focusing processing on the first feature map set to generate a second feature map set; an integration module adapted to integrate the second feature map set to generate a third feature map set; and a detection module adapted to generate a detection result based on the third feature map set and a predetermined detector.
[0013] The present invention also provides a computer-readable storage medium storing one or more instructions, wherein a processor of a risk analysis device within the one or more instructions executes the road defect detection method based on attention mechanism and feature integration as described above.
[0014] The present invention also provides an electronic device, comprising: a memory and a processor; the memory storing at least one program instruction; the processor loading and executing the at least one program instruction to implement the road defect detection method based on attention mechanism and feature integration as described above.
[0015] The beneficial effects of this invention are as follows: This invention provides a road defect detection method based on attention mechanism and feature integration, comprising: acquiring road images; preprocessing the road images; extracting features from the preprocessed road images to generate a first feature map set; performing image feature focusing processing on the first feature map set to generate a second feature map set; integrating the second feature map set to generate a third feature map set; and generating a detection result based on the third feature map set and a predetermined detector. Compared with traditional target detection techniques, the road defect detection method based on attention mechanism and feature integration proposed in this invention can address the specificity of road defect features, fully utilize image information, extract and mix defect features, thereby improving the accuracy of defect detection. Furthermore, by continuously optimizing and training the detection model using a large number of road images, this invention can significantly improve timeliness while effectively ensuring detection accuracy. At the same time, the method proposed in this invention is simple to implement, highly portable, and applicable to most devices. Attached Figure Description
[0016] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0017] Figure 1 This is a flowchart of the road defect detection method based on attention mechanism and feature integration provided in the embodiments of the present invention.
[0018] Figure 2 This is a schematic diagram of the road defect detection device based on attention mechanism and feature integration provided in the embodiments of the present invention.
[0019] Figure 3 This is a network structure diagram of the feature focusing module in this invention.
[0020] Figure 4 This is a network structure diagram of the feature integration module in this invention.
[0021] Figure 5 This is a partial block diagram of the electronic device provided in the embodiments of the present invention. Detailed Implementation
[0022] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0023] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0024] The present invention will now be described in detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0025] Example 1
[0026] Please see Figure 1 The road defect detection method proposed in this invention, based on attention mechanisms and feature integration, addresses the specificity of road defect features by fully utilizing image information to extract and blend defect features, thereby improving detection accuracy. Furthermore, by continuously optimizing and training the detection model using a large number of road images, this invention significantly improves timeliness while effectively ensuring detection accuracy. The proposed method is also simple to implement, highly portable, and applicable to most devices.
[0027] As an example, the road defect detection method based on attention mechanisms and feature integration includes the following steps:
[0028] S110: Obtain road images.
[0029] S120: Preprocess the road image.
[0030] As an example, the preprocessing of the road images includes: generating a first road image training set based on the road images, and performing data transformation on the first road image training set to obtain a second road image training set. Specifically, for the entire road defect dataset used, i.e., the collection of all acquired road images, the dataset is divided proportionally, with the training set accounting for 80% (this training set is the first road image training set) and the validation set accounting for 20%. Data transformation operations are performed on each sample in the training set: flipping, cropping, and brightness and contrast transformation operations to obtain a new dataset (this new dataset is the second road image training set). It should be noted that technicians can set different proportions to divide the dataset between the training and validation sets according to actual needs; no restrictions are imposed here.
[0031] Specifically, given a training set of all samples in the road image training set as {X} (the first road image training set), d data transformation operations, denoted as {R}, are selected to expand the dataset or perturb image features. A series of data transformation operations are then performed on all samples in the training set to obtain an expanded sample set {X'} (the second road image training set), where... i = {1, 2, ..., d}, R i R represents a data transformation operation, where i represents the number of the data transformation operation, and R represents the number of the data transformation operation. i (X) represents R i The resulting road images after processing. Preprocessing the acquired road images can increase the number of images and improve the stability of subsequent training.
[0032] S130: Extract features from the preprocessed road image to generate a first feature map set.
[0033] As an example, the step of extracting features from the preprocessed road image to generate a first feature map set includes: processing the second road image training set using a feature extraction network to obtain a feature map set extracted by each sub-network, wherein the feature extraction network includes ResNet and DarkNet. Specifically, the given image feature extraction network contains c sub-networks, denoted as {T}. The second road image training set is processed using the feature extraction network to obtain a feature map set {F} (the first feature map set) extracted by each sub-network, where F... i =T i (X'), i = {1, 2, ..., c}, T i Representative sub-image feature extraction network, F i Represents T i The result is F. i The sets are disjoint and their union is F, where i represents the corresponding index.
[0034] S140: Perform image feature focusing processing on the first feature map set to generate a second feature map set.
[0035] As an example, the step of performing image feature focusing processing on the first feature map set to generate the second feature map set includes: processing the first feature map set through a channel attention mechanism and a spatial attention mechanism to generate the second feature map set, wherein the channel attention mechanism uses the SE attention mechanism, and the spatial attention mechanism uses the spatial attention mechanism that shares a multilayer perceptron with the SE.
[0036] like Figure 3 As shown, more specifically, the above feature-focusing stage includes:
[0037] The feature focusing phase uses the channel attention mechanism A C Spatial attention mechanism A S This yields a set of feature maps {F'} (the second set of feature maps) focused on the features, where F' = A. S (A C (F)).
[0038] For the channel attention module, we have:
[0039] Given input features F∈R C×H×W Where C represents the number of channels, W represents the feature map width, and H represents the feature map height, the channel feature map F is obtained through max pooling and average pooling. C1 ,F C2 ∈R C×1×1 Using multilayer perceptrons MLP1 and MLP2 to study F C1 ,F C2 Dimensionality reduction and dimensionality increase are performed to mix semantic information from different dimensions, and the ReLU activation function is used to filter redundant information to obtain M. C1 M C2 ∈R C×1×1 M C1 M C2 The features are added together and normalized using the sigmoid function to obtain the channel feature map M. C Finally, the input feature F is compared with the channel feature map M. C The dot product yields the feature map U. The formula is as follows:
[0040]
[0041]
[0042]
[0043]
[0044] in This indicates that an average pooling operation is performed on the input features across the channels. This indicates that a max pooling operation is performed on the input features, where W1 is the parameter of the multilayer perceptron MLP1 and W2 is the parameter of the multilayer perceptron MLP2.
[0045] For the spatial attention module, we have:
[0046] Given input features U∈R C×H×W Where C represents the number of channels, W represents the feature map width, and H represents the feature map height. The feature map is dimensionality-reduced using a multilayer perceptron (MLP1) shared with the channel attention module, resulting in U. M1 ∈R C / r×W×H , where r represents the compression factor. Semantic information from different locations is fused using a convolutional kernel of size 7, and then normalized using the sigmoid function to obtain the spatial feature map M. S Finally, the input features U and the spatial feature map M are compared. S Dot product yields the feature map F'.
[0047]
[0048] A S (U)=sigmoid(f 7×7 (W1(U)))
[0049] Where W1 represents the parameters of the multilayer perceptron MLP1 shared with the channel attention mechanism, f 7×7 This represents a convolution operation with a kernel size of 7.
[0050] Spatial attention focuses on the positional information of the image, which can be described as positional information on a two-dimensional matrix. It focuses on one or more regions of the matrix, while channel attention focuses on the channel information of the image, focusing on one or more channels. Spatial attention processes the two-dimensional matrix (which includes information on length and width), while channel attention processes the channels. Therefore, these two attention mechanisms can process information in three dimensions of the image. By performing image feature focusing processing on the first feature map set to generate the second feature map set, the computer can find its "working area" during the processing of road images using computer vision. That is, when processing road images, the computer can target and process a specific or multiple key regions of the road image, without having to process all regions of the road image. This improves its information representation ability and increases the accuracy of the detection results.
[0051] S150: Integrate the second feature map set to generate a third feature map set.
[0052] As an example, the process of integrating the second feature map set to generate the third feature map set includes: integrating the second feature map set to generate the third feature map set through pyramid convolution.
[0053] Specifically, an image feature integration method, denoted as {G}, is designed for the second feature map set. This method is then applied to all samples in the second feature map set to obtain the integrated image feature map set {F”} (the third feature map set), where F” = G. i (F'), i = {1, 2, ..., c}, G i This represents the image feature integration method, where i represents the corresponding number.
[0054] like Figure 4 As shown, the third feature map set is generated by integrating the second feature map set using pyramid convolution as an example:
[0055] Process F2∈R using pyramid convolutions of different sizes and numbers of groups. C×H×W ,get Where C represents the number of channels, W represents the feature map width, and H represents the feature map height, with j = 3, 4, 5 and h = j-2. M3, M4, M5 are added to F3, F4, F5 and fed into the FPN structure. The smaller feature map is upsampled and concatenated with the larger feature map, then processed by a convolutional kernel of size 1 to fuse the features, resulting in feature map F.
[0056] In processing road images, both local and overall features are crucial. Therefore, in computer vision, we extract both overall and local features of the image and study them simultaneously, rather than separating them for separate analysis. Furthermore, we need to integrate local features and analyze different features at the same time. This simultaneous analysis of local and overall features can greatly improve the accuracy of road defect analysis.
[0057] S160: Generate detection results based on the third feature map set and the predetermined detector.
[0058] As an example, generating detection results based on the third feature map set and the predetermined detector includes: by inputting the third feature map set into the predetermined detector, the detector will automatically output a result vector, which includes the category, the position and size of the bounding box, and the confidence level.
[0059] As an example, after generating the detection results, the detector will be continuously optimized and trained based on these results. Specifically, the network parameters will be optimized based on the predicted results and the true labels, thereby continuously monitoring and improving the detection accuracy of the detector. The network parameters include all learnable parameters from steps S130 to S160. For example, in the detection task of this scheme, convolutional kernels are used to process images, and the learnable parameters refer to the parameters of the convolutional kernels.
[0060] The predetermined detector uses, but is not limited to, the YOLO detector. The detector uses feature maps F (the third feature map set) to predict road defects, and trains the network based on the prediction results and the ground truth labels. The network parameters are updated by optimizing the following objective function:
[0061]
[0062] In the formula, Y c (x j (x, w) represents a disease detection network, where w is the parameter of the detection network, and x... j This is the sample to be tested. j This represents the true label, which is known data obtained through continuous training by relevant technical personnel. loss() represents the loss function.
[0063] In other words, by continuously optimizing network parameters, the accuracy of road defect detection can be continuously improved.
[0064] As an example, when road damage detection is required, the acquired road image is directly input into step S130. Using the trained model, the network modules in steps S130, S140, S150, and S160 are used sequentially to detect damage in the image, predicting the type and location of the damage. Furthermore, after generating the road damage detection results, the accuracy of the detector can be verified based on these results. That is, the network parameters are continuously optimized using the generated detection results and the ground truth labels to improve the detector's accuracy.
[0065] It should be noted that during training, after setting the loss function, gradient backpropagation is performed to update the learnable parameters in S130, S140, S150, and S160. It's important to note that when using a multilayer perceptron (MLP1) in S140, the gradient of parameter W1 is only calculated in the channel attention mechanism, not in the spatial attention mechanism. For example, using binary cross-entropy (BCE) and mean squared error (MSE) as the loss function for location prediction, and mean squared error as the loss function for class prediction, the specific calculation is as follows:
[0066] BCE(y,y′)=-(y×log(y′)+(1-y)×log(1-y′))
[0067]
[0068] Here, y represents the label, which is the type, location, and size of the road defect, and y' represents the prediction result. BCE and MSE are both loss functions used in the fifth step to optimize the network parameters. The detection result includes the type and confidence level of the defect, as well as its size and location. BCE corresponds to optimizing the ability to predict the type and confidence level of the defect. MSE corresponds to optimizing the ability to predict the size and location of the defect.
[0069] Example 2
[0070] Please see Figure 2 This embodiment provides a road defect detection device based on attention mechanism and feature integration, comprising:
[0071] The acquisition module 210 is suitable for acquiring road images.
[0072] The preprocessing module 220 is adapted to preprocess the road images. It includes: a first road image training set generation unit, adapted to generate a first road image training set based on the road images; and a data transformation unit, adapted to perform data transformation on the first road image training set to obtain a second road image training set.
[0073] The feature extraction module 230 is adapted to extract features from the preprocessed road image to generate a first feature map set.
[0074] As an example, the feature extraction module includes: a partitioning unit, adapted to partition the road images into a first road image training set and a road image validation set at a predetermined ratio; and a processing unit, adapted to process the second road image training set using a feature extraction network to obtain a set of feature maps extracted by each sub-network. The feature extraction network includes ResNet and DarkNet.
[0075] The feature focusing processing module 240 is adapted to perform image feature focusing processing on the first feature map set to generate a second feature map set.
[0076] As an example, the second road image training set includes the first road image training set and new road images after data transformation of the first road image training set; the data transformation includes flipping, cropping, brightness and contrast transformation operations.
[0077] As an example, the feature focusing processing module includes a unit for generating a second feature map set, which is adapted to process the first feature map set through a channel attention mechanism and a spatial attention mechanism to generate a second feature map set; wherein the channel attention mechanism uses the SE attention mechanism, and the spatial attention mechanism uses a spatial attention mechanism that shares a multilayer perceptron with the SE.
[0078] The integration module 250 is adapted to integrate the second feature map set to generate a third feature map set.
[0079] As an example, the integration module includes units adapted to integrate the second set of feature maps through pyramid convolution to generate a third set of feature maps.
[0080] The detection module 260 is adapted to generate detection results based on the third feature map set and the predetermined detector.
[0081] As an example, the detection module includes a unit adapted to input the third feature map set into a predetermined detector to generate detection results. The detector is a YOLO detector.
[0082] As an example, the detection module includes units adapted to detect the category and location of road defects based on the third feature map set and a predetermined detector.
[0083] The system further includes, after the detection module: an optimization module adapted to optimize network parameters based on the detection results and the real labels; and a training module adapted to train the predetermined detector based on the optimized network parameters.
[0084] Example 3
[0085] This invention also proposes a storage medium storing a road defect detection program. When executed by a processor, the road defect detection program implements the steps of the road defect detection method described above. Since this storage medium employs all the technical solutions of the above embodiments, it possesses at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be elaborated upon further here.
[0086] Example 4
[0087] Please see Figure 5 The present invention also provides an electronic device, including: a memory and a processor; the memory stores at least one program instruction; the processor loads and executes the at least one program instruction to implement the road defect detection method based on attention mechanism and feature integration provided in Embodiment 1.
[0088] The memory 502 and processor 501 are connected via a bus, which may include any number of interconnecting buses and bridges. The bus connects various circuits of one or more processors 501 and memory 502 together. The bus may also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface between the bus and the transceiver. The transceiver may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 501 is transmitted over a wireless medium via an antenna, which further receives data and transmits it to processor 501.
[0089] Processor 501 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 502 can be used to store data used by processor 501 during operation.
[0090] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the scope of the present invention. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A road disease detection method based on attention mechanism and feature integration, characterized in that, The method comprises: acquiring a road picture; preprocessing the road picture; extracting features from the preprocessed road picture to generate a first feature map set; focusing on picture features of the first feature map set to generate a second feature map set; integrating the second feature map set to generate a third feature map set; generating a detection result based on the third feature map set and a predetermined detector; the focusing on picture features of the first feature map set to generate a second feature map set comprises: processing the first feature map set through a channel attention mechanism and a spatial attention mechanism to generate a second feature map set; wherein the channel attention mechanism uses an SE attention mechanism, and the spatial attention mechanism uses a spatial attention mechanism that shares a multi-layer perceptron with the SE; The feature focusing stage uses a channel attention mechanism and a spatial attention mechanism , resulting in a set of feature maps {F'} that focus on features, which is a second set of feature maps, where F' = F * (C * A + S) ; For the channel attention module, given the input feature where C represents the number of channels, W represents the width of the feature map, and H represents the height of the feature map, the channel feature map is obtained through max pooling and average pooling ; using a multilayer perceptron and a multilayer perceptron to reduce and increase the dimension of to mix semantic information of different dimensions and use an activation function to filter redundant information, obtaining ; add and normalize using a sigmoid function to obtain a channel feature map ; finally, multiply the input feature F with the channel feature map to obtain the feature map U; the formula is as follows: wherein denotes performing an average pooling operation over the channels on the input features; denotes performing a max pooling operation over the channels on the input features, is a parameter of the multilayer perceptron , is a parameter of the multilayer perceptron . For the spatial attention module, given the input feature where C represents the number of channels, W represents the width of the feature map, H represents the height of the feature map, and the multi-layer perceptron shared with the channel attention module reduces the dimension of the feature map to obtain where r represents the compression ratio; the semantic information of different positions is fused through a convolution kernel with a size of 7, and normalized using a sigmoid function to obtain a spatial feature map ; finally, the input feature U is multiplied by the spatial feature map to obtain the feature map F’; wherein represent a multi-layer perceptron shared with the channel attention mechanism parameters of the function, represent a convolution operation with a kernel size of 7; the integrating the second feature map set to generate a third feature map set comprises: integrating the second feature map set through a pyramid convolution to generate a third feature map set; The second feature map set designs a picture feature integration method, denoted as {G}. The picture feature integration method is applied to all samples in the second feature map set to obtain an integrated picture feature map set {F’’}, where F’’= (F’), i={1, 2,…, c}, represents the picture feature integration method, and i represents the corresponding number.
2. The road disease detection method based on attention mechanism and feature integration according to claim 1, wherein, after the generating a detection result based on the third feature map set and a predetermined detector, further comprising: optimizing network parameters based on the detection result and a true label; training the predetermined detector based on the optimized network parameters. 3.The road disease detection method based on attention mechanism and feature integration according to claim 1, wherein, the preprocessing the road picture comprises: generating a first road picture training set based on the road picture; performing data transformation on the first road picture training set to obtain a second road picture training set. 4.The road disease detection method based on attention mechanism and feature integration according to claim 3, wherein, the generating a first road picture training set based on the road picture comprises: dividing the road picture into a first road picture training set and a road picture validation set at a predetermined ratio; the second road picture training set comprises the first road picture training set and new road pictures obtained by performing data transformation on the first road picture training set; the data transformation comprises flipping, cropping, brightness and contrast transformation operations. 5.The road disease detection method based on attention mechanism and feature integration according to claim 1 or 3, wherein, the extracting features from the preprocessed road picture to generate a first feature map set comprises: processing the second road picture training set using a feature extraction network to obtain a feature map set extracted by each sub-network. 6.The road disease detection method based on attention mechanism and feature integration according to claim 1, wherein, the generating a detection result based on the third feature map set and a predetermined detector comprises: detecting the category and location of road diseases based on the third feature map set and the predetermined detector.
7. A road disease detection device based on attention mechanism and feature integration, characterized in that, The device comprises: an acquisition module adapted to acquire a road picture; a preprocessing module adapted to preprocess the road picture; a feature extraction module adapted to extract features from the preprocessed road picture to generate a first feature map set; a feature focusing processing module adapted to focus on picture features of the first feature map set to generate a second feature map set, comprising: processing the first feature map set through a channel attention mechanism and a spatial attention mechanism to generate a second feature map set; wherein the channel attention mechanism uses an SE attention mechanism, and the spatial attention mechanism uses a spatial attention mechanism that shares a multi-layer perceptron with the SE; The feature focusing stage uses a channel attention mechanism and a spatial attention mechanism , resulting in a set of feature maps {F'} that focus on features, being a second set of feature maps, where F' = F * C * S ; For the channel attention module, given the input feature where C represents the number of channels, W represents the width of the feature map, and H represents the height of the feature map, the channel feature map is obtained through max pooling and average pooling ; using a multilayer perceptron and a multilayer perceptron to reduce and increase the dimension of to mix semantic information of different dimensions and use an activation function to filter redundant information, obtaining ; add and normalize using a sigmoid function to obtain a channel feature map ; finally, the input feature F is multiplied by the channel feature map to obtain the feature map U; the formula is as follows: wherein represents performing an average pooling operation on the input features across channels; represents performing a max pooling operation on the input features across channels, is a parameter of the multilayer perceptron , is a parameter of the multilayer perceptron . For the spatial attention module, given the input feature where C represents the number of channels, W represents the width of the feature map, H represents the height of the feature map, and the multi-layer perceptron shared with the channel attention module reduces the dimension of the feature map to obtain where r represents the compression ratio; the semantic information at different positions is fused by processing with a convolution kernel of size 7, and normalized using a sigmoid function to obtain a spatial feature map ; finally, the input feature U is multiplied by the spatial feature map to obtain the feature map F’; wherein representing a multi-layer perceptron shared with the channel attention mechanism parameters of, representing a convolution operation with a kernel size of 7; an integration module adapted to integrate the second feature map set to generate a third feature map set, comprising: integrating the second feature map set through a pyramid convolution to generate a third feature map set; The second feature map set designs a picture feature integration method, denoted as {G}. The picture feature integration method is applied to all samples in the second feature map set to obtain an integrated picture feature map set {F’’}, where F’’= (F’), i={1, 2,…, c}, represents the picture feature integration method, and i represents the corresponding number. A detection module adapted to generate a detection result based on the third feature map set and a predetermined detector.
8. A computer-readable storage medium having stored therein one or more instructions, wherein: The processor of the device for risk analysis within the one or more instructions implements the road disease detection method based on the attention mechanism and feature integration according to any one of claims 1 to 6 when executed.
9. An electronic device, comprising: Comprise: A memory and a processor; At least one program instruction is stored in the memory; The processor implements the road disease detection method based on the attention mechanism and feature integration according to any one of claims 1 to 6 by loading and executing the at least one program instruction.
Citation Information
Patent Citations
Road disease detection method based on candidate area network and machine vision
CN112200143A
Small target detection method based on attention mechanism
CN114202672A
Road disease detection method and system based on convolutional neural network
CN114882474A