Road disease detection method and device and storage medium

By designing a convolution kernel corresponding to the length-to-width ratio of road defects in the convolutional neural network and setting feature fusion weights based on quantitative statistical characteristics, the problem of information loss in road defect detection is solved and the detection accuracy is improved.

CN120635845APending Publication Date: 2025-09-12ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510549498.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies have difficulty accurately identifying road defects, especially inaccurate detection of longitudinal and transverse cracks, which leads to information loss and misidentification, affecting subsequent maintenance decisions.

Method used

By designing multiple convolution kernels in the convolutional neural network to correspond to road defects with different length-to-width ratios, feature fusion weights are set based on quantitative statistical characteristics, and weighted fusion feature extraction and recognition are performed, combined with the output results of the defect detection network.

Benefits of technology

It improves the information perception of road defects with different length-to-width ratios, avoids information loss during the convolution process, and enhances the accuracy of road defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635845A_ABST
    Figure CN120635845A_ABST
Patent Text Reader

Abstract

The invention discloses a road disease detection method and device and a storage medium. The road disease detection method comprises the following steps: extracting quantity statistical characteristics corresponding to road diseases with different length-width ratios in a plurality of marked images; respectively inputting the to-be-detected road image into a plurality of convolution kernels of the convolutional neural network for convolution processing to obtain image features obtained by respective convolution of each convolution kernel; respectively setting a feature fusion weight corresponding to each image feature based on the quantity statistical features; according to the feature fusion weight, performing weighted fusion on each image feature to obtain a weighted fusion feature; and inputting the weighted fusion features into a disease detection network to obtain a road disease detection result, so that the information perceptibility of the road diseases with different length-width ratios can be improved, and the quantity information of the road diseases with different length-width ratios in a real scene is considered during feature fusion, so that the finally obtained weighted fusion features are more accurate, and the accuracy of the road disease detection is improved. And the accuracy of subsequent road recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a road hazard detection method, device, and storage medium. Background Art

[0002] Highways are an important part of the transportation system, but due to factors such as overloading and extreme weather, road surface defects such as cracks, potholes, and fissures may occur during operation. Severe road defects can lead to traffic accidents. Therefore, it is necessary to identify existing road defects on highways to facilitate maintenance and provide warnings.

[0003] How to accurately identify road damage to provide a basis for decision-making is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] In order to solve the above technical problems, the present application at least provides a road disease detection method, equipment and storage medium.

[0005] In a first aspect, the present application provides a road disease detection method, comprising: obtaining a plurality of labeled images containing road diseases, wherein the labeled images are labeled with the aspect ratios of the road diseases in the labeled images, and extracting quantitative statistical features corresponding to road diseases with different aspect ratios in the plurality of labeled images; inputting the road images to be detected into a plurality of convolution kernels of a convolutional neural network for convolution processing to obtain image features convolved by each convolution kernel; wherein the aspect ratios of different convolution kernels correspond one-to-one to different aspect ratios of road diseases; setting feature fusion weights corresponding to the image features convolved by each convolution kernel based on the quantitative statistical features corresponding to road diseases with different aspect ratios; performing weighted fusion on the image features convolved by each convolution kernel according to the feature fusion weights to obtain weighted fusion features; and inputting the weighted fusion features into a disease detection network to obtain a road disease detection result output by the disease detection network.

[0006] In one embodiment, the quantitative statistical features include the total number of road defects with different aspect ratios in multiple labeled images; based on the quantitative statistical features corresponding to the road defects with different aspect ratios, feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set respectively, including: taking any aspect ratio as the ratio to be calculated; based on the total number of road defects containing the ratio to be calculated, feature fusion weights corresponding to the image features obtained by convolving the convolution kernel with the ratio to be calculated are set; wherein the total number of road defects is positively correlated with the feature fusion weight.

[0007] In one embodiment, the marked image is also marked with correct road disease identification or incorrect road disease identification, and the quantitative statistical features contain the number of correct identifications and the number of incorrect identifications corresponding to road diseases with different aspect ratios; based on the quantitative statistical features corresponding to road diseases with different aspect ratios, the feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set respectively, including: taking any aspect ratio as the ratio to be calculated; based on the number of correct identifications and the number of incorrect identifications corresponding to the road diseases with the ratio to be calculated, the recognition accuracy corresponding to the road disease with the ratio to be calculated is calculated; based on the recognition accuracy corresponding to the road disease with the ratio to be calculated, the feature fusion weights corresponding to the image features obtained by convolving the convolution kernel with the ratio to be calculated are set; wherein the recognition accuracy and the feature fusion weight are inversely correlated.

[0008] In one embodiment, the marked image is also marked with a road surface type, and the quantitative statistical features contain the number of road diseases with different length-width ratios corresponding to each road surface type; based on the quantitative statistical features corresponding to the road diseases with different length-width ratios, the feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set respectively, including: taking any length-width ratio as the ratio to be calculated; and identifying the road surface type corresponding to the road image to be detected to obtain the target road surface type; based on the number of road diseases with the ratio to be calculated corresponding to the target road surface type, calculating the correlation between the target road surface type and the road diseases with the ratio to be calculated; wherein, the number of road diseases with the ratio to be calculated and the correlation are positively correlated; based on the correlation between the target road surface type and the road diseases with the ratio to be calculated, setting the feature fusion weights corresponding to the image features obtained by convolving the convolution kernel with the ratio to be calculated; wherein, the correlation and the feature fusion weight are positively correlated.

[0009] In one embodiment, the weighted fusion features are input into a disease detection network to obtain a road disease detection result output by the disease detection network, including: inputting the weighted fusion features into a feature channel fusion network, the feature channel fusion network is used to: split the weighted fusion features in the feature channel dimension to obtain multiple first split features; for some or all of the multiple first split features, fuse one or more other first split features, combine the fused first split features and / or unfused first split features to obtain multiple processed features; splice the multiple processed features to obtain spliced ​​features, and perform convolution processing on the spliced ​​features to obtain target features output by the feature channel fusion network; input the target features into the disease detection network to obtain a road disease detection result output by the disease detection network.

[0010] In one embodiment, multiple processed features are spliced ​​to obtain spliced ​​features, including: splitting the multiple processed features in the feature channel dimension to obtain multiple second split features; inputting the multiple second split features into a single-head self-attention network, performing self-attention calculation on the multiple second split features based on the single-head self-attention network to obtain self-attention calculation results; performing weighted dot multiplication calculation on the multiple second split features based on the self-attention calculation results to obtain spliced ​​features output by the single-head self-attention network.

[0011] In one embodiment, a plurality of feature channel fusion networks of different scales are included; target features are input into a disease detection network to obtain a road disease detection result output by the disease detection network, including: obtaining target features respectively output by feature channel fusion networks of different scales; and inputting the target features respectively output by feature channel fusion networks of different scales into the disease detection network to obtain a road disease detection result output by the disease detection network.

[0012] In one embodiment, the defect detection network includes a defect type detection network and a defect location detection network; inputting target features into the defect detection network to obtain a road defect detection result output by the defect detection network includes: inputting target features into the defect type detection network and the defect location detection network respectively to obtain the defect type output by the defect type detection network and the defect location output by the defect location detection network; combining the defect type and the defect location to obtain a road defect detection result corresponding to the road image to be detected.

[0013] According to a second aspect of the present application, there is provided a road disease detection device, which includes: a quantity statistics module for obtaining multiple labeled images containing road diseases, where the labeled images are marked with the aspect ratios of the road diseases in the labeled images, and extracting the quantity statistical features corresponding to the road diseases with different aspect ratios in the multiple labeled images; a convolution processing module for inputting the road images to be detected into multiple convolution kernels of a convolutional neural network for convolution processing, and obtaining the image features obtained by convolution of each convolution kernel; wherein the aspect ratios of different convolution kernels correspond one-to-one to the different aspect ratios of the road diseases; a weight determination module for setting the feature fusion weights corresponding to the image features obtained by convolution of each convolution kernel based on the quantity statistical features corresponding to the road diseases with different aspect ratios; a weighted fusion module for performing weighted fusion on the image features obtained by convolution of each convolution kernel according to the feature fusion weights, and obtaining weighted fusion features; and a disease detection module for inputting the weighted fusion features into the disease detection network, and obtaining the road disease detection results output by the disease detection network.

[0014] A third aspect of the present application provides an electronic device, comprising a memory and a processor, wherein the processor is configured to execute program instructions stored in the memory to implement the above-mentioned road hazard detection method.

[0015] A fourth aspect of the present application provides a computer-readable storage medium having program instructions stored thereon, which implement the above-mentioned road defect detection method when the program instructions are executed by a processor.

[0016] The above scheme extracts the quantitative statistical features corresponding to road defects with different aspect ratios from multiple labeled images; inputs the road images to be detected into multiple convolution kernels of the convolutional neural network for convolution processing to obtain image features convolved by each convolution kernel; based on the quantitative statistical features corresponding to road defects with different aspect ratios, sets the feature fusion weights corresponding to the image features convolved by each convolution kernel; performs weighted fusion on the image features convolved by each convolution kernel according to the feature fusion weights to obtain weighted fusion features; inputs the weighted fusion features into the defect detection network to obtain the road defect detection results output by the defect detection network, which can improve the information perception of road defects with different aspect ratios and avoid the situation where the short side information is submerged during the convolution process. In addition, the quantitative information of road defects with different aspect ratios in real scenes is taken into account during feature fusion, so that the final weighted fusion features are more accurate, thereby improving the accuracy of subsequent road recognition.

[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0019] Figure 1 is a schematic diagram of a solution implementation environment shown in an exemplary embodiment of the present application;

[0020] Figure 2 is a flow chart of a road damage detection method shown in an exemplary embodiment of the present application;

[0021] Figure 3 is a schematic diagram of a convolutional neural network shown in an exemplary embodiment of the present application;

[0022] Figure 4 is a schematic diagram of a feature channel fusion network shown in an exemplary embodiment of the present application;

[0023] Figure 5 is a schematic diagram of a single-head self-attention module shown in an exemplary embodiment of the present application;

[0024] Figure 6 is a schematic diagram of a disease detection network shown in an exemplary embodiment of the present application;

[0025] Figure 7 is a schematic diagram of a road damage detection model shown in an exemplary embodiment of the present application;

[0026] Figure 8 is a block diagram of a road damage detection device shown in an exemplary embodiment of the present application;

[0027] Figure 9 is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present application;

[0028] Figure 10 It is a schematic diagram of the structure of a computer-readable storage medium shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0029] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0030] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0031] The term "and / or" in this article is merely information describing the association of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0032] The applicant found that due to the large variations in the shape and scale of road defects, the target frame areas of the same type of road defects may differ by dozens of times. In particular, the target frames of longitudinal cracks and transverse cracks are mainly line-shaped with a large length-to-width ratio, which makes feature extraction difficult and information loss occurs, resulting in inaccurate road defect detection.

[0033] Specifically, longitudinal and transverse cracks are primarily thin lines, meaning that the pixel count on one side is very small. This can easily cause information on the shorter side to be lost during the convolution process. Furthermore, crack sizes vary greatly, ranging from over a meter long to just over ten centimeters. Due to this large aspect ratio, conventional convolutional models can easily identify only a portion of the crack defect, failing to fully identify the long side. Alternatively, they can misidentify a long crack defect as several smaller segments. This can lead to statistical errors, potentially affecting subsequent repair decisions or causing false alarms.

[0034] In order to solve the above problems, the present application at least provides a road disease detection method, equipment and storage medium.

[0035] The road damage detection method provided in the embodiments of the present application is described below.

[0036] Please refer to Figure 1 , Figure 1 FIG1 is a schematic diagram of an exemplary embodiment of the present application showing a solution implementation environment, wherein the solution implementation environment may include a terminal 110 and a server 120, and the terminal 110 and the server 120 are in communication connection with each other.

[0037] The number of terminals 110 can be one or more, and the terminals 110 include but are not limited to image acquisition devices, smart phones, smart watches, etc., which store road images. The road images can be collected by the terminal 110 or obtained by the terminal 110 from other databases. This application does not limit this.

[0038] Server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0039] In one example, the server 120 may perform road damage detection processing on the road image obtained from the terminal 110 to obtain a road damage detection result. Of course, the server 120 may store the road damage detection result locally, transmit it back to the terminal 110, or transmit it to other terminals.

[0040] In one example, terminal 110 is installed with a client running a target application. This target application may be an application that provides road damage detection functionality. This target application performs road damage detection processing on captured road images to obtain road damage detection results. Server 120 may be the backend server for this target application, configured to provide backend services for the client of this target application.

[0041] In the road defect detection method provided in the embodiment of the present application, the execution entity of each step can be the terminal 110, such as the client of the target application installed and running in the terminal 110, or the server 120, or the terminal 110 and the server 120 can interact and cooperate to execute the method, that is, part of the steps of the method are executed by the terminal 110 and other steps are executed by the server 120.

[0042] See also Figure 2 , Figure 2 This is a flow chart of a road damage detection method shown in an exemplary embodiment of the present application. The road damage detection method can be applied to Figure 1 It should be understood that the method can also be applied to other exemplary implementation environments and be specifically executed by devices in other implementation environments, and this embodiment does not limit the implementation environment to which the method is applicable.

[0043] like Figure 2 As shown, the road damage detection method includes at least steps S210 to S250, which are described in detail as follows:

[0044] Step S210: Acquire multiple labeled images containing road defects, where the labeled images are labeled with the aspect ratios of the road defects in the labeled images, and extract quantitative statistical features corresponding to road defects with different aspect ratios in the multiple labeled images.

[0045] The marked image can be a road image that has been manually annotated; the marked image can also be an image with road disease detection results. For example, the road disease detection results contain a target frame for marking the location of the road disease. The length-to-width ratio of the road disease is obtained based on the target frame, and the road disease detection results corresponding to each road image within a preset historical time period are recorded. Each road image within the preset historical time period is used as a marked image, and the length-to-width ratio of the road disease in the marked image is obtained based on the road disease detection results.

[0046] The length-to-width ratio of a road defect refers to the ratio between the length and width of the road defect.

[0047] The length-to-width ratio of the road damage in each marked image is calculated, and road damage with different length-to-width ratios is obtained.

[0048] For example, to facilitate the design and calculation of subsequent convolution kernels, the image is divided into a preset number of different aspect ratios, and the aspect ratios are integers. For example, the different aspect ratios of road damage obtained by division include (7:1), (7:2), (7:3), etc.

[0049] Then, the quantitative statistical features corresponding to road defects with different aspect ratios in multiple labeled images are extracted.

[0050] Among them, the quantitative statistical features are used to describe the quantitative attributes of a data set composed of road defects contained in multiple labeled images. They are quantitative indicators obtained by statistical calculation of the data set and are used to describe the key characteristics or distribution patterns of the data set.

[0051] For example, quantitative statistical features include but are not limited to the total number of road defects with different length-to-width ratios in each marked image, and / or the number of correctly identified and incorrectly identified road defects with different length-to-width ratios, and / or the number of road defects with different length-to-width ratios corresponding to each road surface type, etc.

[0052] The type of quantitative statistical features to be extracted can be flexibly selected according to the actual application situation, and this application does not limit this.

[0053] Step S220: Input the road image to be detected into multiple convolution kernels of the convolutional neural network for convolution processing to obtain image features obtained by convolution of each convolution kernel; wherein the length-to-width ratios of different convolution kernels correspond one-to-one to different length-to-width ratios of road defects.

[0054] The road image to be detected refers to a road image for which road damage information needs to be detected.

[0055] For example, a camera is deployed relative to a target road, and the camera periodically captures road images of the target road and sends the road images as road images to be detected to a server. The server performs road disease detection on the received road images to be detected.

[0056] The convolutional neural network contains multiple convolution kernels, and the length-to-width ratios of different convolution kernels correspond to the different length-to-width ratios of road defects.

[0057] Specifically, a convolution kernel is typically a four-dimensional tensor with a shape of (N, C, K_H, K_W), where K_H is the height of the kernel, K_W is the width of the kernel, C is the number of input channels, and N is the number of output channels. The kernel's aspect ratio is the ratio of K_H to K_W.

[0058] For example, the different length-to-width ratios of road defects include (7:1), (7:2), and (7:3). Therefore, the length-to-width ratios of the designed convolution kernels include (1, 7), (2, 7), (3, 7), (7, 1), (7, 2), and (7, 3), which are used to effectively extract features of transverse and longitudinal cracks. Of course, to ensure the recognition of various types of road defects, convolution kernels of other sizes can also be set. For example, the convolutional neural network also includes (3, 3) and (7, 7), which can be used to effectively extract features of other road defects or road defects with characteristic shapes.

[0059] The road image to be detected is input into multiple convolution kernels of the convolutional neural network for convolution processing, and the image features obtained by each convolution kernel are obtained. That is, each convolution kernel performs convolution processing on the road image to be detected, and the image features obtained by each convolution kernel are obtained. Multi-dimensional feature extraction of the road image to be detected is performed to avoid the loss of image information.

[0060] For example, if the convolutional neural network contains the eight convolution kernels of different sizes as mentioned above, the convolution kernel is divided into eight parts according to the channel dimension C corresponding to the road image to be detected. The first seven parts are C divisible by 8 (i.e., C|8), and the last part is C divisible by 8 plus the remainder of C divisible by 8 (i.e., C|8+C%8). In this way, each channel corresponding to the road image to be detected is assigned to each convolution kernel for convolution processing, and the image features output by each convolution kernel are obtained.

[0061] Optionally, before inputting the road image to be detected into the convolutional neural network, the road image to be detected can also be subjected to image preprocessing, which includes but is not limited to image cropping, and / or image resizing, and / or image enhancement, and / or image denoising, etc., which is not limited in this application.

[0062] Step S230: Based on the quantitative statistical features corresponding to road defects with different aspect ratios, the feature fusion weights corresponding to the image features obtained by convolution of each convolution kernel are set.

[0063] According to the quantitative statistical characteristics corresponding to road diseases with different length-width ratios, the feature fusion weights corresponding to the image features obtained by convolution of each convolution kernel are determined respectively.

[0064] Specifically, road defects with different length-to-width ratios correspond to convolution kernels with different length-to-width ratios, and the feature fusion weight of the convolution kernel corresponding to the road defect with any length-to-width ratio is set according to the quantitative statistical characteristics corresponding to the road defect with any length-to-width ratio.

[0065] For example, a road defect with an aspect ratio of (7:1) corresponds to convolution kernels with aspect ratios of (1, 7) and (7, 1). Based on the quantitative statistical characteristics corresponding to the road defect with an aspect ratio of (7:1), the feature fusion weights of the convolution kernels with aspect ratios of (1, 7) and (7, 1) are set.

[0066] For example, the feature fusion weight of each convolution kernel is set according to the total number of road defects with different length-width ratios in each marked image, and / or the number of correctly identified and incorrectly identified road defects with different length-width ratios, and / or the number of road defects with different length-width ratios corresponding to each road surface type.

[0067] Step S240: performing weighted fusion on the image features obtained by convolution of each convolution kernel according to the feature fusion weight to obtain a weighted fusion feature.

[0068] According to the feature fusion weights corresponding to the image features obtained by convolution of each convolution kernel, the image features obtained by convolution of each convolution kernel are weighted fused, and the weighted fusion results are used as weighted fusion features.

[0069] For an example, see Figure 3 , Figure 3 is a schematic diagram of a convolutional neural network shown in an exemplary embodiment of the present application, such as Figure 3 As shown, the convolutional neural network includes a multi-ratio convolution module, a regular convolution module, and a feature fusion module. The aspect ratios of the convolution kernels in the multi-ratio convolution module correspond to the different aspect ratios of road defects. The regular convolution module can be a convolutional network that performs image convolution using existing techniques, such as the convolutional network in the YOLO (You Only Look Once) model. The input can be either the original road image to be inspected or a pre-processed image, and convolution processing is performed on the input using the multi-ratio convolution module and the regular convolution module, respectively. Then, based on the feature fusion weights obtained above, the feature fusion module performs weighted fusion on the image features output by the multi-aspect ratio convolution module and the conventional convolution module to obtain weighted fusion features. Here, the image features output by the multi-aspect ratio convolution module can be first weighted fused, and then the weighted fusion features and the image features output by the conventional convolution module are simply spliced ​​according to the channel dimension based on the splicing network (Concat) to obtain weighted fusion features. Alternatively, the conventional convolution module can be pre-set with feature fusion weights, and all the above image features are weighted fused based on the splicing network to obtain weighted fusion features. Finally, the weighted fusion features are passed through a conventional convolution module to obtain the final weighted fusion features.

[0070] Among them, each feature output by the multi-aspect ratio convolution module and the conventional convolution module can be padded to ensure that the feature shapes of the convolution outputs of different channels are the same by adding features with a value of 0, thereby facilitating feature fusion.

[0071] Of course, you can also delete the conventional convolution module and only use the multi-length-width ratio convolution module. This application does not limit this.

[0072] The above-mentioned convolutional neural network can improve the information perception of road diseases with different length and width ratios by setting convolution kernels that correspond one-to-one to the different length and width ratios of road diseases. In particular, for crack diseases, it can avoid the situation where the short side information is submerged during the convolution process. In addition, the quantitative information of road diseases with different length and width ratios in real scenes is taken into account during feature fusion, making the final weighted fusion feature more accurate and improving the accuracy of subsequent road recognition.

[0073] Step S250: input the weighted fusion features into the disease detection network to obtain the road disease detection results output by the disease detection network.

[0074] The disease detection network is used to output road disease detection results based on the input features, wherein the road disease detection results may include information such as whether there is a road disease, and / or the type of road disease, and / or the location of the road disease.

[0075] Next, some embodiments of the present application are described in detail with examples.

[0076] Based on any of the following embodiments, the feature fusion weight corresponding to the image feature obtained by convolution of each convolution kernel can be set.

[0077] Embodiment 1: The quantitative statistical features contain the total number of road defects with different aspect ratios in multiple labeled images; based on the quantitative statistical features corresponding to the road defects with different aspect ratios, the feature fusion weights corresponding to the image features obtained by convolution of each convolution kernel are set respectively, including: taking any aspect ratio as the ratio to be calculated; based on the total number of road defects containing the ratio to be calculated, the feature fusion weights corresponding to the image features obtained by convolution of the convolution kernel with the ratio to be calculated are set; wherein the total number of road defects is positively correlated with the feature fusion weight.

[0078] The total number of road defects with different length-width ratios in multiple labeled images is counted respectively, and the total number of road defects with different length-width ratios is used as a quantitative statistical feature.

[0079] According to the total number of road defects with different length-width ratios, the feature fusion weights corresponding to the image features obtained by convolution of each convolution kernel are set respectively.

[0080] Specifically, any aspect ratio is used as the ratio to be calculated. Based on the total number of road defects that contain the ratio, the feature fusion weight corresponding to the image features obtained by convolving the convolution kernel with the ratio to be calculated is set. The larger the total number of road defects in the ratio to be calculated, the greater the feature fusion weight corresponding to the convolution kernel for that ratio. Conversely, the smaller the total number of road defects in the ratio to be calculated, the smaller the feature fusion weight corresponding to the convolution kernel for that ratio.

[0081] For example, the total number of road defects with the first aspect ratio is N1, and the total number of road defects with the second aspect ratio is N2. It can be calculated that the feature fusion weight corresponding to the convolution kernel of the first aspect ratio is N1 / (N1+N2), and the feature fusion weight corresponding to the convolution kernel of the second aspect ratio is N2 / (N1+N2).

[0082] Of course, in addition to calculating the proportion relationship between each total number to obtain the feature fusion weight as described above, other methods can also be used to calculate the feature fusion weight, such as calculating the difference between the total number and a preset threshold (the preset threshold can be the mean between each total number, or a value pre-set based on experience), normalizing the difference, and obtaining the feature fusion weight. This application does not limit the specific method of calculating the feature fusion weight based on the total number.

[0083] Through the above embodiment, the feature fusion weights corresponding to each convolution kernel can be adjusted according to the total number of road defects with different aspect ratios, so that more attention can be paid to the image features extracted by the convolution kernel corresponding to road defects that are more common in actual scenes.

[0084] Embodiment 2: The marked image is also marked with correct road disease identification or incorrect road disease identification, and the quantitative statistical features contain the number of correct identifications and the number of incorrect identifications corresponding to road diseases with different aspect ratios; based on the quantitative statistical features corresponding to road diseases with different aspect ratios, the feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set respectively, including: taking any aspect ratio as the ratio to be calculated; based on the number of correct identifications and the number of incorrect identifications corresponding to the road diseases with the ratio to be calculated, the recognition accuracy corresponding to the road disease with the ratio to be calculated is calculated; based on the recognition accuracy corresponding to the road disease with the ratio to be calculated, the feature fusion weights corresponding to the image features obtained by convolving the convolution kernel with the ratio to be calculated are set; wherein the recognition accuracy and the feature fusion weight are inversely correlated.

[0085] In addition to marking the aspect ratio of the road defect in the marked image, the marked image also marks whether the road defect in the marked image is correctly identified or incorrectly identified.

[0086] For example, after obtaining the road damage detection results of the marked image, the road damage detection results can be fed back to other terminals and displayed on other terminals, and feedback information from users on whether the road damage detection results are correct or incorrect can be obtained, and the marked image can be marked accordingly based on the feedback information.

[0087] For another example, multiple different road damage detection models are pre-trained. The road damage detection models can predict the road damage contained in the input image. The marked image is input into multiple different road damage detection models to obtain the road damage detection results output by each road damage detection model. If the road damage detection results are consistent, it is judged that the road damage identification is correct. If the road damage detection results are inconsistent, it is judged that the road damage identification is incorrect.

[0088] The number of correct and incorrect identifications corresponding to road defects with different length and width ratios are counted, and the feature fusion weights corresponding to the image features obtained by convolution of convolution kernels with different length and width ratios are set according to the number of correct and incorrect identifications corresponding to road defects with different length and width ratios.

[0089] Specifically, any length-to-width ratio is taken as the ratio to be calculated; based on the number of correctly identified and incorrectly identified road defects corresponding to the ratio to be calculated, the recognition accuracy rate of the road defects corresponding to the ratio to be calculated is calculated. For example, if the number of correctly identified road defects corresponding to the first length-to-width ratio is n11 and the number of incorrectly identified road defects is n12, the recognition accuracy rate of the road defects corresponding to the first length-to-width ratio is n11 / (n11+n12).

[0090] Then, based on the recognition accuracy of the road damage proportion to be calculated, the feature fusion weight corresponding to the image features obtained by convolving the convolution kernel for that proportion is set. The higher the recognition accuracy of the road damage proportion to be calculated, the smaller the feature fusion weight corresponding to the convolution kernel for that proportion. Conversely, the lower the recognition accuracy of the road damage proportion to be calculated, the larger the feature fusion weight corresponding to the convolution kernel for that proportion.

[0091] Through the above embodiment, the feature fusion weights corresponding to each convolution kernel can be adjusted according to the recognition accuracy corresponding to road defects with different aspect ratios, so that more attention can be paid to the image features extracted by the convolution kernel corresponding to road defects that are easy to be misidentified in actual scenes.

[0092] Embodiment 3: The marked image is also marked with a road surface type, and the quantitative statistical features contain the number of road diseases with different length-width ratios corresponding to each road surface type; based on the quantitative statistical features corresponding to the road diseases with different length-width ratios, the feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set respectively, including: taking any length-width ratio as the ratio to be calculated; and identifying the road surface type corresponding to the road image to be detected to obtain the target road surface type; based on the number of road diseases with the ratio to be calculated corresponding to the target road surface type, calculating the correlation between the target road surface type and the road diseases with the ratio to be calculated; wherein, the number of road diseases with the ratio to be calculated and the correlation are positively correlated; based on the correlation between the target road surface type and the road diseases with the ratio to be calculated, setting the feature fusion weights corresponding to the image features obtained by convolving the convolution kernel with the ratio to be calculated; wherein, the correlation and the feature fusion weight are positively correlated.

[0093] In addition to marking the aspect ratio of the road damage in the marked image, the marked image is also marked with the road surface type of the road in the marked image.

[0094] For example, the pavement types include but are not limited to asphalt pavement, cement concrete pavement, gravel pavement, block pavement, etc.

[0095] The road surface type can be manually marked, or a pre-trained road surface classification model can be used. The road surface classification model can predict the road surface type in the input image. The marked image is input into the road surface classification model to obtain the road surface type output by the road surface classification model. This application does not limit the method for obtaining the road surface type of the marked image.

[0096] The number of road defects with different length-width ratios corresponding to each pavement type is counted. The correlation between each pavement type and road defects with different length-width ratios is calculated based on the statistical results. Based on the correlation, the feature fusion weights corresponding to the image features obtained by convolution of convolution kernels with different length-width ratios are set.

[0097] For example, if the road type corresponding to the road image to be detected is asphalt pavement, statistics are performed on all marked images, and the number of road diseases in the first aspect ratio of the asphalt pavement is d1, and the number of road diseases in the second aspect ratio of the asphalt pavement is d2. Then, based on d1, the correlation between the asphalt pavement and the road diseases in the first aspect ratio can be calculated as d1 / (d1+d2); based on d2, the correlation between the asphalt pavement and the road diseases in the second aspect ratio can be calculated as d2 / (d1+d2).

[0098] Then, based on the correlation degree corresponding to the road defects of the first aspect ratio, the feature fusion weight corresponding to the image features obtained by convolution with the convolution kernel of the first aspect ratio is set; based on the correlation degree corresponding to the road defects of the second aspect ratio, the feature fusion weight corresponding to the image features obtained by convolution with the convolution kernel of the second aspect ratio is set. Specifically, the greater the correlation degree, the greater the feature fusion weight, and the smaller the correlation degree, the smaller the feature fusion weight.

[0099] Through the above embodiment, the feature fusion weights corresponding to each convolution kernel can be adjusted according to the correlation between the road surface type corresponding to the road image to be detected and the road damage ratio to be calculated, so that more attention can be paid to the image features extracted by the convolution kernel corresponding to the road damage with a stronger correlation with the target road surface type.

[0100] In some implementations, the feature fusion weights may be set in combination with the above-mentioned embodiments. For example, according to Embodiment 1, Embodiment 2, and Embodiment 3, three feature fusion weights corresponding to the convolution kernel of the to-be-calculated ratio are calculated, and the mean of the three feature fusion weights is calculated, and the mean calculation result is used as the final feature fusion weight of the convolution kernel of the to-be-calculated ratio.

[0101] Optionally, when calculating feature fusion weights according to the above embodiments, an initial fusion weight may be set for each convolution kernel, a weight adjustment value may be calculated according to one or more of the above embodiments, and then the initial fusion weight and the weight adjustment value may be summed, with the sum being used as the feature fusion weight corresponding to each convolution kernel. The initial fusion weight may be pre-set based on experience.

[0102] In addition, in order to ensure that the feature fusion weights corresponding to each convolution kernel are within a reasonable range, the weight value selection range corresponding to each convolution kernel can be set to detect whether the feature fusion weight is within the weight value selection range. If the feature fusion weight is greater than the maximum value of the weight value selection range, the maximum value of the weight value selection range is taken as the final feature fusion weight; if the feature fusion weight is less than the minimum value of the weight value selection range, the minimum value of the weight value selection range is taken as the final feature fusion weight.

[0103] Through the above embodiment, more accurate weighted fusion features are extracted, and the weighted fusion features are input into the disease detection network to obtain the road disease detection results output by the disease detection network.

[0104] In some embodiments, in step S250, the weighted fusion features are input into the disease detection network to obtain the road disease detection results output by the disease detection network, including:

[0105] Step S251: input the weighted fusion features into the feature channel fusion network, which is used to: split the weighted fusion features in the feature channel dimension to obtain multiple first split features; for some or all of the split features in the multiple first split features, fuse one or more other first split features, combine the fused first split features and / or the unfused first split features to obtain multiple processed features; splice the multiple processed features to obtain spliced ​​features, and perform convolution processing on the spliced ​​features to obtain target features output by the feature channel fusion network.

[0106] For an example, see Figure 4 , Figure 4 is a schematic diagram of a feature channel fusion network shown in an exemplary embodiment of the present application. Figure 4 As shown in the figure, the direction of the arrow represents the data flow. Assuming that the shape of the input feature X is Cx*Wx*Hx, it is divided into three equal parts according to the feature channel dimension to obtain multiple first split features X1, X2, and X3. The shape of the first split features is (Cx / 3)*Wx*Hx.

[0107] Then, for some or all of the multiple first split features, one or more other first split features are fused to obtain multiple processed features.

[0108] For example, see Figure 4 For X2, X2 and X1 are added (Add). Specifically, X2 and X1 are added element by element to achieve feature fusion. The added features are input into a two-dimensional convolution (Conv2d) for feature extraction (without changing the feature shape). The output of the two-dimensional convolution is used as the processed features corresponding to X2. Similarly, for X3, the processed features corresponding to X3 and X2 are added, and then input into a two-dimensional convolution for feature extraction to obtain the processed features corresponding to X3. X1 is not fused with other first-split features, and X1 is directly used as the processed feature corresponding to X1.

[0109] Then, multiple processed features are spliced ​​together to obtain spliced ​​features.

[0110] Exemplarily, multiple processed features are spliced ​​to obtain spliced ​​features, including: splitting the multiple processed features in the feature channel dimension to obtain multiple second split features; inputting the multiple second split features into a single-head self-attention network, performing self-attention calculation on the multiple second split features based on the single-head self-attention network to obtain self-attention calculation results; performing weighted point multiplication calculation on the multiple second split features based on the self-attention calculation results to obtain spliced ​​features output by the single-head self-attention network.

[0111] Continue to see Figure 4Each processed feature is input into the Channel Split module, which splits all input features by channel to obtain multiple second-split features. These second-split features are then input into the Single Head Self Attention (SHSA) module. SHSA performs self-attention calculations on all inputs, dynamically focusing on the input feature vectors. This allows the model to establish global dependencies between different positions, thereby capturing long- and short-range dependencies in the sequence. The Single Head Self Attention module does not change the feature shape and concatenates all channel features. The output of the concatenated feature has a shape of Cx*Wx*Hx.

[0112] Specifically, see Figure 5 , Figure 5 This is a schematic diagram of a single-head self-attention module shown in an exemplary embodiment of the present application. Figure 5 As shown in the figure, three linear networks are used to map each input channel feature (i.e., the second split feature) three times, resulting in three new features: Q, K, and V, which represent query, key, and value, respectively. Matrix dot product (Matrix Dot Product) is then performed on Q and K, and normalized using a softmax function to obtain weights. This weight is used as the result of the self-attention calculation to perform a weighted dot product on V. Finally, the output is mapped through a linear network to obtain the concatenated feature.

[0113] Then, the concatenated features are input into the two-dimensional convolution for convolution operation to obtain the new feature X4. X4 is input into the mean pooling module (Mean pooling), and the average pooling operation is performed according to the channel dimension to obtain a one-dimensional vector of dimension Cx*1, which is used to perform weighted operations on different channels of X4 in the scale adjustment module (Scale), that is, multiplying the elements of the corresponding channels.

[0114] The above-mentioned feature channel fusion network can filter out useless background features in the road image to be detected and can effectively extract target features of different sizes.

[0115] Finally, the target features output by the feature channel fusion network are obtained.

[0116] Step S252: Input the target features into the disease detection network to obtain the road disease detection results output by the disease detection network.

[0117] Exemplarily, the disease detection network includes a disease type detection network and a disease location detection network; inputting the target features into the disease detection network to obtain the road disease detection results output by the disease detection network, including: inputting the target features into the disease type detection network and the disease location detection network respectively to obtain the disease type output by the disease type detection network and the disease location output by the disease location detection network; combining the disease type and the disease location to obtain the road disease detection results corresponding to the road image to be detected.

[0118] For an example, see Figure 6 , Figure 6 is a schematic diagram of a disease detection network shown in an exemplary embodiment of the present application, such as Figure 6 As shown in the figure, the direction of the arrow represents the data flow. If the input target feature has a shape of c*w*h, the target feature is flattened row by row and then column by column according to the channel dimension to obtain a two-dimensional feature vector. The number of two-dimensional feature vectors is w*h. The two-dimensional feature vector is then input into the single-head self-attention module for feature extraction and then input into two linear networks for linear mapping. One linear network outputs the predicted result of the disease type, with a shape of w*h*Num, where Num is the number of predicted categories. The other linear network outputs the predicted result of the disease location, with a shape of w*h*4, where 4 represents the horizontal coordinate of the center point of the prediction box, the vertical coordinate of the center, the length of the prediction box, and the width of the prediction box.

[0119] Combined with the disease type and disease location, the road disease detection result corresponding to the road image to be detected is obtained.

[0120] Of course, depending on the actual application scenario, it is also possible to only identify the disease type or disease location, and this application does not limit this.

[0121] Exemplarily, a plurality of feature channel fusion networks of different scales are included; target features are input into a disease detection network to obtain a road disease detection result output by the disease detection network, including: obtaining target features outputted respectively by feature channel fusion networks of different scales; target features outputted respectively by feature channel fusion networks of different scales are inputted into the disease detection network to obtain a road disease detection result outputted by the disease detection network.

[0122] Setting up feature channel fusion networks of different scales to input the target features output by the feature channel fusion networks of different scales into the disease detection network, and obtaining the road disease detection results output by the disease detection network can effectively suppress false reporting of road diseases and improve the accuracy of road disease identification.

[0123] For example, a large road defect target (such as a large pothole or crack with smaller potholes or cracks inside) may detect smaller targets (small potholes or cracks inside) at the same time as the large target (large pothole or crack), resulting in the output of multiple road defect detection results, when in fact there is only one road defect. This will lead to errors in road defect statistics, and in turn, lead to the subsequent generation of incorrect road defect maintenance decisions or subsequent generation of incorrect warnings. Therefore, when performing road defect recognition on the road image to be detected, comprehensive detection is performed in combination with the target features output by the multi-scale feature channel fusion network to avoid incorrect recognition in the above situations and reduce the output of useless prediction boxes.

[0124] In some embodiments, the above-mentioned convolutional neural network, feature channel fusion network, disease detection network, etc. together constitute a road disease detection model.

[0125] For an example, see Figure 7 , Figure 7 is a schematic diagram of a road damage detection model shown in an exemplary embodiment of the present application. Figure 7 As shown in the figure, by improving the YOLO11 model, a road disease detection model is obtained. Specifically, the YOLO11 network structure consists of three parts: the backbone network (Backbone), the neck network (Neck) and the head network (Head).

[0126] Among them, the convolutional neural network composed of multiple convolution kernels with different aspect ratios in this application is used as a multi-aspect ratio convolutional network, and the Backbone is composed of the original conventional convolutional network of YOLO11, a multi-aspect ratio convolutional network, a feature channel fusion network, the original pyramid pooling network of YOLO11 (Spatial Pyramid Pooling-Fast, SPPF), and the original attention cascade network of YOLO11 (Cascade Partial Self-Attention, C2PSA).

[0127] The network structure and data processing steps of the multi-aspect ratio convolutional network and the feature channel fusion network have been described in detail in the above embodiments and will not be repeated here.

[0128] Specifically, the first conventional convolutional network increases the input channels from 3 to 16, and the length and width are halved. The remaining conventional convolutional networks all increase the input channels to twice the original and halve the length and width.

[0129] The role of SPPF is to perform pooling operations of different sizes on the input features and perform feature splicing to improve the detection ability of targets of different sizes without changing the shape of the input features.

[0130] The role of C2PSA is to enhance the spatial attention and feature extraction capabilities of the model without changing the shape of the input features.

[0131] Neck includes YOLO11's original image upsampling network (Upsample), YOLO11's original splicing network (Concat), feature channel fusion network, YOLO11's original conventional convolutional network, and multi-length and width ratio convolutional network.

[0132] Upsamle is used to upsample features, maintain the number of feature channels, and change the length and width of the features.

[0133] The head module includes a defect detection network (Trans Head), which receives target features output by feature channel fusion networks at different scales. It combines target features at different scales to identify road defect detection results, realizes semantic-level deep feature extraction of target features at different scales, and improves the recognition accuracy of the model.

[0134] The road damage detection model is trained using sample images. The sample images are marked with sample labels. The sample labels contain the true category and true frame of the road damage. By inputting the sample images into the road damage detection model to be trained, the predicted category and predicted frame output by the road damage detection model are obtained. Then, the parameters of the road damage detection model are adjusted by calculating the model loss based on the true category, true frame, predicted category, and predicted frame to obtain the trained road damage detection model. The trained road damage detection model can be deployed on a server or terminal to realize road damage detection.

[0135] Exemplarily, the model loss includes category prediction loss and target box prediction loss, wherein the category prediction loss can be calculated using cross entropy loss (Cross Entropy Loss), mean squared error loss (Mean Squared Error Loss), logarithmic loss (Log Loss), etc.

[0136] The target box prediction loss can be calculated using the following formulas 1 to 4:

[0137]

[0138] Among them, IOU is the intersection-union ratio of the predicted box and the real box, b, b gt Represents the center coordinates of the predicted box and the real box respectively, p 2 (b,b gt ) represents the Euclidean distance between the center points of the predicted box and the true box; w gt 、h gtRepresent the width and height of the real box respectively; w and h represent the width and height of the predicted box respectively; C represents the diagonal length of the minimum closed area (minimum bounding rectangle) of the predicted box and the real box.

[0139] The final target frame prediction loss WH-CIOU is calculated by formula 1 to formula 4. Compared with the related art that only considers the intersection of the prediction frame and the target frame, this application takes into account the fact that the shapes of road disease targets are variable, such as transverse cracks, longitudinal cracks, etc., that is, the length of the real frame is much larger than the width, or the width is much larger than the length. Therefore, in formula 4, the difference between the length and width of the prediction frame and the length and width of the real frame is calculated, and the loss is calculated from the length and width dimensions of the prediction frame and the real frame to accelerate model convergence and improve the model's target frame prediction accuracy for longitudinal cracks and transverse cracks.

[0140] The road defect detection method provided by the present application extracts quantitative statistical features corresponding to road defects with different aspect ratios from multiple labeled images; inputs the road images to be detected into multiple convolution kernels of a convolutional neural network for convolution processing to obtain image features convolved by each convolution kernel; based on the quantitative statistical features corresponding to road defects with different aspect ratios, sets feature fusion weights corresponding to the image features convolved by each convolution kernel; performs weighted fusion on the image features convolved by each convolution kernel according to the feature fusion weights to obtain weighted fusion features; inputs the weighted fusion features into a defect detection network to obtain road defect detection results output by the defect detection network, which can improve the information perception of road defects with different aspect ratios and avoid the situation where short side information is submerged during the convolution process. In addition, the quantitative information of road defects with different aspect ratios in real scenes is taken into account during feature fusion, so that the final weighted fusion features are more accurate, thereby improving the accuracy of subsequent road recognition.

[0141] Figure 8 FIG. 1 is a block diagram of a road damage detection device according to an exemplary embodiment of the present application. Figure 8 As shown, the exemplary road disease detection device 800 includes:

[0142] The quantitative statistics module 810 is used to obtain multiple labeled images containing road defects, where the labeled images are labeled with the aspect ratios of the road defects in the labeled images, and extract quantitative statistical features corresponding to the road defects with different aspect ratios in the multiple labeled images;

[0143] The convolution processing module 820 is used to input the road image to be detected into multiple convolution kernels of the convolutional neural network for convolution processing, thereby obtaining image features obtained by convolution of each convolution kernel. The aspect ratios of different convolution kernels correspond to the aspect ratios of different road defects.

[0144] The weight determination module 830 is used to set the feature fusion weight corresponding to the image feature obtained by convolution of each convolution kernel based on the quantitative statistical features corresponding to road diseases with different aspect ratios;

[0145] The weighted fusion module 840 is used to perform weighted fusion on the image features obtained by convolution of each convolution kernel according to the feature fusion weight to obtain a weighted fusion feature;

[0146] The disease detection module 850 is used to input the weighted fusion features into the disease detection network to obtain the road disease detection results output by the disease detection network.

[0147] It should be noted that the road defect detection device provided in the above-described embodiment and the road defect detection method provided in the above-described embodiment share the same concept. The specific manner in which each module and unit performs its operations has been described in detail in the method embodiments and will not be repeated here. In actual applications, the road defect detection device provided in the above-described embodiment can, as needed, allocate the aforementioned functions to different functional modules. This means that the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. This is not a limitation herein.

[0148] See also Figure 9 , Figure 9 This is a schematic diagram of the structure of an embodiment of an electronic device of the present application. Electronic device 900 includes memory 901 and processor 902. Processor 902 is configured to execute program instructions stored in memory 901 to implement the steps of any of the aforementioned road defect detection method embodiments. In a specific implementation scenario, electronic device 900 may include, but is not limited to, a microcomputer and a server. Furthermore, electronic device 900 may also include mobile devices such as laptops and tablet computers, without limitation herein.

[0149] Specifically, the processor 902 is used to control itself and the memory 901 to implement the steps in any of the above-mentioned road defect detection method embodiments. The processor 902 can also be called a central processing unit (CPU). The processor 902 may be an integrated circuit chip with signal processing capabilities. The processor 902 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 902 can be implemented by an integrated circuit chip.

[0150] See also Figure 10 , Figure 10 The computer-readable storage medium 1000 stores program instructions 1010 that can be executed by a processor, and the program instructions 1010 are used to implement the steps of any of the above-mentioned road defect detection method embodiments.

[0151] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0152] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0153] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0154] In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A road disease detection method, characterized in that: The method comprises: Acquire a plurality of marked images containing road damage, wherein the marked images are marked with the aspect ratios of the road damage in the marked images, and extract quantitative statistical features corresponding to road damages with different aspect ratios in the plurality of marked images; The road image to be detected is input into multiple convolution kernels of the convolutional neural network for convolution processing, and the image features obtained by convolution of each convolution kernel are obtained. The length-to-width ratios of different convolution kernels correspond to the different length-to-width ratios of road defects. Based on the quantitative statistical features corresponding to the road diseases with different aspect ratios, respectively, the feature fusion weights corresponding to the image features obtained by convolution of each convolution kernel are set; According to the feature fusion weight, the image features obtained by convolution of each convolution kernel are weightedly fused to obtain a weighted fusion feature; The weighted fusion features are input into a road damage detection network to obtain a road damage detection result output by the road damage detection network.

2. The method according to claim 1, characterized in that The quantitative statistical features include the total number of road defects with different aspect ratios in the multiple marked images; and based on the quantitative statistical features corresponding to the road defects with different aspect ratios, the feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set, including: Take any aspect ratio as the ratio to be calculated; Based on the total number of road defects containing the proportion to be calculated, a feature fusion weight corresponding to the image feature obtained by convolving the convolution kernel of the proportion to be calculated is set; wherein the total number of road defects is positively correlated with the feature fusion weight.

3. The method according to claim 1, characterized in that The marked image is further marked with whether the road damage is correctly identified or incorrectly identified, and the quantitative statistical features include the number of correctly identified road damages and the number of incorrectly identified road damages corresponding to different aspect ratios. The feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set based on the quantitative statistical features corresponding to the road damages with different aspect ratios, including: Take any aspect ratio as the ratio to be calculated; Calculating the recognition accuracy rate corresponding to the proportion of road damage to be calculated based on the number of correctly identified and the number of incorrectly identified corresponding to the proportion of road damage to be calculated; Based on the recognition accuracy corresponding to the road damage proportion to be calculated, the feature fusion weight corresponding to the image feature obtained by convolution of the convolution kernel of the proportion to be calculated is set; wherein the recognition accuracy and the feature fusion weight are inversely correlated.

4. The method according to claim 1, wherein The marked image is also marked with a road surface type, and the quantitative statistical features include the number of road defects with different length-to-width ratios corresponding to each road surface type; and based on the quantitative statistical features corresponding to the road defects with different length-to-width ratios, the feature fusion weights corresponding to the image features obtained by convolving each convolution kernel are set, including: Taking any aspect ratio as the ratio to be calculated; and identifying the road surface type corresponding to the road image to be detected to obtain a target road surface type; Calculating a correlation between the target pavement type and the road disease proportion to be calculated based on the number of road disease proportions to be calculated corresponding to the target pavement type; wherein the number of road disease proportions to be calculated is positively correlated with the correlation; Based on the correlation between the target road surface type and the road disease proportion to be calculated, a feature fusion weight corresponding to the image feature obtained by convolving the convolution kernel of the proportion to be calculated is set; wherein the correlation degree and the feature fusion weight are positively correlated.

5. The method according to any one of claims 1 to 4, characterized in that Inputting the weighted fusion features into a road damage detection network to obtain a road damage detection result output by the road damage detection network includes: The weighted fusion features are input into a feature channel fusion network, which is used to: Splitting the weighted fusion feature in the feature channel dimension to obtain a plurality of first split features; For some or all of the multiple first split features, fuse one or more other first split features, and combine the fused first split features and / or unfused first split features to obtain multiple processed features; Splicing the multiple processed features to obtain spliced ​​features, and performing convolution processing on the spliced ​​features to obtain target features output by the feature channel fusion network; The target features are input into a road damage detection network to obtain a road damage detection result output by the road damage detection network.

6. The method according to claim 5, characterized in that The step of splicing the plurality of processed features to obtain a spliced ​​feature includes: Splitting the plurality of processed features respectively in the feature channel dimension to obtain a plurality of second split features; Inputting the plurality of second split features into a single-head self-attention network, and performing self-attention calculation on the plurality of second split features based on the single-head self-attention network to obtain a self-attention calculation result; Based on the self-attention calculation result, a weighted point multiplication calculation is performed on the multiple second split features to obtain the splicing features output by the single-head self-attention network.

7. The method according to claim 5, characterized in that A feature channel fusion network containing multiple different scales; inputting the target features into a disease detection network to obtain a road disease detection result output by the disease detection network, including: Obtain target features output by feature channel fusion networks at different scales; The target features respectively output by the feature channel fusion networks of different scales are respectively input into the disease detection network to obtain the road disease detection results output by the disease detection network.

8. The method according to claim 5, characterized in that The disease detection network includes a disease type detection network and a disease location detection network; inputting the target feature into the disease detection network to obtain the road disease detection result output by the disease detection network includes: Inputting the target features into the disease type detection network and the disease location detection network respectively, to obtain the disease type output by the disease type detection network and the disease location output by the disease location detection network; In combination with the defect type and the defect location, a road defect detection result corresponding to the road image to be detected is obtained.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, and the processor is used to execute program instructions stored in the memory to implement the steps in the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and the program instructions can be executed by a processor to implement the steps in the method according to any one of claims 1 to 8.