Method and device for generating road surface detail description file based on lightweight pavement crack segmentation, equipment and medium

Through lightweight neural network model and edge computing technology, the problem of insufficient accuracy of road crack detection on resource-constrained devices is solved, and efficient and accurate real-time detection and monitoring are achieved.

CN120451977APending Publication Date: 2025-08-08WESTERN CHINA SCI CITY INNOVATION CENT OF INTELLIGENT & CONNECTED VEHICLES (CHONGQING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446033.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing deep learning models are difficult to operate efficiently on resource-constrained embedded devices, and they cannot accurately distinguish different categories of details in complex scenarios, resulting in insufficient detection accuracy of road cracks.

Method used

The lightweight neural network model is adopted, including the first convolutional network, encoding network, neck network and decoding network connected in sequence. Multi-scale features are extracted through the axial depth separation convolution and mixed pooling attention mechanism, and real-time detection is performed in combination with edge computing technology.

Benefits of technology

It improves the reliability and accuracy of road crack detection, realizes real-time road condition monitoring and reporting, and reduces the computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451977A_ABST
    Figure CN120451977A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method, a device, equipment and a medium for generating a road surface detail description file based on lightweight pavement crack segmentation, and the method comprises the steps: obtaining a road image collected by a vehicle, and extracting pavement crack information from the road image based on a trained lightweight neural network model, the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network and a second convolutional network which are connected in sequence, the encoding network comprises a plurality of encoders which are connected in sequence, the decoding network comprises a plurality of decoders which are connected in sequence, the number of the encoders is the same as that of the decoders, and the number of the neck network is the same as that of the decoders. Each encoder is connected with each decoder; and generating a pavement description file according to the road geometric structure information of the current position of the vehicle and the pavement crack information of the current position. By adopting the technical scheme, the reliability and the accuracy of a pavement crack detection result are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a method, device, equipment and medium for generating a road surface detail description file based on lightweight pavement crack segmentation. Background Art

[0002] With the continuous development of road construction and the increasing requirements for road maintenance quality, as well as more thoughts from all walks of life on the precise detection and efficient maintenance of road safety and durability, more and more road engineering-related companies and scientific research institutions have begun to enter the field of intelligent road detection, trying to seize the initiative in the precise identification and efficient analysis of road diseases.

[0003] The road surface detail description file generation device based on lightweight pavement crack segmentation has the following important functions: 1. Accurate disease assessment: It can accurately identify detailed information such as the location, shape, length and width of pavement cracks, and provide accurate disease data for road maintenance departments. 2. Improve maintenance efficiency: By quickly generating road surface detail description files, the time for manual inspection and data collation is reduced. 3. Reduce maintenance costs: Accurate crack detection can avoid excessive repairs and unnecessary road closures, thereby reducing maintenance costs. 4. Ensure road safety: Timely detection and repair of pavement cracks can effectively prevent road damage and traffic accidents caused by crack expansion. 5. Promote intelligent road management: The generated road surface detail description file can be integrated with the road management system to achieve real-time monitoring and intelligent management of road conditions.

[0004] Existing road surface detection methods often use deep learning models such as U-Net (a convolutional neural network-based architecture) and SegNet (an efficient image semantic segmentation network model). These models contain a large number of parameters, which requires high computing resources and storage space, making them difficult to run efficiently on resource-constrained embedded or mobile devices. When dealing with complex scenes, existing semantic segmentation models cannot accurately distinguish between details of different categories, resulting in insufficient segmentation accuracy and affecting the reliability of detection results. Summary of the Invention

[0005] The embodiments of the present invention provide a method, device, equipment and medium for generating a road surface detail description file based on lightweight pavement crack segmentation, thereby improving the reliability and accuracy of pavement detection results.

[0006] In a first aspect, the present invention provides a method for generating a road surface detail description file based on lightweight pavement crack segmentation, the method comprising:

[0007] Acquiring a road image collected by the vehicle;

[0008] Extracting pavement crack information from the road image based on a trained lightweight neural network model, wherein the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network, and a second convolutional network connected in sequence, the encoding network includes a plurality of encoders connected in sequence, the decoding network includes a plurality of decoders connected in sequence, the encoders and decoders are the same in number, and each encoder is connected to each decoder;

[0009] A road surface description file is generated based on road geometry information at a current position of the vehicle and road surface crack information at the current position.

[0010] Optionally, the lightweight neural network model is trained in the cloud and deployed to the vehicle after the training is completed. The sample data used to train the lightweight neural network model is obtained in the following manner:

[0011] Collecting images of road cracks;

[0012] Preprocessing the pavement crack image, wherein the preprocessing includes: grayscale processing, image denoising processing, image enhancement processing and image cropping processing;

[0013] Morphological processing is performed on the pre-processed pavement crack image, and the image data after morphological processing is used as sample data.

[0014] Optionally, each encoder includes:

[0015] The first residual axial convolution module is used to extract preliminary features of the road image;

[0016] A convolution module, configured to extract spatial features from the preliminary features;

[0017] A hybrid pooling attention mechanism module is used to extract multi-scale features from the spatial features.

[0018] Optionally, the first residual axial convolution module includes:

[0019] a first axial depth-separable convolution layer and a second axial depth-separable convolution layer, wherein the signal input ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are both signal output ends of a previous module, and the signal output ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are connected to the convolution module through a first signal superposition module, and the signal input end is connected to the first signal superposition module. When the current encoder is an encoder directly connected to the first convolution network in the encoding network, the previous module is the first convolution network; when the current encoder is not an encoder directly connected to the first convolution network in the encoding network, the previous module is a previous encoder connected to the current encoder, and the first superposition module is used to perform a sum operation on the received signal;

[0020] The first axial depth-wise separable convolutional layer includes a first horizontal convolutional layer and a first vertical convolutional layer connected in sequence, and the second axial depth-wise separable convolutional layer includes a second vertical convolutional layer and a second horizontal convolutional layer connected in sequence, wherein the size of the convolution kernel of each horizontal convolutional layer is 1×k, and the size of the convolution kernel of each vertical convolutional layer is k×1.

[0021] Optionally, the hybrid pooling attention mechanism module includes: a wavelet transform layer, a maximum pooling layer, a first sampling layer, a second sampling layer, a third sampling layer and a normalization layer, wherein,

[0022] The signal input end of the wavelet transform layer is the signal output end of the filtering module. The signal input end of the wavelet transform layer is connected to the second signal superposition module through the maximum pooling layer. The first output end of the wavelet transform layer is connected to the second signal superposition module through the first sampling layer. The second output end of the wavelet transform layer is connected to the signal input ends of the second sampling layer and the third sampling layer respectively. The signal output ends of the second sampling layer and the third sampling layer are both connected to the second signal superposition module, and the superimposed signals are output to the normalization layer. The second superposition module is used to perform a sum operation on the received signals.

[0023] The signal output by the normalization layer and the signal output by the first sampling layer are multiplied and input into the second signal superposition module. The output end of the second signal superposition module is connected to the subsequent module connected to the current encoder, wherein, when the current encoder is the last encoder of the encoding network, the subsequent module is the neck network; when the current encoder is not the last encoder of the encoding network, the subsequent module is the next encoder connected to the current encoder;

[0024] The operation of the first sampling layer includes two downsampling operations performed sequentially;

[0025] The operation of the second sampling layer includes a downsampling operation and an upsampling operation performed successively;

[0026] The operation of the third sampling layer includes an upsampling operation and a downsampling operation performed successively.

[0027] Optionally, the loss function used by the lightweight neural network model is expressed by the following formula:

[0028]

[0029] The output of the model is f(x), y represents the true label of the sample, y = 1 represents a positive sample, and y = 0 represents a negative sample; x pos is the mean of the eigenvectors of the positive samples, x neg is the mean of the feature vector of the negative sample, and θ is a threshold used to control the triggering condition of the negative sample loss.

[0030] Each decoder includes:

[0031] The upsampling module, the point-by-point convolution module, and the second residual axial convolution module are connected in sequence.

[0032] In a second aspect, an embodiment of the present invention further provides a device for generating a road surface detail description file based on lightweight pavement crack segmentation, the device comprising:

[0033] a road image acquisition module, configured to acquire a road image collected by the vehicle;

[0034] a pavement crack information extraction module configured to extract pavement crack information from the road image based on a trained lightweight neural network model, wherein the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network, and a second convolutional network connected in sequence, the encoding network includes a plurality of encoders connected in sequence, the decoding network includes a plurality of decoders connected in sequence, the encoders and decoders are the same in number, and each encoder is connected to each decoder;

[0035] The road surface description file generating module is configured to generate a road surface description file according to the road geometry information of the current position of the vehicle and the road surface crack information of the current position.

[0036] Optionally, the lightweight neural network model is trained in the cloud and deployed to the vehicle after the training is completed. The sample data used to train the lightweight neural network model is obtained in the following manner:

[0037] Collecting images of road cracks;

[0038] Preprocessing the pavement crack image, wherein the preprocessing includes: grayscale processing, image denoising processing, image enhancement processing and image cropping processing;

[0039] Morphological processing is performed on the pre-processed pavement crack image, and the image data after morphological processing is used as sample data.

[0040] Optionally, each encoder includes:

[0041] The first residual axial convolution module is used to extract preliminary features of the road image;

[0042] A convolution module, configured to extract spatial features from the preliminary features;

[0043] A hybrid pooling attention mechanism module is used to extract multi-scale features from the spatial features.

[0044] Optionally, the first residual axial convolution module includes:

[0045] a first axial depth-separable convolution layer and a second axial depth-separable convolution layer, wherein the signal input ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are both signal output ends of a previous module, and the signal output ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are connected to the convolution module through a first signal superposition module, and the signal input end is connected to the first signal superposition module. When the current encoder is an encoder directly connected to the first convolution network in the encoding network, the previous module is the first convolution network; when the current encoder is not an encoder directly connected to the first convolution network in the encoding network, the previous module is a previous encoder connected to the current encoder, and the first superposition module is used to perform a sum operation on the received signal;

[0046] The first axial depth-wise separable convolutional layer includes a first horizontal convolutional layer and a first vertical convolutional layer connected in sequence, and the second axial depth-wise separable convolutional layer includes a second vertical convolutional layer and a second horizontal convolutional layer connected in sequence, wherein the size of the convolution kernel of each horizontal convolutional layer is 1×k, and the size of the convolution kernel of each vertical convolutional layer is k×1.

[0047] Optionally, the hybrid pooling attention mechanism module includes: a wavelet transform layer, a maximum pooling layer, a first sampling layer, a second sampling layer, a third sampling layer and a normalization layer, wherein,

[0048] The signal input end of the wavelet transform layer is the signal output end of the filtering module. The signal input end of the wavelet transform layer is connected to the second signal superposition module through the maximum pooling layer. The first output end of the wavelet transform layer is connected to the second signal superposition module through the first sampling layer. The second output end of the wavelet transform layer is connected to the signal input ends of the second sampling layer and the third sampling layer respectively. The signal output ends of the second sampling layer and the third sampling layer are both connected to the second signal superposition module, and the superimposed signals are output to the normalization layer. The second superposition module is used to perform a sum operation on the received signals.

[0049] The signal output by the normalization layer and the signal output by the first sampling layer are multiplied and input into the second signal superposition module. The output end of the second signal superposition module is connected to the subsequent module connected to the current encoder, wherein, when the current encoder is the last encoder of the encoding network, the subsequent module is the neck network; when the current encoder is not the last encoder of the encoding network, the subsequent module is the next encoder connected to the current encoder;

[0050] The operation of the first sampling layer includes two downsampling operations performed sequentially;

[0051] The operation of the second sampling layer includes a downsampling operation and an upsampling operation performed successively;

[0052] The operation of the third sampling layer includes an upsampling operation and a downsampling operation performed successively.

[0053] Optionally, the loss function used by the lightweight neural network model is expressed by the following formula:

[0054]

[0055] The output of the model is f(x), y represents the true label of the sample, y = 1 represents a positive sample, and y = 0 represents a negative sample; x pos is the mean of the eigenvectors of the positive samples, x neg is the mean of the feature vector of the negative sample, and θ is a threshold used to control the triggering condition of the negative sample loss.

[0056] Optionally, each decoder includes:

[0057] The upsampling module, the point-by-point convolution module, and the second residual axial convolution module are connected in sequence.

[0058] In a third aspect, an embodiment of the present invention further provides an in-vehicle electronic device, including:

[0059] a memory storing executable program code;

[0060] a processor coupled to the memory;

[0061] The processor calls the executable program code stored in the memory to execute the method for generating a road surface detail description file based on lightweight pavement crack segmentation provided by any embodiment of the present invention.

[0062] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating a road surface detail description file based on lightweight pavement crack segmentation provided by any embodiment of the present invention.

[0063] The technical solution provided by the embodiment of the present invention is that the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network and a second convolutional network connected in sequence. The encoding network includes multiple encoders connected in sequence, and the decoding network includes multiple decoders connected in sequence. The number of encoders and decoders is the same, and each encoder is connected to each decoder. By constructing the above-mentioned lightweight neural network model, training the model, and deploying the trained model to the vehicle, the reliability and accuracy of the road crack information detection results can be improved.

[0064] The innovative features of the embodiments of the present invention include:

[0065] 1. By constructing the above-mentioned lightweight neural network model, which is composed of a first convolutional network, an encoding network, a neck network, a decoding network, and a second convolutional network connected in sequence, the encoding network includes multiple encoders connected in sequence, and the decoding network includes multiple decoders connected in sequence. The number of encoders and decoders is the same, and each encoder is connected to each decoder. By training this model, the reliability and accuracy of pavement crack information detection results can be improved, which is one of the innovations of the embodiments of the present invention.

[0066] 2. By adopting wavelet downsampling technology, the multi-scale features of the image can be effectively extracted. By adopting the axial depth-separable convolution layer, it can help capture more contextual information in different directions, improve the segmentation accuracy while retaining image details, improve the reliability and accuracy of the pavement crack detection results, and help improve the segmentation performance and robustness of the model in complex scenarios. This is one of the innovations of the embodiment of the present invention.

[0067] 3. By leveraging edge computing technology and real-time data processing capabilities, after the trained lightweight neural network model is deployed to the vehicle's edge detection equipment, crack detection and segmentation can be performed locally on the vehicle, enabling real-time road condition monitoring and reporting. This improves the real-time nature and response speed of road inspections, which is one of the innovations of the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0069] Figure 1a A training flow chart of a lightweight neural network model provided in Example 1 of the present invention;

[0070] Figure 1b This is a diagram showing the effect of the cropping process provided in the first embodiment of the present invention;

[0071] Figure 1c A structural block diagram of a lightweight neural network model provided in Example 1 of the present invention;

[0072] Figure 1d A schematic structural diagram of a specific lightweight neural network model provided in Example 1 of the present invention;

[0073] Figure 1e A schematic structural diagram of an encoder provided in accordance with the first embodiment of the present invention;

[0074] Figure 1f This is a structural block diagram of a first residual axial convolution module provided in Example 1 of the present invention;

[0075] Figure 1g A schematic structural block diagram of a specific first residual axial convolution module provided in the first embodiment of the present invention;

[0076] Figure 1h This is a structural block diagram of a hybrid pooling attention mechanism module provided in Example 1 of the present invention;

[0077] Figure 2 A flowchart of a method for generating a road surface detail description file based on lightweight pavement crack segmentation provided in the second embodiment of the present invention;

[0078] Figure 3 This is a structural block diagram of a device for generating a road surface detail description file based on lightweight pavement crack segmentation provided in a third embodiment of the present invention;

[0079] Figure 4 This is a structural diagram of a vehicle-mounted electronic device provided in Example 4 of the present invention. DETAILED DESCRIPTION

[0080] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0081] It should be noted that the terms "including," "having," and any variations thereof in the embodiments of the present invention and the accompanying drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0082] The present invention discloses a method, device, apparatus, and medium for generating a road surface detail description file based on lightweight pavement crack segmentation. The following describes the technical solution provided by the present invention in detail from the perspectives of training and application of a lightweight neural network model.

[0083] Example 1

[0084] Figure 1a This is a training flow chart of a lightweight neural network model provided by the first embodiment of the present invention. The lightweight neural network model can be trained and updated in the cloud without deploying high-power computing equipment. Figure 1a As shown, the method provided in this embodiment includes:

[0085] S110: Collect road crack images.

[0086] In this embodiment, a road crack image can be collected by a vehicle-mounted image acquisition device, wherein the vehicle-mounted image acquisition device can be a vehicle-mounted camera, a vehicle-mounted radar, etc. The road crack image refers to a road image containing crack information.

[0087] S120: Preprocess the pavement crack image.

[0088] In this embodiment, pavement crack image preprocessing is an important preliminary step for crack detection and analysis, and mainly includes five steps: image grayscale conversion, image denoising, image enhancement, image cropping, and image output.

[0089] Among them, for image grayscale, during the image grayscale process, the color RGB image can be converted into a grayscale image by taking a weighted average of the values on the RGB (red, green, and blue color mode) three channels.

[0090] Specifically, let the pixel coordinates be (x, y), the values corresponding to the red, green, and blue (R, G, B) channels be R(x, y), G(x, y), and B(x, y), respectively, and let the grayscale value be Gray(x, y). The weights of the three channels are set. Based on engineering experience, they are generally set to R: 0.299, G: 0.587, and B: 0.114R. The grayscale value of the image can then be calculated:

[0091] Gray(x,y)=0.299·R(x,y)+0.587·G(x,y)+0.114R(x,y)

[0092] After obtaining the grayscale image, the noise in the image can be effectively removed by Gaussian filtering. Let a frame of image be I, and the filter window size be (2k+1)×(2k+1) (k is a positive integer, for example, when k=1, the filter window size is 3x3), then:

[0093]

[0094] Among them, G(m,n) is the discretized Gaussian function value, which needs to be obtained by discrete sampling of the continuous two-dimensional Gaussian function. In practical applications, in order to ensure that the total pixel value of the filtered image remains unchanged (energy conservation), the Gaussian function is usually normalized so that

[0095] For image enhancement processing, the grayscale histogram of the image can be adjusted to make the grayscale distribution of the image more uniform. Assume that the grayscale range of the original image is [0, L-1] (for 8-bit grayscale image, L=256), n i Represents the number of pixels with gray level, N is the total number of pixels in the image, then the probability p of gray level appearing in the original image i for:

[0096] The cumulative distribution function CDF (Cumulative Distribution Function) is:

[0097] After histogram equalization, the mapping relationship between the new gray level j and the original gray level i is:

[0098] j = round((L-1)×CDF(i))

[0099] Here, round() represents the rounding function, which is used to convert the calculated result into an integer grayscale value. The function of this formula is to remap the grayscale values of the original image according to its cumulative distribution function, so that the grayscale distribution of the new image is more uniform.

[0100] The cropping process of the image is to select a quadrilateral area in the ground image directly in front of the vehicle. The range can be directly calibrated on the vehicle, and then the rectangular image required for training is obtained through perspective transformation. Figure 1b This is a cutting effect diagram provided by the first embodiment of the present invention, such as Figure 1b As shown in FIG, for a selected area in the ground image directly in front of the vehicle, a corresponding rectangular image can be obtained through perspective transformation.

[0101] S130 , performing morphological processing on the pre-processed pavement crack image, and using the image data after morphological processing as sample data.

[0102] Morphological processing includes dilation and erosion. The morphological processing process is further divided into opening and closing operations. Opening involves performing erosion first, followed by dilation. Closing involves performing dilation first, followed by erosion. In the embodiment of the present invention, closing is selected for processing. Finally, a thinning algorithm is used to thin objects (such as cracks) in the image into lines with a single pixel width, while maintaining the topological structure of the objects and extracting the centerline of the crack in the image.

[0103] S140. Train the lightweight neural network model based on the sample data.

[0104] In this embodiment, Figure 1c As shown, the lightweight neural network model is composed of a first convolutional network 1, an encoding network 2, a neck network 3, a decoding network 4 and a second convolutional network 5 connected in sequence, wherein the encoding network includes a plurality of encoders connected in sequence, and the decoding network includes a plurality of decoders connected in sequence. The number of encoders and decoders is the same, and each encoder is connected to each decoder. The embodiment of the present invention does not specifically limit the number of encoders and decoders.

[0105] In the lightweight neural network structure provided in this embodiment, the network image is first feature extracted by the encoder to obtain semantic information. Before decoding, a bridging operation will be performed through the neck network. In order to avoid an increase in the number of parameters, the neck network provided in this embodiment can use residual axial convolution for multi-branch parallel processing. Each decoder is mainly composed of an upsampling module, a point-by-point convolution module and a second residual axial convolution module connected in sequence, wherein the second residual axial convolution module can be set to the same structure as the first residual axial convolution module, and this embodiment does not specifically limit this.

[0106] Specifically, Figure 1d This is a schematic diagram of the structure of a specific lightweight neural network model provided in Example 1 of the present invention. Figure 1dFive encoders (Encoder1~Encoder5) and five decoders (Decoder1~Decoder5) are shown in the figure. In the structure diagram, Image is the input picture, Conv7×7 is the first convolutional network, which is used to perform convolution processing on the input image (Image), Conv1×1 is the second convolutional network, which is used to perform convolution processing on the image data output by the encoder, BottleNeck is the neck network, Predicted Mask is the predicted binary crack image, and H and W are the width and height of the image after encoding, respectively.

[0107] The core structure of the lightweight neural network model is described in detail below.

[0108] Figure 1e This is a schematic diagram of the structure of an encoder provided in the first embodiment of the present invention, as shown in FIG. Figure 1e As shown, the encoder provided by this embodiment includes a first residual axial convolution module 21, a convolution module 22 and a mixed pooling attention mechanism module 23, wherein,

[0109] The first residual axial convolution module 21 is used to extract preliminary features of the road image;

[0110] The convolution module 22 is used to extract spatial features from the preliminary features. Specifically, the convolution kernel of the convolution module can be 1×1.

[0111] The hybrid pooling attention mechanism module 23 is used to extract multi-scale features from the spatial features.

[0112] Among them, for the first residual axial convolution module 21, Figure 1f This is a structural block diagram of a first residual axial convolution module provided in the first embodiment of the present invention, such as Figure 1f As shown, the first residual axial convolution module includes: a first axial depth-separable convolution layer 211 and a second axial depth-separable convolution layer 212, wherein the signal input ends of the first axial depth-separable convolution layer 211 and the second axial depth-separable convolution layer 212 are both signal output ends of the previous module, and the signal output ends of the first axial depth-separable convolution layer 211 and the second axial depth-separable convolution layer 212 are connected to the convolution module 22 through the first signal superposition module, and the signal input ends of the first axial depth-separable convolution layer 211 and the second axial depth-separable convolution layer 212 are connected to the first signal superposition module. The function of the first signal superposition module is to perform a sum operation on the signals input to the module.

[0113] It should be noted that, in the case where the current encoder is the first encoder directly connected to the first convolutional network in the encoding network, the previous module connected to the signal input ends of the first axial depth-separable convolutional layer 211 and the second axial depth-separable convolutional layer 212 is the first convolutional network; in the case of an encoder directly connected to the first convolutional network in the current encoder non-encoding network, the previous module connected to the signal input ends of the first axial depth-separable convolutional layer 211 and the second axial depth-separable convolutional layer 212 is the previous encoder connected to the current encoder.

[0114] Specifically, such as Figure 1f As shown, the first axial depth-separable convolution layer 211 includes a first horizontal convolution layer 2111 and a first vertical convolution layer 2112 connected in sequence, and the second axial depth-separable convolution layer 212 includes a second vertical convolution layer 2121 and a second horizontal convolution layer 2122 connected in sequence.

[0115] Figure 1g A schematic diagram of the structure of a specific first residual axial convolution module provided in the first embodiment of the present invention is shown as follows: Figure 1g As shown, the horizontal convolution layers and vertical convolution layers in the first axial depth-separable convolution layer 211 and the second axial depth-separable convolution layer 212 are both convolution modules, wherein the sizes of the convolution kernels of the first horizontal convolution layer and the second horizontal convolution layer can be set to 1×k, and the sizes of the convolution kernels of the first vertical convolution layer and the second vertical convolution layer can be set to k×1.

[0116] In this embodiment, during initialization, two depth-separable convolution layers, a first axial depth-separable convolution layer and a second axial depth-separable convolution layer, are defined. Each depth-separable convolution layer is used to perform convolution operations on the input tensor in the horizontal and vertical directions, respectively. In the forward propagation, the input tensor passes through the depth-separable convolution layers in the horizontal and vertical directions, and then is added to the original input tensor to obtain the final output. The advantage of this axial depth-separable convolution structure is that it can effectively reduce the number of parameters and computational complexity of the model. By sharing weights, the convolution operations in the horizontal and vertical directions can share the same convolution kernel, thereby greatly reducing the number of parameters. In addition, the introduction of residual connections ensures that the gradient of the network can be better propagated, speeds up the training speed, and also helps to alleviate the problem of gradient disappearance or explosion. This structure has low computational cost and memory consumption while ensuring model performance, and is suitable for lightweight and efficient deep learning model design.

[0117] In addition, for the hybrid pooling attention mechanism module 23 in the encoder, the module includes: a wavelet transform layer 231, a maximum pooling layer 232, a first sampling layer 233, a second sampling layer 234, a third sampling layer 235 and a normalization layer 236. Specifically, Figure 1h This is a structural block diagram of a hybrid pooling attention mechanism module provided in Example 1 of the present invention, such as Figure 1h As shown,

[0118] The signal input end of the wavelet transform layer 231 is the signal output end of the filtering module in the encoder. The signal input end of the wavelet transform layer 231 is connected to the second signal superposition module through the maximum pooling layer 232. The first output end of the wavelet transform layer 231 is connected to the second signal superposition module through the first sampling layer. The second output end of the wavelet transform layer 231 is connected to the signal input ends of the second sampling layer 234 and the third sampling layer 235 respectively. The signal output ends of the second sampling layer 234 and the third sampling layer 235 are both connected to the second signal superposition module, and the superimposed signal is output to the normalization layer 236; the signal output of the normalization layer 236 and the signal output of the first sampling layer 233 are multiplied and input to the second signal superposition module. The output end of the second signal superposition module is connected to the next module connected to the current encoder. The function of the second signal superposition module is to perform a sum operation on the signals input thereto.

[0119] It should be noted that, when the current encoder is the last encoder of the encoding network, the subsequent module connected to the current encoder is the neck network in the lightweight neural network model; when the current encoder is not the last encoder of the encoding network, the subsequent module connected to the current encoder is the next encoder connected to the current encoder.

[0120] In this embodiment, the operation of the first sampling layer includes two downsampling operations performed sequentially, the operation of the second sampling layer includes a downsampling operation and an upsampling operation performed sequentially, and the operation of the third sampling layer includes an upsampling operation and a downsampling operation performed sequentially.

[0121] In this embodiment, a hybrid wavelet pooling attention mechanism is used as a downsampling module. Wavelet pooling downsampling has greater advantages than maximum pooling and average pooling in terms of feature extraction, information retention, noise resistance and compression effect, and is suitable for signal processing and image processing tasks that require multi-scale analysis and frequency localization processing. Wavelet downsampling can extract features at different scales, while maximum pooling and average pooling can only extract features within a fixed window. By mixing wavelet downsampling and average pooling, this embodiment is more conducive to capturing multi-scale features in signals or images, thereby helping to improve the diversity and richness of feature extraction.

[0122] In the lightweight neural network structure provided in this embodiment, the encoder is first used to extract features from the network image to obtain semantic information. Before decoding, a bridge operation is performed through the neck network. In order to avoid an increase in the number of parameters, the neck network provided in this embodiment is also based on Figure 1fThe residual axial convolution shown in the figure performs multi-branch parallel operation. Each decoder mainly comprises an upsampling module, a point-by-point convolution module, and a second residual axial convolution module connected in sequence. The second residual axial convolution module can be set to the same structure as the first residual axial convolution module, which is not specifically limited in this embodiment.

[0123] In this embodiment, in pavement crack detection, the number of crack pixels (positive samples) may be far less than the number of non-crack pixels (negative samples). Considering that traditional loss functions may cause the model to overfit the majority class (negative samples) in this case, while having poor recognition ability for the minority class (positive samples), this embodiment can select and construct a similarity imbalance loss function through similarity measurement to solve this problem.

[0124] Specifically, for two vectors x and y, the cosine similarity formula is Where x·y is the vector dot product, and |x| and |y| are the moduli of vectors x and y, respectively. Assume the model output is f(x), and the true label is y (y = 1 for a positive sample, y = 0 for a negative sample). For positive samples, we want the model output to have a high similarity to the positive sample feature vector; for negative samples, we want a low similarity. The loss function can be designed as follows:

[0125]

[0126] where x pos is the mean of the eigenvectors of the positive samples, x neg is the mean of the feature vectors of negative samples, and θ is a threshold used to control the triggering condition of the negative sample loss. This loss function encourages the model to produce high similarity output for positive samples and low similarity output for negative samples, and adjusts the sensitivity to imbalanced data through the threshold.

[0127] In the technical solution provided in this embodiment, a lightweight neural network model is constructed. The model is composed of a first convolutional network, an encoding network, a neck network, a decoding network, and a second convolutional network connected in sequence. The encoding network includes multiple encoders connected in sequence, and the decoding network includes multiple decoders connected in sequence. The number of encoders and decoders is the same, and each encoder is connected to each decoder. By constructing the above-mentioned lightweight neural network model and training the model, the reliability and accuracy of the pavement crack information detection results can be improved.

[0128] Furthermore, the trained lightweight neural network model can be deployed on the edge computing device on the vehicle, and the road crack segmentation algorithm can be updated regularly through OTA (Over-the-Air Technology).

[0129] Next, the process of using a lightweight neural network model to obtain pavement crack information and generate a pavement description file is introduced.

[0130] Example 2

[0131] Figure 2 This is a flowchart of a method for generating a road surface detail description file based on lightweight pavement crack segmentation provided in the second embodiment of the present invention. This method can be applied to vehicle-mounted terminals such as vehicle-mounted computers and vehicle-mounted edge computing devices, and can also be applied to servers. This embodiment of the present invention does not limit this. The method provided in this embodiment can be executed by a device for generating a road surface detail description file based on lightweight pavement crack segmentation, and the device can be implemented in software and / or hardware. Figure 2 As shown, the method provided in this embodiment specifically includes:

[0132] S210: Acquire a road image collected by a vehicle.

[0133] S220. Extract pavement crack information from the road image based on the trained lightweight neural network model.

[0134] Among them, the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network and a second convolutional network connected in sequence. The encoding network includes multiple encoders connected in sequence, and the decoding network includes multiple decoders connected in sequence. The number of encoders and decoders is the same, and each encoder is connected to each decoder. The specific structure and function of each network in the lightweight neural network model, as well as the training process of the lightweight neural network model, can be specifically referred to the description of the above embodiment and will not be repeated here.

[0135] S230: Generate a road surface description file based on the road geometry information at the current position of the vehicle and the road surface crack information at the current position.

[0136] In this embodiment, the vehicle's current location information (including longitude, latitude, and altitude) can be obtained through the vehicle-mounted positioning device, and the road geometry information of the current location (such as the specific location of the road, road geometry dimensions, etc.) can be obtained through a high-precision area.

[0137] For road images captured by the vehicle, the specific location information corresponding to the road image can be determined based on real-time positioning information obtained from the onboard positioning device. Accurate road crack information can then be extracted from the road image using a lightweight neural network model deployed on the vehicle. Attributes can then be added to this crack information, including the specific location of the road surface and the presence of cracks. Based on the constructed road geometry and texture information, a road surface description file can be rendered and generated. For example, an OpenCRG (Open Curved Regular Grid) file can be generated. This file is an open format mesh standard designed for complex surface modeling and data analysis, primarily used in engineering and scientific fields requiring high-precision surface descriptions.

[0138] In this embodiment, with the help of edge computing technology and real-time data processing capabilities, road crack detection and segmentation can be performed locally on the vehicle based on a lightweight neural network model deployed to the vehicle, realizing real-time road condition monitoring and reporting, and improving the real-time and response speed of road inspections.

[0139] Example 3

[0140] Figure 3 This is a structural block diagram of a device for generating a road surface detail description file based on lightweight pavement crack segmentation provided in the third embodiment of the present invention, such as Figure 3 As shown, the device includes: a road image acquisition module 310, a road crack information extraction module 320 and a road description file generation module 330, wherein,

[0141] A road image acquisition module 310 is configured to acquire a road image collected by the vehicle;

[0142] a pavement crack information extraction module 320 configured to extract pavement crack information from the road image based on a trained lightweight neural network model, wherein the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network, and a second convolutional network connected in sequence, the encoding network including a plurality of encoders connected in sequence, the decoding network including a plurality of decoders connected in sequence, the encoders and decoders being equal in number, and each encoder being connected to each decoder;

[0143] The road surface description file generating module 330 is configured to generate a road surface description file according to the road geometry information of the current position of the vehicle and the road surface crack information of the current position.

[0144] Optionally, the lightweight neural network model is trained in the cloud and deployed to the vehicle after the training is completed. The sample data used to train the lightweight neural network model is obtained in the following manner:

[0145] Collecting images of road cracks;

[0146] Preprocessing the pavement crack image, wherein the preprocessing includes: grayscale processing, image denoising processing, image enhancement processing and image cropping processing;

[0147] Morphological processing is performed on the pre-processed pavement crack image, and the image data after morphological processing is used as sample data.

[0148] Optionally, each encoder includes:

[0149] The first residual axial convolution module is used to extract preliminary features of the road image;

[0150] A convolution module, configured to extract spatial features from the preliminary features;

[0151] A hybrid pooling attention mechanism module is used to extract multi-scale features from the spatial features.

[0152] Optionally, the first residual axial convolution module includes:

[0153] a first axial depth-separable convolution layer and a second axial depth-separable convolution layer, wherein the signal input ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are both signal output ends of a previous module, and the signal output ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are connected to the convolution module through a first signal superposition module, and the signal input end is connected to the first signal superposition module. When the current encoder is an encoder directly connected to the first convolution network in the encoding network, the previous module is the first convolution network; when the current encoder is not an encoder directly connected to the first convolution network in the encoding network, the previous module is a previous encoder connected to the current encoder, and the first superposition module is used to perform a sum operation on the received signal;

[0154] The first axial depth-wise separable convolutional layer includes a first horizontal convolutional layer and a first vertical convolutional layer connected in sequence, and the second axial depth-wise separable convolutional layer includes a second vertical convolutional layer and a second horizontal convolutional layer connected in sequence, wherein the size of the convolution kernel of each horizontal convolutional layer is 1×k, and the size of the convolution kernel of each vertical convolutional layer is k×1.

[0155] Optionally, the hybrid pooling attention mechanism module includes: a wavelet transform layer, a maximum pooling layer, a first sampling layer, a second sampling layer, a third sampling layer and a normalization layer, wherein,

[0156] The signal input end of the wavelet transform layer is the signal output end of the filtering module. The signal input end of the wavelet transform layer is connected to the second signal superposition module through the maximum pooling layer. The first output end of the wavelet transform layer is connected to the second signal superposition module through the first sampling layer. The second output end of the wavelet transform layer is connected to the signal input ends of the second sampling layer and the third sampling layer respectively. The signal output ends of the second sampling layer and the third sampling layer are both connected to the second signal superposition module, and the superimposed signals are output to the normalization layer. The second superposition module is used to perform a sum operation on the received signals.

[0157] The signal output by the normalization layer and the signal output by the first sampling layer are multiplied and input into the second signal superposition module. The output end of the second signal superposition module is connected to the subsequent module connected to the current encoder, wherein, when the current encoder is the last encoder of the encoding network, the subsequent module is the neck network; when the current encoder is not the last encoder of the encoding network, the subsequent module is the next encoder connected to the current encoder;

[0158] The operation of the first sampling layer includes two downsampling operations performed sequentially;

[0159] The operation of the second sampling layer includes a downsampling operation and an upsampling operation performed successively;

[0160] The operation of the third sampling layer includes an upsampling operation and a downsampling operation performed successively.

[0161] Optionally, the loss function used by the lightweight neural network model is expressed by the following formula:

[0162]

[0163] The output of the model is f(x), y represents the true label of the sample, y = 1 represents a positive sample, and y = 0 represents a negative sample; x pos is the mean of the eigenvectors of the positive samples, x neg is the mean of the feature vector of the negative sample, and θ is a threshold used to control the triggering condition of the negative sample loss.

[0164] Optionally, each decoder includes:

[0165] The upsampling module, the point-by-point convolution module, and the second residual axial convolution module are connected in sequence.

[0166] Example 4

[0167] See also Figure 4 , Figure 4This is a structural diagram of a vehicle-mounted electronic device provided by the fourth embodiment of the present invention. Figure 4 As shown, the vehicle-mounted electronic equipment may include:

[0168] A memory 701 storing executable program code;

[0169] a processor 702 coupled to the memory 701;

[0170] The processor 702 calls the executable program code stored in the memory 701 to execute the method for generating a road surface detail description file based on lightweight pavement crack segmentation provided by any embodiment of the present invention.

[0171] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the method for generating a road surface detail description file based on lightweight pavement crack segmentation provided by any embodiment of the present invention.

[0172] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the above-mentioned processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0173] In the embodiments provided herein, it should be understood that "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.

[0174] In addition, the functional units in the embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0175] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests for causing a computer device (which can be a personal computer, server, or network device, specifically a processor in the computer device) to execute some or all of the steps of the above-mentioned methods of various embodiments of the present invention.

[0176] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0177] Those skilled in the art will appreciate that the accompanying drawings are merely schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.

[0178] Those skilled in the art will appreciate that the modules in the apparatuses of the embodiments may be distributed in the apparatuses of the embodiments as described in the embodiments, or may be located in one or more apparatuses different from the embodiments with corresponding changes. The modules in the above embodiments may be combined into one module or further divided into multiple sub-modules.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a road surface detail description file based on lightweight pavement crack segmentation, characterized in that: include: Acquiring a road image collected by the vehicle; Extracting pavement crack information from the road image based on a trained lightweight neural network model, wherein the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network, and a second convolutional network connected in sequence, the encoding network includes a plurality of encoders connected in sequence, the decoding network includes a plurality of decoders connected in sequence, the encoders and decoders are the same in number, and each encoder is connected to each decoder; A road surface description file is generated based on road geometry information at a current position of the vehicle and road surface crack information at the current position.

2. The method according to claim 1, characterized in that The lightweight neural network model is trained in the cloud and deployed to the vehicle after training is completed; The sample data used to train the lightweight neural network model is obtained in the following way: Collecting images of road cracks; Preprocessing the pavement crack image, wherein the preprocessing includes: grayscale processing, image denoising processing, image enhancement processing and image cropping processing; Morphological processing is performed on the pre-processed pavement crack image, and the image data after morphological processing is used as sample data.

3. The method according to claim 1 or 2, characterized in that Each encoder includes: The first residual axial convolution module is used to extract preliminary features of the road image; A convolution module, configured to extract spatial features from the preliminary features; A hybrid pooling attention mechanism module is used to extract multi-scale features from the spatial features.

4. The method according to claim 3, characterized in that The first residual axial convolution module includes: a first axial depth-separable convolution layer and a second axial depth-separable convolution layer, wherein the signal input ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are both signal output ends of a previous module, and the signal output ends of the first axial depth-separable convolution layer and the second axial depth-separable convolution layer are connected to the convolution module through a first signal superposition module, and the signal input end is connected to the first signal superposition module. When the current encoder is an encoder directly connected to the first convolution network in the encoding network, the previous module is the first convolution network; when the current encoder is not an encoder directly connected to the first convolution network in the encoding network, the previous module is a previous encoder connected to the current encoder, and the first superposition module is used to perform a sum operation on the received signal; The first axial depth-wise separable convolutional layer includes a first horizontal convolutional layer and a first vertical convolutional layer connected in sequence, and the second axial depth-wise separable convolutional layer includes a second vertical convolutional layer and a second horizontal convolutional layer connected in sequence, wherein the size of the convolution kernel of each horizontal convolutional layer is 1×k, and the size of the convolution kernel of each vertical convolutional layer is k×1.

5. The method according to claim 3, characterized in that The hybrid pooling attention mechanism module includes: a wavelet transform layer, a maximum pooling layer, a first sampling layer, a second sampling layer, a third sampling layer and a normalization layer, wherein, The signal input end of the wavelet transform layer is the signal output end of the filtering module. The signal input end of the wavelet transform layer is connected to the second signal superposition module through the maximum pooling layer. The first output end of the wavelet transform layer is connected to the second signal superposition module through the first sampling layer. The second output end of the wavelet transform layer is connected to the signal input ends of the second sampling layer and the third sampling layer respectively. The signal output ends of the second sampling layer and the third sampling layer are both connected to the second signal superposition module, and the superimposed signals are output to the normalization layer. The second superposition module is used to perform a sum operation on the received signals. The signal output by the normalization layer and the signal output by the first sampling layer are multiplied and input into the second signal superposition module. The output end of the second signal superposition module is connected to the subsequent module connected to the current encoder, wherein, when the current encoder is the last encoder of the encoding network, the subsequent module is the neck network; when the current encoder is not the last encoder of the encoding network, the subsequent module is the next encoder connected to the current encoder; The operation of the first sampling layer includes two downsampling operations performed sequentially; The operation of the second sampling layer includes a downsampling operation and an upsampling operation performed successively; The operation of the third sampling layer includes an upsampling operation and a downsampling operation performed successively.

6. The method according to claim 1, characterized in that The loss function used in the lightweight neural network model is expressed by the following formula: The output of the model is f(x), y represents the true label of the sample, y = 1 represents a positive sample, and y = 0 represents a negative sample; x pos is the mean of the eigenvectors of the positive samples, x neg is the mean of the feature vector of the negative sample, and θ is a threshold used to control the triggering condition of the negative sample loss.

7. The method according to claim 1, characterized in that Each decoder includes: The upsampling module, the point-by-point convolution module, and the second residual axial convolution module are connected in sequence.

8. A device for generating a road surface detail description file based on lightweight pavement crack segmentation, characterized in that: include: a road image acquisition module, configured to acquire a road image collected by the vehicle; a pavement crack information extraction module configured to extract pavement crack information from the road image based on a trained lightweight neural network model, wherein the lightweight neural network model is composed of a first convolutional network, an encoding network, a neck network, a decoding network, and a second convolutional network connected in sequence, the encoding network includes a plurality of encoders connected in sequence, the decoding network includes a plurality of decoders connected in sequence, the encoders and decoders are the same in number, and each encoder is connected to each decoder; The road surface description file generating module is configured to generate a road surface description file according to the road geometry information of the current position of the vehicle and the road surface crack information of the current position.

9. An in-vehicle electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a road surface detail description file based on lightweight pavement crack segmentation as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for generating a road surface detail description file based on lightweight pavement crack segmentation as described in any one of claims 1 to 7 is implemented.