Tire defect detection method
By constructing the MAC module of the parallel cavity convolution module to replace the pooling operation in the Unet model and constructing the DMA Unet model, the problems of low efficiency and low accuracy of tire defect detection in the prior art are solved, and efficient and accurate tire defect detection are achieved.
Patent Information
- Application Number
- CN202510660324.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
In the prior art, tire defect detection methods are inefficient and have low accuracy, making it difficult to accurately identify defects in tire patterns.
Build a MAC module, including four parallel hollow convolution modules, a convolution module and output module, replace the pooling operation in the Unet model, build a DMA Unet model, and detect defects on the tire surface.
By retaining detailed information through the hollow convolution module, it can better handle subtle crack characteristics, improve the ability to identify cracks in complex backgrounds, and achieve high efficiency and accuracy tire defect detection.
Smart Images

Figure CN120182275A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and specifically to a method for detecting tire defects. Background Art
[0002] The tire is one of the important components of an automobile, and it bears the weight of the entire vehicle body. Therefore, the quality of the tire is directly related to the overall safety of the automobile. There may be relatively serious defects on the tread part of the automobile tire. Common defects include: severe wear, rubber chunks missing, severe cracks or fissures. These defects may cause a tire blowout during high-speed driving or emergency braking. Therefore, before the tire is manufactured and leaves the factory, as well as during the daily inspection and maintenance of the automobile, the defects on the tire will be detected.
[0003] In the prior art, many repair shops or factories detect tire defects based on physical methods, and detect whether there are defects on the tire through tests of light transmittance, airtightness or water tightness. When implementing this physical method, the tire must be taken out separately and placed on a special detection device, which is very inconvenient and has low detection efficiency. With the popularization of image recognition algorithms and neural network models, there are many image recognition algorithms based on neural networks in the field of civil engineering for identifying defects such as cracks on road surfaces, bridge decks or mine cave walls. However, in order to increase the grip and anti-slip drainage, many patterns are designed on the tire surface. These patterns and defects both have three-dimensional dimensions, and the occurrence positions of the tire defects on the tire are irregular, and the sizes are also irregular. The difficulty of tire defect recognition lies in distinguishing these defect problems among a large number of patterns. Therefore, if the crack recognition algorithms in other fields are directly applied to the crack recognition on the tire surface, the recognition accuracy is very low. Summary of the Invention
[0004] In order to solve the problems of low efficiency or inaccurate detection results of the existing methods for detecting tire defects, the present invention provides a method for detecting tire defects, which can identify the defects on the tire surface with high efficiency and accuracy.
[0005] The technical solution of the present invention is as follows: A method for detecting tire defects, characterized in that it includes the following steps: S1: Construct a MAC module; The MAC module includes: an input module, four parallel atrous convolution modules, a convolution module, and an output module that are connected in sequence; Four parallel dilated convolution modules are respectively set with dilation rates. The input feature map is simultaneously fed into the four parallel dilated convolution modules, and the output channels are adjusted and the feature information of the input feature map is extracted through different dilation rates. Then, the feature information extracted by the four-way parallel dilated convolution is superimposed together based on the Concat operation, and then fed into a 1*1 convolution module to adjust the number of channels of the feature map. Finally, the feature map is output based on the output module; S2: Based on the Unet model, replace the original pooling operation in the Unet model with the MAC module to construct the DMAUnet model; S3: Build a tire defect detection model based on the DMA Unet model; The tire defect detection model includes: a feature input module, a DMA Unet model, and a detection result output module connected in sequence; The DMA Unet model detects and enhances the defects in the input feature map to obtain an enhanced feature map. The output of the DMA Unet model includes: the enhanced feature map and the feature value corresponding to each pixel point in the enhanced feature map; The detection result output module judges the defect features existing in the enhanced feature map. The specific judgment method is: Compare the feature value corresponding to each pixel point with a preset defect judgment threshold. When the feature value of the pixel point is greater than the defect judgment threshold, it is judged that the pixel point has a defect; S4: Build a training data set and a validation data set based on various tire shooting data; Use the training data set to train the tire defect detection model to obtain the trained tire defect detection model; S5: Based on the image acquisition device, collect tire images in real time and feed them into the trained tire defect detection model. The tire defect detection model outputs the detection results for each input image; The detection results include: the enhanced feature map and the feature value corresponding to each pixel point in the enhanced feature map.
[0006] It is further characterized in that: The DMA Unet model includes: an encoder and a decoder, and the encoder and the decoder are connected through a bottleneck layer; The encoder includes N downsampling modules, and each of the downsampling modules includes: a densely connected convolutional block, a convolutional layer, and a MAC Module connected in sequence; the decoder includes N upsampling modules, and each of the upsampling modules includes: an UpSampling layer and a convolutional layer connected in sequence, and the UpSampling layer is implemented based on transposed convolution; where N≥1; a cross-layer skip connection is established between the convolutional layer in the same-level upsampling module and the convolutional layer in the downsampling module; The bottleneck layer is implemented based on M consecutive convolutional layers, where M≥1; The original input image is fed into the encoder to gradually extract features through multiple convolutional layers and pooling layers, obtaining N-level feature maps. At the same time, each downsampling module uses a densely connected manner based on the densely connected convolutional block to transfer each level of feature map to the next layer, so that the input obtained by each layer in the encoder is not only the output of the previous layer, but the output of all network layers before the current layer; the bottleneck layer is after the encoder, and further extracts the feature information of the deepest layer through M convolutional operations; the decoder uses the UpSampling layer to gradually restore the spatial size of the image, and at the same time uses skip connections to fuse the feature information from the encoder; In step S5, the image acquisition device is implemented based on an in-vehicle camera; The trained tire defect detection model is preset in the intelligent vehicle system; The feature input module includes: an image clarity determination module; in the image clarity determination module, the timing of receiving the image transmitted by the in-vehicle camera is determined according to the real-time vehicle speed, and whether the received input image is available is determined according to the clarity of the tire tread; if it is determined that the image is available, the image collected by the camera with the tire surface is sent into the tire defect detection model in real time to detect the defects on the tire surface in real time; The detection result output module further includes a defect determination condition, and the defect determination condition is: the defective pixel points are continuous and the number of pixel points is greater than a preset continuous threshold; when the defect determination condition is met, it is determined that there are defects on the tire surface corresponding to the input image; The dilation rates of the four dilated convolutional modules in the MAC module are respectively set to: 6, 12, 18, 24; The value of N in the decoder and the encoder is 3; the value of M in the bottleneck layer is 2; The defect judgment threshold is set to 0.5; The defects include: wear, rubber chunks missing, and cracks; The detection result output module further includes: a marking module, once the defective pixel points are detected, the position of the defect will be highlighted in the enhanced feature map.
[0007] A tire defect detection method provided by this application constructs a MAC module based on four parallel atrous convolution modules, uses the MAC module to replace the original pooling operation in the Unet model, and constructs a DMA Unet model to detect defects on the tire surface; atrous convolution can retain more detailed information. Compared with the original pooling layer in the Unet model, it can better process subtle crack features; because defects such as cracks in tire crack detection may be highly similar to the features of the tread, different atrous rates are set for the four-way atrous convolution in this application to expand the receptive field. The performance of defects such as cracks may vary at different scales. Therefore, through four-way parallel atrous convolution, the network can perform feature extraction simultaneously at multiple scales, which helps to detect tire surface defects of different sizes and shapes; the detection of defects such as cracks may be interfered by background noise. Four-way parallel atrous convolution can combine context information at multiple scales and improve the ability to identify cracks in complex backgrounds; after four-way atrous convolution in the MAC module, a 1*1 convolution is used once to integrate the obtained feature information and adjust the number of channels at the same time, effectively avoiding the loss of feature information caused by the pooling layer in the network structure, enhancing the feature expression ability, and ensuring that the model can identify defects on the tire surface with high efficiency and accuracy. Description of the Drawings
[0008] Figure 1 is a schematic structural diagram of the MAC module; Figure 2 is an architecture diagram of the DMA Unet network; Figure 3 is a comparison diagram of tire crack detection effects; Figure 4 is an application structure example of the tire defect detection model on an intelligent vehicle; Figure 5 is a diagram of the in-vehicle alarm prompt screen. Detailed Implementation Manner
[0009] This application includes a tire defect detection method, which includes the following steps.
[0010] S1: Construct a MAC (multiple addition convolution) module; as Figure 1 shown, the MAC module includes: an input module (represented as an input feature image in the figure), four parallel atrous convolution modules, a convolution module, and an output module (represented as an output feature image in the figure) connected in sequence.
[0011] In tire crack detection, the patterns of the tire itself pose a significant challenge to the detection accuracy. The characteristics of tire patterns are highly similar to those of cracks. A single porosity rate cannot effectively extract crack feature information, nor can it effectively distinguish defects such as cracks from tire tread patterns. In this application, the input feature image (InputFeature Map) is fed into four parallel atrous convolution modules simultaneously. The four parallel atrous convolution modules are set with different porosity rates to adjust the output channels and extract the feature information of the input feature map through different porosity rates.
[0012] The area of the input image that a certain layer of neurons in the convolutional neural network (or a certain point on the feature map) can "see" is called the receptive field. Using atrous convolution to increase the spacing of the convolutional kernel directly expands the receptive field. While not losing feature information, it captures more extensive context information.
[0013] While atrous convolution increases the receptive field, its computational cost is relatively small compared to traditional convolution. However, as the porosity rate increases, the network may lose some sensitivity to details because the core of atrous convolution may "skip" some important local information within the perception range. Therefore, in this application, a balance is made between the accuracy of crack detection and computational efficiency by choosing different atrous convolution rates. Specifically, the atrous convolution blocks Conv2d in the MAC module all use 3*3 convolutional kernels, and the porosity rates are set to 6, 12, 18, and 24 respectively. In this application, the atrous convolution module is set to 4 channels, which can avoid the loss of feature information caused by the pooling layer within an appropriate computational cost and meet the computational performance requirements of intelligent vehicles. The four atrous convolution modules are set in parallel, which can save computational time, improve computational efficiency, and enhance the real-time performance of this method.
[0014] Then, the feature information extracted by the four-way parallel atrous convolution is superimposed together based on the Concat operation and fed into a 1*1 convolutional module Conv2d to adjust the number of channels of the feature map to the number of channels required for the next layer. Finally, the output feature image (Output Feature Map) is obtained.
[0015] In tire defect detection, high-precision detection of small cracks is often required. Dilated convolution can retain more detailed information and preserve detailed features. That is, using dilated convolution in a tire defect detection model can help the network better capture the shape and structure of cracks and better handle subtle crack features. Cracks may appear differently at different scales. Through four-way parallel dilated convolution, the network can perform feature extraction simultaneously at multiple scales, which helps detect cracks of different sizes and shapes. In detection tasks, defects such as cracks may be interfered by background noise. Four-way parallel dilated convolution can combine context information at multiple scales, improve the ability to identify cracks in complex backgrounds, and enhance context understanding.
[0016] S2: Based on the Unet model, replace the original pooling operation in the Unet model with the MAC module to construct the DMA (direct multiple addition) Unet model.
[0017] Considering that multiple pooling layers in the Unet network can cause loss of features and information, which affects the network itself for detection tasks. Using the MAC Module to replace the pooling layer, the original design intention lies in the selection of the dilation rate of the four-way dilated convolution. Since the cracks in tire crack detection may be highly similar to the features of the tread, different dilation rates are selected to expand the receptive field, so as to more accurately extract the crack feature information that the network is intended to learn. After the four-way dilated convolution, a 1*1 convolution is used once to integrate the obtained feature information and adjust the number of channels at the same time. Although the computational cost will increase, the loss of feature information caused by the pooling layer is avoided in the network structure. Compared with the pooling operation, the MAC Module with 4-way dilated convolution can retain the spatial resolution, expand the receptive field, perform feature extraction at multiple scales at the same time, avoid the loss of feature information, and enhance the feature expression ability.
[0018] As Figure 2 shown, the DMA Unet model includes: an encoder and a decoder, which are connected by a bottleneck layer between the encoder and the decoder.
[0019] The encoder is used to extract features. The encoder includes N downsampling modules, and each downsampling module includes: a densely connected convolutional block, a convolutional layer, and a MAC Module connected in sequence.
[0020] The encoder contains multiple Dense Blocks. Each Dense Block has multiple convolutional layers. The convolutional kernel size for the convolution operation is 3*3, the stride is 1, and the padding method is same. In the figure, p is short for padding, s is short for stride, and k is short for kernel. The meaning of p=same is that after the convolution operation, the size of the feature map is padded to be the same as the input feature. Figure 1 After each Dense Block, the MAC Module is used to reduce the image size.
[0021] In this method, the dense connections implemented based on the Dense Block are introduced into the model, and then the feature maps obtained after dense connections at different levels in the encoder are skipped and linked to the decoder.
[0022] In this embodiment, the value of N is 3, that is, the encoder includes 3 downsampling modules. The first downsampling module includes: Dense Block1, Conv2d1, and MAC Module1 connected in sequence. The second downsampling module includes: DenseBlock2, Conv2d2, and MAC Module2 connected in sequence. The third downsampling module includes: Dense Block3, Conv2d3, and MAC Module3 connected in sequence.
[0023] The decoder uses a transposed convolutional layer (upsampling layer, UpSampling layer) to gradually restore the spatial size of the image, and at the same time uses skip connections to fuse the feature information from the encoder. The decoder includes N upsampling modules; each upsampling module includes: an UpSampling layer and a convolutional layer connected in sequence, and the UpSampling layer is implemented based on transposed convolution; where N≥1; a cross-layer skip connection is established between the convolutional layer in each level of the upsampling module and the convolutional layer in the corresponding level of the downsampling module. Specifically, when implementing, the UpSampling layer can be implemented based on the technologies that can perform upsampling in the prior art.
[0024] Specifically, the decoder includes 3 upsampling modules. The first upsampling module includes: UpSampling1 and Conv2d6. The second upsampling module includes: UpSampling2 and Conv2d7. The third upsampling module includes: UpSampling3 and Conv2d8. The feature map after being processed by the decoder is output after being processed by the convolution Conv2d9.
[0025] The bottleneck layer is implemented based on M consecutive convolutional layers, where M ≥ 1. The bottleneck layer is located between the encoder and the decoder. In this embodiment, the value of M is 2. Specifically, it consists of two 3*3 convolutional layers Conv2d4 and Conv2d5, with a stride of 1 and a padding method of same. The bottleneck layer is set after the encoder to further extract the feature information of the deepest level.
[0026] In the V-shaped network structure constructed based on the Unet model, the upsampling module 1 and the downsampling module 3 are at the same level, the upsampling module 2 and the downsampling module 2 are at the same level, and the upsampling module 3 and the downsampling module 1 are at the same level.
[0027] The decoder is used to restore the image size and perform the segmentation task. By using the UpSampling layer for upsampling, the detailed feature information in the corresponding layer in the encoder is skipped and transferred to the decoder, so as to improve the spatial resolution of the feature map, restore the size of the feature map, and then use the convolutional layer for feature extraction. After passing through the UpSampling layer, the corresponding layer in the decoder can be concatenated (Concat) with the output of the corresponding layer in the encoder due to the consistency of the feature map size, in order to restore the spatial information of the image.
[0028] Figure 2 For the three skip connections in, due to their same feature dimensions in the network output results and the inspiration of the Unet results, the convolutional output results after the dense convolutional blocks: Dense Block1, Dense Block 2, and Dense Block 3 in the encoder can be skip-linked to the deep convolutional layers Conv2d6, Conv2d7, and Conv2d8 in the decoder. Through this skip-link method, the network can obtain more detailed information, especially in terms of the edge information and structural details of the image. By combining dense links and skip links, each layer of the DenseUnet network can receive the features from both the encoder part and the features from the previous layers in the decoder at the same time, thus better fusing multi-scale features.
[0029] The original input image is fed into the encoder to gradually extract features through multiple convolutional layers and pooling layers, obtaining N-level feature maps. Each downsampling module in each layer uses the dense link method based on the dense convolutional block to transfer each level of feature map to the next layer, so that the input obtained by each layer in the encoder is not only the output of the previous layer, but the output of all network layers before the current level. The bottleneck layer is after the encoder and further extracts the feature information of the deepest level through M convolutional operations. The decoder uses the UpSampling layer to gradually restore the spatial size of the image, and at the same time uses skip links to fuse the feature information from the encoder.
[0030] The DMA Unet introduces the dense connection idea of DenseNet, passing the output of each layer to all subsequent network layers. That is, the input obtained by each layer is not only the output of the previous layer, but the output of all network layers before the current layer. This allows the feature map to accumulate more low-level and high-level features at each layer. At the same time, this dense connection method improves the information flow of the feature network, enabling each layer to benefit from deeper features and avoiding the problem of gradient disappearance, reducing a certain amount of parameter redundancy.
[0031] S3: Construct a tire defect detection model based on the DMA Unet model; As Figure 4 shown, the tire defect detection model includes a feature input module, a DMA Unet model, and a detection result output module connected in sequence; The DMA Unet model detects and enhances the defects in the input feature map to obtain an enhanced feature map. The output of the DMA Unet model includes: the enhanced feature map and the feature value corresponding to each pixel point in the enhanced feature map.
[0032] The defects in this application include: wear, rubber chunks missing, and cracks. The detection result output module judges the defect features existing in the enhanced feature map. The specific judgment method is: Compare the feature value corresponding to each pixel point with a preset defect judgment threshold. When the feature value of the pixel point is greater than the defect judgment threshold, it is judged that there is a defect at that place. In this embodiment, the defect judgment threshold is set to 0.5.
[0033] The specific display method of the feature value is as follows. In this embodiment, the feature values of the enhanced feature map with pixels of 4*3 (width * height) are displayed as follows: [0.2 0.3 0.3 0.2 0.5 0.6 0.7 0.6 0.4 0.3 0.2 0.4]; Compare the feature value of each pixel with the defect judgment threshold, and the defect judgment result is: there are defects at the points with pixel values of 0.6, 0.7, and 0.6.
[0034] In another example, the feature values of the enhanced feature map with pixels of 5*3 are displayed as follows: [0.2 0.3 0.4 0.3 0.2 0.5 0.6 0.5 0.6 0.7 0.3 0.2 0.4 0.3 0.2].
[0035] Compare the feature value of each pixel with the defect judgment threshold, and the defect judgment result is: there are defects at the points with pixel values of 0.6, 0.6, and 0.7.
[0036] In specific implementation, for better visualization effects, a marking module can also be set in the detection result output module. Once defective pixel points are detected, the positions of the defects can be highlighted and marked red on the enhanced feature map and then output after enhanced display, so as to obtain an enhanced feature map with marks.
[0037] S4: Based on various tire capture data, label the cracks on the tire to construct training sample data; Construct a training data set and a validation data set based on various training sample data; use the training data set to train the tire defect detection model to obtain a trained tire defect detection model. The specific construction methods of the training data set and the validation data set, as well as the training method of the model can be implemented based on the existing technology.
[0038] S5: Based on the image acquisition device, capture tire images in real time and send them into the trained tire defect detection model. The tire defect detection model outputs the detection results for each input image; The detection results include: defect judgment results, enhanced feature maps, and the feature values corresponding to each pixel point in the enhanced feature maps.
[0039] This method is used to detect the defects on the tire surface based on image recognition and machine learning models. Only the trained model needs to be saved in the processor, and the images captured by the image acquisition device are sent into the model to complete the detection. There is no need to specifically manufacture exclusive detection equipment. Compared with the existing physical detection methods, the implementation efficiency is high and the cost is relatively low. It can be applied to various scenarios such as tire defect detection in repair shops or factories.
[0040] Especially worth mentioning is that this method has good real-time performance and can be applied to intelligent vehicles to detect vehicle tires in real time.
[0041] Such as Figure 4 As shown, the image acquisition device uses the surround-view camera in the vehicle-mounted camera, or installs an image acquisition camera specifically for each tire. The trained tire defect detection model is preset in the vehicle-mounted system to receive the images including the tire surface sent by the vehicle-mounted camera in real time and detect the defects on the tire surface in real time. An image clarity determination module is set in the feature input module to determine the timing of receiving images according to the real-time vehicle speed. For example, when the vehicle speed is lower than 5 km / h, the camera can capture clear and available tire surface images.
[0042] In actual application, a defect determination condition is also set in the detection result output module to avoid the interference of noise points in the image on the judgment. For example, the defect determination condition is set as follows: the number of pixel points of the defective image is greater than or equal to 8, and the defective pixel points are continuous, that is, any defective pixel point has an adjacent relationship with at least one other defective pixel point. When all defective pixel points are successively and continuously adjacent, it indicates the existence of a thin linear crack; when the defective pixel points are not successively and continuously adjacent, it indicates the existence of a defect of a non-linear crack, such as a crack with a complex shape, a block defect or a sheet wear. If the defect determination condition is met, it is determined that there is a defect in the tire position corresponding to the original input image, and the detection result is output, and the determination result is also output for the original input image. The determination results include: there is a defect and there is no defect.
[0043] For the tire surface image with tire defects, the output of the corresponding tire defect detection model includes: the determination result, the enhanced feature map with marks, and the feature value corresponding to each pixel point in the enhanced feature map.
[0044] According to the tire tread, determine whether the image is available. If the image is determined to be available, the image captured by the camera with the tire surface is sent to the model for detection. Once the determination result is that there are tire defects in the input image, a defect notification is sent to the intelligent vehicle system. The intelligent vehicle system can, according to the determination result, remind the driver to pay attention to defects such as cracks in the tire through text and voice warnings. For the specific prompt screen, refer to the appendix Figure 5 . Through the real-time defect detection of the tire surface by this method, the probability of vehicle accidents caused by tire defects can be effectively reduced. The specific installation position of the camera and the image processing method can be realized based on the existing technology.
[0045] In order to verify the effectiveness of this method for tire surface defect detection, the recognition results of the DMA Unet model in this application are compared with the detection results of other image detection models. The comparison methods include: Canny operator, Sobel operator, SegNet model, Mask R-CNN model.
[0046] As Figure 3 shown, the first column is 3 original input images, the second column is the enhanced feature maps respectively output by the DMA Unet model in this application for the 3 original input images, and then the enhanced feature maps output by other comparison models.
[0047] The first set of horizontal detection results are the enhanced feature maps corresponding to the detection results output by each method for the first original input image. From the five detection result images in the first horizontal group, it can be seen that compared with the other four, DMA-Unet extracts richer crack details and the detected crack features are more obvious. From the second horizontal group of detection result images, it can be seen that although DMA-Unet's performance in detecting cracks on the tire cross-section is slightly inferior to CANNY and SOBEL, its detection effect has less noise than SegNet and has stronger anti-interference ability for vertical stripes on the tire cross-section. On the tire side, especially for fine and dense cracks, the detection ability of DMA-Unet is significantly better than CANNY and SOBEL, and it further demonstrates its anti-interference ability compared with SegNet. Generally speaking, DMA-Unet can have higher detection accuracy than other algorithms in the face of complex and fine and dense cracks.
[0048] After using the technical solution of the present invention, by introducing a four-way parallel atrous convolution module into the Unet model in this application, a DMA Unet model can be constructed to achieve a balance between detail and macroscopic information. At a lower atrous convolution rate (i.e., a smaller spacing between convolution kernels), the network can usually better capture the detailed information of the image, such as the fine parts of cracks. While a higher atrous convolution rate helps to capture more macroscopic feature information in the image, such as the long-term trend of cracks or large-scale damage areas. Therefore, different atrous convolution rates will affect the degree of attention of the network to detail and macroscopic information during crack detection. In the DMA Unet model of this method, the atrous rates of the four-way atrous convolution are set to 6, 12, 18, and 24 respectively. At the same time, setting different atrous rates for the four-way atrous convolution module can also extract local features and global features simultaneously; the convolution with a lower atrous rate is more inclined to extract local features, while the convolution with a higher atrous rate helps to extract global features. In the crack detection task, the convolution with a low atrous rate is more suitable for detecting local cracks; but if the cracks show a long-distance distribution or there are larger damaged areas, the convolution with a high atrous rate may perform better because it can capture more global information.
Claims
1. A method for detecting tire defects, characterized in that, It includes the following steps: S1: Construct a MAC module; The MAC module includes: an input module, four parallel atrous convolution modules, a convolution module, and an output module that are connected in sequence; The four parallel atrous convolution modules are respectively set with dilation rates. The input feature map is simultaneously fed into the four parallel atrous convolution modules. The output channels are adjusted through different dilation rates and the feature information of the input feature map is extracted. Then, the feature information extracted by the four-way parallel atrous convolution is stacked together based on the Concat operation, and then fed into a 1*1 convolution module to adjust the number of channels of the feature map. Finally, the feature map is output based on the output module; S2: Based on the Unet model, replace the original pooling operation in the Unet model with the MAC module to construct a DMA Unet model; S3: Construct a tire defect detection model based on the DMA Unet model; The tire defect detection model includes: a feature input module, a DMA Unet model, and a detection result output module that are connected in sequence; The DMA Unet model detects and enhances the defects in the input feature map to obtain an enhanced feature map. The output of the DMA Unet model includes: the enhanced feature map and the feature value corresponding to each pixel point in the enhanced feature map; The detection result output module judges the defect features existing in the enhanced feature map to determine whether there are defects; S4: Construct a training data set and a validation data set based on various tire capture data; Use the training data set to train the tire defect detection model to obtain the trained tire defect detection model; S5: Based on an image acquisition device, a tire image is captured in real time and fed into the trained tire defect detection model. The tire defect detection model outputs the detection result for each input image; The detection result includes: a defect judgment result, an enhanced feature map, and the feature value corresponding to each pixel point in the enhanced feature map.
2. The method for detecting tire defects according to claim 1, characterized in that: The DMA Unet model includes: an encoder and a decoder, and the encoder and the decoder are connected through a bottleneck layer; The encoder includes N downsampling modules. Each downsampling module includes: a dense convolution block, a convolution layer, and a MAC Module connected in sequence; the decoder includes N upsampling modules. Each upsampling module includes: an UpSampling layer and a convolution layer connected in sequence. The UpSampling layer is implemented based on transposed convolution; where N≥1; a cross-layer skip connection is established between the convolution layer in the same-level upsampling module and the convolution layer in the downsampling module; The bottleneck layer is implemented based on M consecutive convolution layers, M≥1; The original input image is fed into the encoder to gradually extract features through multiple convolutional layers and pooling layers, obtaining N-level feature maps. At the same time, each downsampling module in each layer uses a densely connected manner based on the dense convolutional block to transfer each level of feature map to the next layer, so that the input obtained by each layer in the encoder is not only the output of the previous layer, but the output of all network layers before the current layer; the bottleneck layer is after the encoder and further extracts the feature information of the deepest layer through M convolutional operations; the decoder uses the UpSampling layer to gradually restore the spatial size of the image and uses skip connections to fuse the feature information from the encoder.
3. The method for detecting tire defects according to claim 1, characterized in that: In step S5, the image acquisition device is implemented based on an in-vehicle camera. The trained tire defect detection model is preset in the intelligent vehicle system. The feature input module includes: an image clarity determination module; in the image clarity determination module, the timing of receiving the image transmitted by the in-vehicle camera is determined according to the real-time vehicle speed, and whether the received input image is available is determined according to the clarity of the tire tread; if the image is determined to be available, the image captured by the camera with the tire surface is sent into the tire defect detection model in real time to detect the defects on the tire surface in real time. The detection result output module also includes a defect determination condition, and the defect determination condition is: the defective pixel points are continuous and the number of pixel points is greater than a preset continuous threshold; when the defect determination condition is met, it is determined that there are defects on the tire surface corresponding to the input image.
4. The method for detecting tire defects according to claim 1, characterized in that: The dilation rates of the four dilated convolutional modules in the MAC module are set to 6, 12, 18, and 24 respectively.
5. The method for detecting tire defects according to claim 2, characterized in that: The value of N in the decoder and the encoder is 3; the value of M in the bottleneck layer is 2.
6. The method for detecting tire defects according to claim 1, characterized in that: The defect judgment threshold is set to 0.
5.
7. The method for detecting tire defects according to claim 1, characterized in that: The defects include: wear, rubber chunks missing, and cracks.
8. The method for detecting tire defects according to claim 1, characterized in that: The detection result output module also includes: a marking module. Once defective pixel points are detected, the location of the defect will be highlighted in the enhanced feature map.
9. The method for detecting tire defects according to claim 1, characterized in that: In step S3, the detection result output module judges the defect features existing in the enhanced feature map. The specific judgment method is as follows: The feature value corresponding to each pixel point is compared with a preset defect judgment threshold. When the feature value of the pixel point is greater than the defect judgment threshold, it is determined that the pixel point has a defect.
Citation Information
Patent Citations
Image recognition method, device, storage medium and system
CN115908173A
Automobile tire tread pattern defect detection system and method
CN115908304A
Toll vehicle type identification method and system based on expressway passing image
CN119992483A