Defect detection method and system for engineering signboard based on deep learning
By constructing a deep learning-based defect detection network model, the problems of low efficiency and low accuracy of traditional manual inspection methods have been solved, enabling efficient automatic detection of tilted installations and surface defects on engineering signs, thus improving detection accuracy and adaptability.
Patent Information
- Application Number
- CN202510915137.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Traditional manual inspection methods are inefficient and inaccurate in the installation of engineering signs and the detection of surface defects, and it is difficult to detect installation tilt or surface defects in a timely manner.
A deep learning-based defect detection method is adopted. By constructing a defect detection network model, including an initial feature extraction module, a feature segmentation module, a feature enhancement module, a feature fusion module, an attention module, and a multi-scale fusion module, the method performs preprocessing, feature extraction, segmentation, enhancement, and fusion on engineering sign images to achieve automatic detection.
It improves the efficiency and accuracy of defect detection, enhances the ability to identify defects, and improves the model's adaptability to complex scenarios and detection accuracy.
Smart Images

Figure CN120894282A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of defect detection, and particularly relates to a defect detection method and system for engineering signboards based on deep learning BACKGROUND With the continuous development of traffic, road and construction engineering, engineering signboards play an important role in traffic safety and engineering management. Ensuring the correct installation and integrity of signboards is a key link to ensure safety and improve management efficiency; however, traditional manual inspection methods have the disadvantages of labor-intensive, low efficiency, strong subjectivity, etc., and it is difficult to find installation inclination or surface defects in time.
[0002] In recent years, in order to improve the detection efficiency and accuracy, deep learning technology has gradually become a research hotspot; using computer vision, image processing and machine learning technologies, automatic detection of the inclination state and surface defects of engineering signboards after installation has become an important direction of industry development. SUMMARY
[0003] The purpose of the embodiments of the application is to provide a defect detection method and system for engineering signboards based on deep learning, which can solve the technical problems of low efficiency and low accuracy of traditional manual inspection methods in the prior art.
[0004] In order to solve the above technical problems, the application is implemented as follows: In a first aspect, the embodiments of the application provide a defect detection method for engineering signboards based on deep learning, which comprises: Obtaining a plurality of engineering signboard images to construct an engineering signboard dataset according to the plurality of engineering signboard images, and pre-processing the engineering signboard dataset to obtain a target dataset; Constructing a defect detection network model, wherein the defect detection network model comprises an initial feature extraction module, a feature segmentation module, a feature enhancement module, a feature fusion module, an attention module and a multi-scale fusion module; Training the defect detection network model according to part of the data in the target dataset, and testing the defect detection network model according to another part of the data; Obtaining an engineering signboard image to be detected, processing the engineering signboard image to be detected according to the tested defect detection network model, and obtaining a defect detection result of the engineering signboard image to be detected.
[0005] As an optional implementation manner of the first aspect of the application, the pre-processing of the engineering signboard dataset to obtain a target dataset is specifically: Resizing each of the engineering signboard images in the engineering signboard dataset to obtain a corresponding standard image for each of the engineering signboard images; Performing noise reduction and smoothing processing on each of the standard images to obtain each of the smooth images corresponding to each of the standard images; Performing enhancement and sharpening processing on each of the smooth images to obtain each of the sharpened images corresponding to each of the smooth images; Performing cropping on the areas other than the engineering signboard and the ground in each of the sharpened images to obtain each of the target images corresponding to each of the sharpened images, and constructing the target dataset according to each of the target images. As an optional implementation form of the first aspect of the present application, the processing of the to-be-detected engineering signboard image according to the defect detection network model after the test to obtain the defect detection result of the to-be-detected engineering signboard image is specifically: Performing initial feature extraction processing on the to-be-detected engineering signboard image according to the initial feature extraction module to obtain an initial feature map corresponding to the to-be-detected engineering signboard image; Performing segmentation processing on the initial feature map according to the feature segmentation module to obtain a signboard feature map and a ground feature map; Performing enhancement on the signboard feature map and the ground feature map according to the feature enhancement module to obtain a signboard enhanced feature map and a ground enhanced feature map; Performing fusion processing on the signboard enhanced feature map, the ground enhanced feature map, and the initial feature map according to the feature fusion module to obtain a first fusion feature map; Performing processing on the first fusion feature map according to the attention module to obtain a fusion attention feature map and a signboard attention feature map; Performing multi-scale fusion processing on the signboard attention feature map according to the multi-scale fusion module to obtain a signboard multi-scale fusion feature map, and performing fusion on the signboard enhanced feature map, the signboard multi-scale fusion feature map, and the initial feature map to obtain a second fusion feature map; Obtaining a signboard installation defect result of the to-be-detected engineering signboard image according to the first fusion feature, and obtaining a signboard damage defect result of the to-be-detected engineering signboard image according to the second fusion feature. As an optional implementation form of the first aspect of the present application, the segmentation processing on the initial feature map according to the feature segmentation module to obtain a signboard feature map and a ground feature map is specifically: Performing edge detection on the initial feature map to obtain a signboard edge feature map and a ground edge feature map in the initial feature map, and performing pixel-level cropping on the initial feature map to obtain a plurality of pixel feature maps with the same size; Smoothly connect the center in the pixel feature map containing the signboard edge to obtain a signboard pixel feature map, and smoothly connect the center in the pixel feature map containing the ground edge to obtain a ground pixel feature map; Fuse the signboard pixel feature map and the signboard edge feature map to obtain a signboard target edge feature map, and fuse the ground pixel feature map and the ground edge feature map to obtain a ground target edge feature map. Segment the initial feature map according to the ground target edge feature map and the signboard target edge feature map to obtain the signboard feature map and the ground feature map. As an optional implementation of the first aspect of the application, the signboard feature map and the ground feature map are enhanced according to the feature enhancement module to obtain a signboard enhanced feature map and a ground enhanced feature map; specifically: The signboard feature map and the ground feature map are respectively subjected to global maximum pooling processing to obtain a signboard pooling feature map and a ground pooling feature map. The signboard pooling feature map and the ground pooling feature map are respectively subjected to 1×1 convolution processing to obtain a signboard spatial feature map and a ground spatial feature map. The signboard spatial feature map and the ground spatial feature map are respectively subjected to 3×3 depth separable convolution processing to obtain a signboard texture feature map and a ground texture feature map. The signboard spatial feature map and the ground spatial feature map are respectively subjected to 1×1 shallow convolution processing to obtain a signboard color feature map and a ground color feature map. The signboard spatial feature map, the signboard texture feature map, and the signboard color feature map, and the ground spatial feature map, the ground texture feature map, and the ground color feature map are respectively subjected to batch normalization and activation processing in sequence. The activated signboard spatial feature map, the signboard texture feature map, and the signboard color feature map are respectively subjected to residual connection, and the activated ground spatial feature map, the ground texture feature map, and the ground color feature map are respectively subjected to residual connection to obtain a signboard enhanced feature map and a ground enhanced feature map. As an optional implementation of the first aspect of the application, the first fusion feature map is processed according to the attention module to obtain a fusion attention feature map and a signboard attention feature map; specifically: The first fusion feature map is subjected to feature segmentation to obtain a signboard fusion feature map, and the signboard fusion feature map is subjected to global average pooling processing to obtain a fusion pooling feature map. The global fusion feature map and the local fusion feature map are subjected to batch normalization and activation processing to obtain a global activation feature map and a local activation feature map. The global activation feature map and the local activation feature map are subjected to residual connection according to a residual attention mechanism to obtain the fusion attention feature map. The global activation feature map and the local activation feature map are subjected to residual connection according to a residual attention mechanism to obtain the fusion attention feature map. The fusion attention feature map is subjected to 3x3 convolution, batch normalization, activation and maximum pooling processing in sequence to obtain the signboard attention feature map. As an optional implementation of the first aspect of the application, the multi-scale fusion module performs multi-scale fusion processing on the signboard attention feature map to obtain a signboard multi-scale fusion feature map, specifically: The signboard attention feature map is subjected to multiple down-sampling to obtain multiple signboard scale feature maps of different scales; The multiple signboard scale feature maps are randomly divided into two groups to obtain a first scale feature set and a second scale feature set, respectively; The first scale feature set is subjected to cross fusion according to the second scale feature set to obtain each cross scale feature map corresponding to each signboard scale feature map in the first scale feature set; Each cross scale feature map is subjected to multi-scale feature fusion processing to obtain a cross fusion feature map; Each signboard scale feature map in the second scale feature set is subjected to strengthening according to the cross fusion feature map to obtain each strengthened scale feature map corresponding to each signboard scale feature map in the second scale feature set; Each strengthened scale feature map is subjected to multi-scale fusion processing to obtain a strengthened fusion feature map, and the strengthened fusion feature map is spliced with the cross fusion feature map to obtain the signboard multi-scale fusion feature map. In a second aspect, the embodiments of the application provide an engineering signboard defect detection system based on deep learning, which comprises: An acquisition module acquires multiple engineering signboard images to construct an engineering signboard dataset according to the multiple engineering signboard images, and pre-processes the engineering signboard dataset to obtain a target dataset; A construction module constructs a defect detection network model, wherein the defect detection network model comprises an initial feature extraction module, a feature segmentation module, a feature enhancement module, a feature fusion module, an attention module and a multi-scale fusion module. The training module trains the defect detection network model according to part of data in the target data set, and tests the defect detection network model according to another part of data; The prediction module acquires an image of an engineering sign to be detected, processes the image of the engineering sign to be detected according to the tested defect detection network model, and obtains a defect detection result of the image of the engineering sign to be detected.
[0006] In a third aspect, an electronic device is provided, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, and the program or instructions are executed by the processor to implement the steps of the method according to the first aspect.
[0007] In a fourth aspect, a readable storage medium is provided, and the readable storage medium stores a program or instructions, and the program or instructions are executed by a processor to implement the steps of the method according to the first aspect.
[0008] In the embodiments of the present application, compared with the prior art, the following beneficial effects are achieved: (1) The initial feature map is segmented by the feature segmentation module to segment out the sign feature and the ground feature, and then the feature enhancement module is used to enhance the sign feature map and the ground feature map respectively, and the separate enhancement helps to highlight the key features and important features in the sign feature map and the ground feature map; (2) The sign enhanced feature map, the ground enhanced feature map, and the initial feature map are fused by the feature fusion module, so that the fused feature map has both the initial feature and the region enhanced feature, thereby enhancing the defect recognition ability of the defect detection network model; (3) The attention module is used to globally capture the sign information of the fused attention feature map and locally enhance the edge positioning, and the double branches are cooperated to improve the adaptability of the defect detection network model to complex scenes; (4) The multi-scale fusion module is used to cross-fuse and enhance the sign scale feature maps of different scales, to enhance the feature expression ability of the sign scale feature maps of different scales, so that the defect detection network model can accurately capture the defects, and the detection precision of the defect detection network model is enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0009] Fig. 1 is a flowchart of a defect detection method for an engineering sign based on deep learning provided by some embodiments of the present application; Fig. 2 is a structure diagram of a defect detection network model in a defect detection method for an engineering sign based on deep learning provided by some embodiments of the present application. DETAILED DESCRIPTION
[0010] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present application.
[0011] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in a "or" relationship.
[0012] The deep learning-based engineering sign defect detection method and system provided by the embodiments of the present application will be described in detail below in conjunction with the drawings and specific embodiments and application scenarios.
[0013] Embodiments A deep learning-based engineering sign defect detection method includes the following steps: S100: Obtain a plurality of engineering sign images to construct an engineering sign dataset according to the plurality of engineering sign images, and pre-process the engineering sign dataset to obtain a target dataset; It should be understood that a plurality of engineering sign images are first obtained to construct an engineering sign dataset according to the plurality of engineering sign images, and the engineering sign dataset contains engineering sign images of different light, angle, and defect types, so that the subsequent model can be better trained, and then the engineering sign dataset is pre-processed to obtain a target dataset.
[0014] Further, the engineering sign dataset in S100 is pre-processed to obtain a target dataset, specifically: S110: Adjust the size of each engineering sign image in the engineering sign dataset to obtain each corresponding standard image; S120: Perform noise reduction and smoothing processing on each standard image to obtain each corresponding smooth image of each standard image; S130: Perform enhancement and sharpening processing on each smooth image to obtain each corresponding sharpened image of each smooth image; S140: cropping the region other than the engineering signboard and the ground in each sharpened image to obtain each target image corresponding to each sharpened image, and constructing a target data set according to each target image.
[0015] It needs to be understood that first, the size of each engineering signboard image in the engineering signboard data set is adjusted to obtain each corresponding standard image, ensuring that the spatial scales of all standard images are consistent, thereby facilitating stable learning in the subsequent model training process and avoiding the influence of scale difference on model learning; then each standard image is denoised and smoothed to obtain each smooth image corresponding to each standard image; by smoothing and denoising the standard image, the noise and irregular details in the standard image can be reduced, thereby obtaining a smooth image that can highlight key edge and structural features; then each smooth image is enhanced and sharpened to obtain each sharpened image corresponding to each smooth image; through sharpening and enhancement processing, the details in the smooth image can be more prominent; finally, the region other than the engineering signboard and the ground in each sharpened image is cropped to obtain each target image corresponding to each sharpened image, and a target data set is constructed according to each target image. By recognizing the engineering signboard region and the ground region in the sharpened image, and then according to the edge information and color information of the recognized engineering signboard region and ground region, the other regions in the sharpened image are cropped to obtain target images that only retain the engineering signboard and the ground, and then a target data set is constructed according to all target images; by cropping the other regions in the sharpened image, the interference information is removed, which helps to improve the efficiency of the subsequent model.
[0016] S200: constructing a defect detection network model, the defect detection network model including an initial feature extraction module, a feature segmentation module, a feature enhancement module, a feature fusion module, an attention module, and a multi-scale fusion module; S300: training the defect detection network model according to part of the data in the target data set, and testing the defect detection network model according to another part of the data; It needs to be understood that first, the target images in the target data set are divided into a training set and a test set according to a ratio of 3:1, and then a loss function is constructed, the defect detection network model is trained for multiple rounds according to the constructed loss function and the training set, the loss function continuously adjusts the parameters of the model during the training process, so that after multiple rounds of training, the parameters of the model are optimized to convergence, and then the training of the defect detection network model is stopped, the trained defect detection network model is tested through the test set, and the performance of the defect detection network model is verified according to the test result, and if the verification result is qualified, the model training and testing are completed.
[0017] S400: Obtain the to-be-detected engineering signboard image, and process the to-be-detected engineering signboard image according to the tested defect detection network model to obtain a defect detection result of the to-be-detected engineering signboard image.
[0018] It should be understood that, first, the to-be-detected engineering signboard image is obtained, and the to-be-detected engineering signboard image is preprocessed through the preprocessing steps S110-S140; and then the preprocessed to-be-detected engineering signboard image is processed according to the tested defect detection network model, so as to obtain that the engineering signboard in the to-be-detected engineering signboard image has the tilt defect and the surface damage defect.
[0019] Further, in S400, the to-be-detected engineering signboard image is processed according to the tested defect detection network model to obtain a defect detection result of the to-be-detected engineering signboard image, specifically: S410: Perform initial feature extraction processing on the to-be-detected engineering signboard image according to the initial feature extraction module to obtain an initial feature map corresponding to the to-be-detected engineering signboard image; S420: Perform segmentation processing on the initial feature map according to the feature segmentation module to obtain a signboard feature map and a ground feature map; S430: Perform enhancement on the signboard feature map and the ground feature map according to the feature enhancement module to obtain a signboard enhanced feature map and a ground enhanced feature map; S440: Perform fusion processing on the signboard enhanced feature map, the ground enhanced feature map, and the initial feature map according to the feature fusion module to obtain a first fusion feature map; S450: Process the first fusion feature map according to the attention module to obtain a signboard attention feature map; S460: Perform multi-scale fusion processing on the signboard attention feature map according to the multi-scale fusion module to obtain a signboard multi-scale fusion feature map, and fuse the signboard enhanced feature map, the signboard multi-scale fusion feature map, and the initial feature map to obtain a second fusion feature map; S470: Obtain a signboard tilt defect result of the to-be-detected engineering signboard image according to the first fusion feature, and obtain a signboard surface damage defect result of the to-be-detected engineering signboard image according to the second fusion feature.
[0020] It needs to be understood that first, the initial feature extraction module extracts the initial feature map of the to-be-detected engineering sign image through the convolution layer, and the extraction of the initial feature map provides basic information for subsequent processing; second, the feature segmentation module performs segmentation processing on the initial feature map to obtain a sign feature map and a ground feature map; by segmenting out the sign features and ground features in the initial feature map, it is helpful for the feature enhancement module to separately enhance the sign feature map and the ground feature map; then, the feature enhancement module enhances the sign feature map and the ground feature map to obtain a sign enhanced feature map and a ground enhanced feature map; by separately enhancing the sign feature map and the ground feature map, it is helpful to highlight the key features and important features in the sign feature map and the ground feature map, thereby suppressing the redundant information in the sign feature map and the ground feature map, so that the defect detection network model can enhance the perception ability of the defect features in the sign enhanced feature map and the ground enhanced feature map; then, the feature fusion module fuses the sign enhanced feature map, the ground enhanced feature map and the initial feature map through feature splicing to obtain a first fusion feature map; by fusing the sign enhanced feature map, the ground enhanced feature map and the initial feature map, the first fusion feature map retains the initial feature and increases the sign enhanced feature and the ground enhanced feature map, thereby enhancing the recognition ability of the defect detection network model to defects; the attention module processes the first fusion feature map to obtain a sign attention feature map; the attention module can automatically adjust the importance of each region in the first fusion feature map, so that the defect detection network model can increase the attention to potential defect information; then, the multi-scale fusion module performs multi-scale fusion processing on the sign attention feature map to obtain a sign multi-scale fusion feature map, and fuses the sign enhanced feature map, the sign multi-scale fusion feature map and the initial feature map to obtain a second fusion feature map; multi-scale fusion processing on the sign attention feature map can capture different scale defect information, and fusing the sign enhanced feature map, the sign multi-scale fusion feature map and the initial feature map can significantly improve the expression of the features; finally, the tilt classifier determines whether the sign of the to-be-detected engineering sign image is tilted according to the first fusion feature, and the surface damage classifier determines whether the surface of the sign of the to-be-detected engineering sign image has a damage defect according to the second fusion feature.
[0021] Further, in S420, the initial feature map is segmented by the feature segmentation module to obtain a sign feature map and a ground feature map; specifically: S421: performing edge detection on the initial feature map to obtain a sign edge feature map and a ground edge feature map in the initial feature map, and performing pixel-level cropping on the initial feature map to obtain a plurality of pixel feature maps with the same size; S422: Smoothly connect the center of the pixel feature map containing the signboard edge to obtain a signboard pixel feature map, and smoothly connect the center of the pixel feature map containing the ground edge to obtain a ground pixel feature map; S423: Fuse the signboard pixel feature map and the signboard edge feature map to obtain a signboard target edge feature map, and fuse the ground pixel feature map and the ground edge feature map to obtain a ground target edge feature map; S424: Segment the initial feature map according to the ground target edge feature map and the signboard target edge feature map to obtain a signboard feature map and a ground feature map.
[0022] Specifically, the signboard feature map and the ground feature map are represented by the following formula: wherein, signboard feature map is represented by, ground feature map is represented by, initial feature map is represented by element-wise multiplication is represented by, morphological dilation operation is represented by, signboard target edge feature map is represented by, ground target edge feature map is represented by, and learnable weight is represented by, activation function is represented by, signboard pixel feature map is represented by, ground pixel feature map is represented by, signboard edge feature map is represented by, ground edge feature map is represented by, morphological closing operation is represented by, a set of all pixel feature maps containing signboard edges is represented by, a set of all pixel feature maps containing ground edges is represented by.
[0023] It needs to be understood that first, the initial feature map is subjected to edge detection by using the Canny operator to obtain the signboard edge feature map and the ground edge feature map in the initial feature map, and the initial feature map is subjected to grid pixel-level cropping according to a fixed size to obtain a plurality of pixel feature maps of the same size; through edge detection, the edge information of the signboard and the ground in the initial feature map is extracted, and the regions in the initial feature map are divided according to the edge information of the signboard and the ground, so as to obtain the signboard edge feature map and the ground edge feature map. The edge detection can ensure that the contours of the signboard and the ground are clear, provide accurate area boundaries for subsequent pixel set cropping, and improve the segmentation accuracy; secondly, the center of the pixel feature map containing the signboard edge is smoothly connected by using morphological closing operation to obtain the signboard pixel feature map, and the center of the pixel feature map containing the ground edge is smoothly connected to obtain the ground pixel feature map; through the smooth connection, the obtained signboard pixel feature map and ground pixel feature map have complete contours, so as to effectively make up for the discontinuity problem in edge detection; then, the signboard pixel feature map and the signboard edge feature map are fused according to channel weighting to obtain the signboard target edge feature map, and the ground pixel feature map and the ground edge feature map are fused to obtain the ground target edge feature map; by fusing the signboard pixel feature map and the signboard edge feature map, the boundary contour features of the signboard can be more prominent, and by fusing the ground pixel feature map and the ground edge feature map, the boundary contour features of the ground can be more prominent; finally, morphological dilation operation is used to segment the initial feature map according to the ground target edge feature map and the target edge feature map to obtain the signboard feature map and the ground feature map; by segmenting the initial feature map according to the ground target edge feature map and the signboard target edge feature map with significant edge features, the segmentation accuracy in the feature segmentation process can be increased, so that more accurate signboard feature map and ground feature map are obtained.
[0024] Further, the signboard feature map and the ground feature map are enhanced according to the feature enhancement module in S430 to obtain the signboard enhanced feature map and the ground enhanced feature map; specifically: S431: the signboard feature map and the ground feature map are subjected to global maximum pooling processing respectively to obtain the signboard pooling feature map and the ground pooling feature map; S432: the signboard pooling feature map and the ground pooling feature map are subjected to 1x1 convolution processing respectively to obtain the signboard spatial feature map and the ground spatial feature map respectively; S433: the signboard spatial feature map and the ground spatial feature map are subjected to 3x3 depth separable convolution processing respectively to obtain the signboard texture feature map and the ground texture feature map respectively; S434: 1x1 shallow convolution processing is respectively performed on the sign space feature map and the ground space feature map, and sign color feature map and ground color feature map are respectively obtained; S435: batch normalization and activation processing are sequentially performed on the sign space feature map, the sign texture feature map and the sign color feature map, and the ground space feature map, the ground texture feature map and the ground color feature map, respectively; S436: residual connection is respectively performed on the activated sign space feature map, the sign texture feature map and the sign color feature map, and residual connection is performed on the activated ground space feature map, the ground texture feature map and the ground color feature map, and sign enhanced feature map and ground enhanced feature map are respectively obtained.
[0025] Further, the sign enhanced feature map and the ground enhanced feature map are represented by the following formula: , Wherein, sign enhanced feature map is represented by, ground enhanced feature map is represented by, sign space feature map is represented by, sign texture feature map is represented by, sign color feature map is represented by, ground space feature map is represented by, ground texture feature map is represented by, ground color feature map is represented by, 1x1 convolution is represented by, element-wise addition is represented by, and learning weight is represented by.
[0026] It needs to be understood that first, global max pooling processing is performed on the signboard feature map and the ground feature map respectively to obtain a signboard pooling feature map and a ground pooling feature map; the global max pooling processing can capture the most significant information in the signboard feature map and the ground feature map, and strengthen the global expression; then, 1x1 convolution processing is performed on the signboard pooling feature map and the ground pooling feature map respectively to obtain a signboard spatial feature map and a ground spatial feature map; the channel dimension of the signboard pooling feature map and the ground pooling feature map is adjusted through 1x1 convolution, cross-channel information is fused, and compact signboard spatial feature map and ground spatial feature map are generated; secondly, 3x3 depth separable convolution processing is performed on the signboard spatial feature map and the ground spatial feature map respectively to obtain signboard texture features and ground texture features; through 3x3 depth separable convolution, it is decomposed into channel-by-channel convolution and point-by-point convolution, while more texture features are extracted, the calculation amount is greatly reduced, and thus the processing speed of the defect detection network model is improved; then, 1x1 shallow convolution processing is performed on the signboard spatial feature map and the ground spatial feature map respectively to obtain a signboard color feature map and a ground color feature map; through 1x1 shallow convolution, global color distribution features can be extracted to obtain a signboard color feature map and a ground color feature map; the parameter amount of 1x1 shallow convolution is low, and the color information in the signboard spatial feature map and the ground spatial feature map can be extracted while the calculation time is short; then, the obtained signboard spatial feature map, signboard texture feature map and signboard color feature map, and ground spatial feature map, ground texture feature map and ground color feature map are sequentially subjected to batch normalization and activation processing; through batch normalization processing and activation processing, convergence can be accelerated and gradient explosion can be relieved; finally, the activated signboard spatial feature map, signboard texture feature map and signboard color feature map are connected by residual connection, and the activated ground spatial feature map, ground texture feature map and ground color feature map are connected by residual connection to obtain a signboard enhanced feature map and a ground enhanced feature map; through residual connection, the original spatial feature map is added to the enhanced color feature map and texture feature map, which can suppress the problem of gradient vanishing while retaining low-level details; the feature enhancement module balances the demand for detection accuracy and efficiency in the defect detection process through the design of multi-branch feature enhancement combined with residual retention, and ensures the accuracy and efficiency in real-time detection.
[0027] Further, in S450, the first fusion feature map is processed according to the attention module to obtain a signboard attention feature map; specifically: S451: performing feature segmentation on the first fusion feature map to obtain a signboard fusion feature map, and performing global average pooling processing on the signboard fusion feature map to obtain a fusion pooling feature map; S452: Global attention processing and local attention processing are performed on the fusion attention feature map respectively, and global fusion feature map and local fusion feature map are obtained respectively; S453: Batch normalization and activation processing are performed on the global fusion feature map and the local fusion feature map, and global activation feature map and local activation feature map are obtained; S454: The global activation feature map and the local activation feature map are connected by residual attention mechanism, and the fusion attention feature map is obtained; S455: The fusion attention feature map is sequentially processed by 3x3 convolution, batch normalization, activation and max pooling, and the signboard attention feature map is obtained.
[0028] Specifically, the signboard attention feature map is represented by the following formula: , Wherein, represents the signboard attention feature map, represents the max pooling processing with a step of 2, represents the activation function, represents the batch normalization processing, represents the 3x3 convolution, represents the fusion attention feature map, represents the convolution kernel, represents the offset of the convolution kernel in the row direction, represents the offset of the convolution kernel in the column direction, represents the row index, represents the column index, represents the global activation feature map, represents the local activation feature map, represents the signboard fusion feature map, , and represent the learnable weights.
[0029] It needs to be understood that first, the first fusion feature map is subjected to feature segmentation to obtain a signboard fusion feature map, and the signboard fusion feature map is subjected to global average pooling processing to obtain a fusion pooling feature map; by segmenting the first fusion feature map, the signboard in the first fusion feature map is segmented to obtain the signboard fusion feature map, and then the global average pooling processing can compress the spatial dimension, suppress noise and enhance the corresponding key channel, so that the fusion pooling feature map has global statistical information of the channel dimension; then, the fusion attention feature map is subjected to global attention processing and local attention processing respectively, and the channel relationship is established through the full connection layer in the process of global attention processing; the local attention processing realizes the interaction of focusing on the spatial local features through the spatial convolution, so as to extract fine-grained details, and meanwhile, the calculation complexity of the self-attention can be reduced, so that the global fusion feature map and the local fusion feature are obtained respectively; by globally capturing the signboard information and locally enhancing the edge positioning, the adaptability of the defect detection network model to complex scenes is improved; secondly, the global fusion feature map and the local fusion feature map are subjected to batch normalization and activation processing to obtain the global activation feature map and the local activation feature map; the batch normalization processing and the activation processing can accelerate the convergence and relieve the gradient explosion; then, the residual attention mechanism is used for residual connection of the global activation feature map and the local activation feature map to obtain a fusion attention feature map; through the residual connection, the fusion attention feature map not only retains the information in the original signboard fusion feature map, but also strengthens the discriminative region; and by assigning a learnable weight to the global activation feature map, the local activation feature map and the signboard fusion feature map, the fusion attention feature map can highlight the key features and suppress the interference features; finally, the fusion attention feature map is subjected to 3×3 convolution, batch normalization, activation and maximum pooling processing in sequence to obtain a signboard attention feature map; by performing 3×3 convolution on the fusion attention feature map, the dimension of the fusion attention feature map can be adjusted, the cross-channel information can be fused and the spatial correlation can be enhanced; and the maximum pooling processing is to retain the response features through spatial downsampling to improve the robustness of the defect detection model.
[0030] Further, the multi-scale fusion module in S460 performs multi-scale fusion processing on the signboard attention feature map to obtain a signboard multi-scale fusion feature map, specifically: S461: The signboard attention feature map is subjected to multiple downsampling to obtain multiple signboard scale feature maps of different scales; S462: The multiple signboard scale feature maps are randomly divided into two groups to obtain a first scale feature set and a second scale feature set respectively; S463: The first scale feature set is subjected to cross fusion according to the second scale feature set to obtain each cross scale feature map corresponding to each signboard scale feature map in the first scale feature set; S464: Perform multi-scale feature fusion processing on each cross-scale feature map to obtain a cross-fused feature map; S465: Enhance each sign scale feature map in the second scale feature set based on the cross-fusion feature map to obtain each enhanced scale feature map corresponding to each sign scale feature map in the second scale feature set. S466: Perform multi-scale fusion processing on each enhanced scale feature map to obtain an enhanced fusion feature map, and then stitch the enhanced fusion feature map with the cross-fusion feature map to obtain the sign multi-scale fusion feature map.
[0031] Specifically, the multi-scale fusion feature map of the sign is represented by the following formula: , in, This represents the multi-scale fused feature map of the sign. This indicates a feature concatenation operation. Represents the cross-fusion feature map. This represents a 1×1 convolution + activation + batch normalization. Indicates the first A scaled feature map, Indicates bilinear interpolation scaling. Indicates altitude, Indicates width, Represents the cross-fusion feature map. Indicates spatial dimensions. Represents the second-scale feature set. A diagram illustrating the scale and features of a sign. Indicates the first Each cross-scale feature map This represents a 3×3 convolution + activation + batch normalization. Represents the first scale feature set. A diagram illustrating the scale and features of a sign. This represents the second-scale feature set.
[0032] It is important to understand that, firstly, the attention feature map of the sign is downsampled multiple times to obtain multiple sign scale feature maps at different scales. Through multiple downsampling, detailed information and global contextual information of the attention feature map of the sign at different scales can be captured, which helps to enhance the defect detection network model's ability to recognize signs of different sizes. Then, the multiple signboard scale feature maps are randomly divided into two groups to obtain a first scale feature set and a second scale feature set respectively; through random grouping, the generalization ability of the defect detection network model can be enhanced, and support is provided for subsequent cross fusion; secondly, cross fusion is performed on the first scale feature set according to the second scale feature set to obtain each cross scale feature map corresponding to each signboard scale feature map in the first scale feature set; through cross fusion, the interaction and complementation of signboard scale feature maps of different scales can be realized, and the feature expression ability of signboard scale feature maps of different scales is strengthened; then, each cross scale feature map is subjected to multi-scale feature fusion processing; through multi-scale fusion of all cross scale feature maps, the feature information of signboard scale feature maps under different scales is fused, so that a cross fusion feature map with rich feature information is obtained; then, each signboard scale feature map in the second scale feature set is strengthened according to the cross fusion feature map to obtain each strengthened scale feature map corresponding to each signboard scale feature map in the second scale feature set; through reverse strengthening of the second scale feature set by the cross fusion feature map, the feature expression ability of signboard scale feature maps of different scales can be enhanced; finally, each strengthened scale feature map is subjected to multi-scale fusion processing to obtain a strengthened fusion feature map, and the strengthened fusion feature map and the cross fusion feature map are spliced to obtain a signboard multi-scale fusion feature map; through secondary fusion, a strengthened fusion feature map capable of highlighting key region features is generated, and finally through a third fusion, the strengthened fusion feature map and the cross fusion feature map are spliced to obtain a signboard multi-scale fusion feature map that integrates global and local information.
[0033] According to the defect detection method for engineering signboards based on deep learning, the initial feature map is segmented by the feature segmentation module to segment out the signboard features and ground features, and then the feature enhancement module is used to strengthen the signboard feature map and the ground feature map respectively. The separate strengthening helps to highlight the key features and important features in the signboard feature map and the ground feature map. The feature fusion module is used to fuse the signboard enhanced feature map, the ground enhanced feature map and the initial feature map, so that the fused feature map has both the initial features and the regionally strengthened features, thereby enhancing the defect recognition ability of the defect detection network model. The attention module is used to globally capture the signboard information of the fused attention feature map and locally enhance the edge positioning, and the double-branch cooperation improves the adaptability of the defect detection network model to complex scenes. The multi-scale fusion module is used to cross-fuse and strengthen the signboard scale feature maps of different scales, thereby enhancing the feature expression ability of the signboard scale feature maps of different scales, so that the defect detection network model can accurately capture defects and enhance the detection precision of the defect detection network model.
[0034] It should be noted that the execution subject of the defect detection method for engineering signboard based on deep learning provided in the embodiment of the present application can be a defect detection system for engineering signboard based on deep learning, or a control module in the defect detection system for engineering signboard based on deep learning for executing the defect detection method for engineering signboard based on deep learning. In the embodiment of the present application, the defect detection method for engineering signboard based on deep learning is executed by the defect detection system for engineering signboard based on deep learning as an example to illustrate the defect detection method for engineering signboard based on deep learning provided in the embodiment of the present application.
[0035] The defect detection system for engineering signboard based on deep learning comprises the following modules: The acquisition module acquires a plurality of engineering signboard images to construct an engineering signboard dataset according to the plurality of engineering signboard images and pre-process the engineering signboard dataset to obtain a target dataset. The construction module constructs a defect detection network model, wherein the defect detection network model comprises an initial feature extraction module, a feature segmentation module, a feature enhancement module, a feature fusion module, an attention module and a multi-scale fusion module. The training module trains the defect detection network model according to a part of data in the target dataset and tests the defect detection network model according to another part of data. The prediction module acquires an engineering signboard image to be detected, processes the engineering signboard image to be detected according to the tested defect detection network model and obtains a defect detection result of the engineering signboard image to be detected.
[0036] The defect detection system for engineering signboard based on deep learning in the embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an ios operating system or other possible operating systems, which are not limited in the embodiment of the present application.
[0037] The defect detection system for engineering signboard based on deep learning provided in the embodiment of the present application can realize each process of the defect detection method for engineering signboard based on deep learning in the method embodiment and achieve the same technical effects. To avoid repetition, each process will not be described here again. Figs. 1-2 The defect detection system for engineering signboard based on deep learning provided in the embodiment of the present application can realize each process of the defect detection method for engineering signboard based on deep learning in the method embodiment and achieve the same technical effects. To avoid repetition, each process will not be described here again.
[0038] Optionally, the embodiment of the present application further provides an electronic device, which comprises a processor, a memory, a program or instructions stored on the memory and executable on the processor. When the program or instructions are executed by the processor, each process of the defect detection method for engineering signboard based on deep learning in the method embodiment is realized, and the same technical effects are achieved. To avoid repetition, each process will not be described here again.
[0039] The embodiment of the application further provides a readable storage medium, which stores a program or instructions, and the program or instructions are executed by a processor to realize each process of the above-mentioned deep learning-based engineering signboard defect detection method embodiment and achieve the same technical effects. To avoid repetition, details are not described herein.
[0040] The processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0041] It should be noted that, in this document, the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or apparatus including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element. In addition, it should be pointed out that the scope of the method and apparatus in the embodiments of the application is not limited to the order of functions shown or discussed, and can also include functions performed in a substantially simultaneous manner or in reverse order, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.
[0042] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, air conditioner or network device, etc.) execute the method described in each embodiment of the application.
[0043] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.
Claims
1. A defect detection method for engineering signs based on deep learning, characterized in that, The method includes: Multiple images of construction signs are acquired to construct a construction sign dataset based on the multiple images of construction signs, and the construction sign dataset is preprocessed to obtain a target dataset; A defect detection network model is constructed, which includes an initial feature extraction module, a feature segmentation module, a feature enhancement module, a feature fusion module, an attention module, and a multi-scale fusion module. The defect detection network model is trained using a portion of the target dataset, and tested using another portion of the dataset. The image of the project sign to be inspected is acquired, and the image is processed according to the defect detection network model after testing to obtain the defect detection result of the project sign image.
2. The method for defect detection of engineering signs based on deep learning according to claim 1, characterized in that: The preprocessing of the engineering signage dataset to obtain the target dataset specifically involves: The size of each engineering sign image in the engineering sign dataset is adjusted to obtain a corresponding standard image. Each of the standard images is subjected to noise reduction and smoothing processing to obtain each smoothed image corresponding to each of the standard images; Each of the smoothed images is enhanced and sharpened to obtain each sharpened image corresponding to each of the smoothed images; The area outside the engineering sign and ground in each of the sharpened images is cropped to obtain each target image corresponding to each sharpened image. The target dataset is constructed based on each target image.
3. The method for defect detection of engineering signs based on deep learning according to claim 1, characterized in that, The defect detection network model, after testing, is used to process the image of the engineering sign to be inspected, and the defect detection result of the image of the engineering sign to be inspected is obtained, specifically as follows: The initial feature extraction module performs initial feature extraction processing on the image of the engineering sign to be inspected to obtain the initial feature map corresponding to the image of the engineering sign to be inspected. The initial feature map is segmented according to the feature segmentation module to obtain a sign feature map and a ground feature map; The feature enhancement module enhances the sign feature map and the ground feature map to obtain an enhanced sign feature map and an enhanced ground feature map. The feature fusion module fuses the enhanced feature map of the sign, the enhanced feature map of the ground, and the initial feature map to obtain a first fused feature map. The first fused feature map is processed by the attention module to obtain the sign attention feature map; The multi-scale fusion module performs multi-scale fusion processing on the sign attention feature map to obtain a sign multi-scale fusion feature map, and then fuses the sign enhanced feature map, the sign multi-scale fusion feature map and the initial feature map to obtain a second fusion feature map; The tilting defect result of the signboard in the image of the project signboard to be inspected is obtained based on the first fusion feature, and the surface damage defect result of the signboard in the image of the project signboard to be inspected is obtained based on the second fusion feature.
4. The method for defect detection of engineering signs based on deep learning according to claim 3, characterized in that, The initial feature map is segmented according to the feature segmentation module to obtain a sign feature map and a ground feature map; specifically: Edge detection is performed on the initial feature map to obtain the sign edge feature map and the ground edge feature map in the initial feature map, and the initial feature map is cropped at the pixel level to obtain multiple pixel feature maps of the same size; The centers of the pixel feature maps containing the edges of the sign are smoothly connected to obtain the sign pixel feature map, and the centers of the pixel feature maps containing the edges of the ground are smoothly connected to obtain the ground pixel feature map. The sign pixel feature map and the sign edge feature map are fused to obtain the sign target edge feature map, and the ground pixel feature map and the ground edge feature map are fused to obtain the ground target edge feature map; The initial feature map is segmented based on the ground target edge feature map and the sign target edge feature map to obtain the sign feature map and the ground feature map.
5. The method for defect detection of engineering signs based on deep learning according to claim 3, characterized in that, The feature enhancement module enhances the sign feature map and the ground feature map to obtain an enhanced sign feature map and an enhanced ground feature map; specifically: Global max pooling is performed on the sign feature map and the ground feature map respectively to obtain the sign pooled feature map and the ground pooled feature map; The sign pooling feature map and the ground pooling feature map are respectively processed by 1×1 convolution to obtain the sign spatial feature map and the ground spatial feature map, respectively. The sign spatial feature map and the ground spatial feature map are respectively processed by 3×3 depth separable convolution to obtain the sign texture feature map and the ground texture feature map, respectively. The spatial feature map of the sign and the spatial feature map of the ground are processed by 1×1 shallow convolution to obtain the color feature map of the sign and the color feature map of the ground, respectively. Batch normalization and activation processing are performed on the sign spatial feature map, sign texture feature map, and sign color feature map, as well as the ground spatial feature map, ground texture feature map, and ground color feature map, respectively. Residual links are performed on the activated sign spatial feature map, sign texture feature map, and sign color feature map, respectively, and residual links are also performed on the activated ground spatial feature map, ground texture feature map, and ground color feature map, respectively, to obtain the sign enhanced feature map and the ground enhanced feature map.
6. The method for defect detection of engineering signs based on deep learning according to claim 3, characterized in that, The process of processing the first fused feature map using the attention module to obtain the sign attention feature map is as follows: The first fused feature map is segmented to obtain a sign fused feature map, and the sign fused feature map is subjected to global average pooling to obtain a fused pooled feature map. The fused attention feature map is subjected to global attention processing and local attention processing respectively to obtain a global fused feature map and a local fused feature map respectively; Batch normalization and activation processing are performed on the global fusion feature map and the local fusion feature map to obtain a global activation feature map and a local activation feature map; The global activation feature map and the local activation feature map are residually connected according to the residual attention mechanism to obtain the fused attention feature map; The fused attention feature map is sequentially processed with 3×3 convolution, batch normalization, activation, and max pooling to obtain the sign attention feature map.
7. The method for defect detection of engineering signs based on deep learning according to claim 3, characterized in that, The multi-scale fusion module performs multi-scale fusion processing on the sign attention feature map to obtain a multi-scale fused feature map of the sign, specifically: The attention feature map of the sign was downsampled multiple times to obtain multiple sign scale feature maps at different scales; The scale feature maps of the multiple signs are randomly divided into two groups to obtain the first scale feature set and the second scale feature set, respectively. Based on the second scale feature set, the first scale feature set is cross-fused to obtain each cross-scale feature map corresponding to each sign scale feature map in the first scale feature set. Perform multi-scale feature fusion processing on each of the cross-scale feature maps to obtain a cross-fused feature map; Based on the cross-fusion feature map, each sign scale feature map in the second scale feature set is enhanced to obtain each enhanced scale feature map corresponding to each sign scale feature map in the second scale feature set. Each of the enhanced scale feature maps is subjected to multi-scale fusion processing to obtain an enhanced fusion feature map, and the enhanced fusion feature map is then spliced with the cross-fusion feature map to obtain the signboard multi-scale fusion feature map.
8. A deep learning-based defect detection system for engineering signs, implementing the deep learning-based defect detection method for engineering signs as described in any one of claims 1-7, characterized in that, The system includes: Acquisition module: Acquires multiple images of construction signs, constructs a construction sign dataset based on the multiple images of construction signs, and preprocesses the construction sign dataset to obtain a target dataset; Construction Module: Constructs a defect detection network model, which includes an initial feature extraction module, a feature segmentation module, a feature enhancement module, a feature fusion module, an attention module, and a multi-scale fusion module; Training module: Trains the defect detection network model using a portion of the target dataset, and tests the defect detection network model using another portion of the dataset; Prediction module: Acquires the image of the project sign to be inspected, processes the image of the project sign to be inspected according to the tested defect detection network model, and obtains the defect detection result of the image of the project sign to be inspected.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of a deep learning-based defect detection method for engineering signs as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of a deep learning-based defect detection method for engineering signs as described in any one of claims 1-7.
Citation Information
Patent Citations
Method and device for joint segmentation and 3D reconstruction of a scene
CN110121733A
Aluminum profile surface defect classification method and device based on deep learning
CN116188361A
Ceramic tile surface defect segmentation method based on improved U2-Net
CN116645514A
Railway track surface defect detection method based on multi-scale context information feature fusion
CN118761955A
RCSOSA-AFGC traffic sign detection method and system fused with lightweight detection head
CN119741673A