New Vehicle Surface Defect Detection Method and System Based on the FamsYOLO Network

By building the FamsYOLO network, combining the DBL-3 module, residual structure and SEnet module, the loss function is optimized, and the problem of small and medium-sized object detection of surface defects in new cars is solved, and efficient surface defect detection of new cars is achieved, ensuring factory quality and speed.

CN120070429BActive Publication Date: 2025-07-22EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510535095.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-22
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing new car surface defect detection methods have problems such as insufficient accuracy and slow detection speed when dealing with small defects, making it difficult to achieve efficient automated detection.

Method used

The FamsYOLO network is built, including the backbone, neck and output part. Through the combination of DBL-3 module, residual structure, Fasterc3k2 module and SEnet module, the loss function is optimized, and the image data set is obtained and annotated is used with a gantry with a camera, and the network is trained to achieve high-precision and fast defect detection.

Benefits of technology

The surface defect detection of new cars with high accuracy and robustness in small target detection is achieved, which improves detection efficiency, reduces calculation amount and memory access, enhances the feature extraction and recognition ability of small targets, and ensures the factory quality and speed of new cars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070429B_ABST
    Figure CN120070429B_ABST
Patent Text Reader

Abstract

The present invention relates to a new vehicle surface defect detection method and system based on the FamsYOLO network, which includes the following steps: acquiring images of the front, rear, left, right, and top surfaces of different types of new vehicles before leaving the factory, and identifying vehicle surface defects in the images to construct an image dataset; constructing the FamsYOLO network, the backbone part of which is serially composed of the head DBL-3 module and the middle residual structure from top to bottom; the neck part is successively the DBL-1 module, upsampling, the first feature splicing, the Fasterc3k2 module, Reshape transformation, the second feature splicing, the Fasterc3k2 module, and the SEnet module from bottom to top; training the FamsYOLO network using the image dataset, and the trained network is used for the detection of new vehicle surface defects. It can maintain high accuracy and robustness in the case of small defects and improve the detection performance when the vehicle surface defects are small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of small target detection, and in particular to a new car surface defect detection method and system based on the FamsYOLO network. Background Art

[0002] In the technical field of the automotive manufacturing industry, it is crucial to ensure the ex-factory quality and speed of products. In automotive manufacturing, the inspection of vehicles is the last step before leaving the factory. Precise target detection to identify the surface condition of new cars can prevent unqualified products from entering the market, helping to eliminate potential economic losses and legal disputes; while rapid target detection to identify the surface condition of new cars can promptly discover quality problems and improve the overall production efficiency. With the development of technology, image-based vehicle surface defect detection technology has made remarkable progress, but still faces various challenges. Especially when the vehicle surface defects are small, the accuracy and speed of defect detection technology still need to be further improved.

[0003] In traditional surface defect detection methods, there is a certain effect in identifying larger defects in vehicles, but often performs poorly when dealing with small defects, and there is a phenomenon of difficulty in identification. In the automotive manufacturing industry, the defects of new cars are generally relatively small. When traditional surface defect detection methods are used for new car defect detection, there are disadvantages such as insufficient accuracy, incomplete recognition, and room for improvement in detection speed.

[0004] Therefore, developing a new car surface defect detection method that can effectively handle small targets and realize automatic detection of surface defects during new car ex-factory has important practical significance and application value. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the purpose of the present invention is to provide a new car surface defect detection method and system based on the FamsYOLO network, which can maintain high accuracy and robustness when the defects are small, and at the same time have a relatively fast processing ability to improve the detection performance when the vehicle surface defects are small.

[0006] To solve the above technical problems, the technical solution of the present invention is as follows:

[0007] In the first aspect, the present invention provides a new car surface defect detection method based on the FamsYOLO network, and the detection method includes the following steps:

[0008] Step 1: Obtain images of the front, rear, left, right, and top surfaces of vehicles before leaving the factory for different types of new cars, and mark the vehicle surface defects in the images to construct an image dataset;

[0009] Step 2: Construct the FamsYOLO network,

[0010] The FamsYOLO network includes three parts: the backbone, the neck, and the output. The backbone part is serially composed of the head DBL-3 module and the middle residual structure from top to bottom. The middle residual structure is serially composed of a one-layer residual module and three two-layer residual modules.

[0011] The neck part, from bottom to top, is successively the DBL-1 module, upsampling, the first feature concatenation, the Fasterc3k2 module, Reshape transformation, the second feature concatenation, the Fasterc3k2 module, and the SEnet module.

[0012] The output of the third two-layer residual module in the backbone part is connected to the DBL-1 module in the neck. The output of the DBL-1 module is upsampled and then feature concatenated with the output of the second two-layer residual module in the backbone part. After being processed by the Fasterc3k2 module, it is then feature concatenated with the result of the output of the first two-layer residual module in the backbone part after being processed by the Reshape transformation in the neck. After being processed by another Fasterc3k2 module and the SEnet module, the output of the neck is obtained.

[0013] The Fasterc3k2 module includes a 1×1 convolutional kernel, a feature segmentation operation, two FasterC3 modules, a feature concatenation, and a 1×1 convolutional kernel connected in sequence. Part of the result of the feature segmentation operation is simultaneously feature concatenated with the outputs of the first FasterC3 module and the second FasterC3 module, and then processed by the 1×1 convolutional kernel to obtain the output of the Fasterc3k2 module.

[0014] The FasterC3 module includes a 1×1 convolutional kernel, a FasterBlock module, a feature concatenation, and a 1×1 convolutional kernel connected in sequence. The output of the first 1×1 convolutional kernel is feature concatenated with the output of the FasterBlock module and then processed by the second 1×1 convolutional kernel to obtain the output of the FasterC3 module.

[0015] The FasterBlock module includes a partial 3×3 convolutional kernel, a 1×1 convolutional kernel, BN normalization, the activation function ReLu, and a 1×1 convolutional kernel connected in sequence. The input of the partial 3×3 convolutional kernel is connected to the output of the last 1×1 convolutional kernel for residual connection to obtain the output of the FasterBlock module.

[0016] The output part, from top to bottom, is successively the DBL-1 module, the DBL-3 module, the DBL-1 module, the DBL-3 module, and a 1×1 convolutional kernel. By processing the output feature map of the neck, the surface defects of vehicles in an image can be recognized, and the category and bounding box of the detection target can be displayed in the image.

[0017] So far, the construction of the FamsYOLO network is completed;

[0018] Step 3: Use the image dataset in Step 1 to train the FamsYOLO network, and the trained FamsYOLO network is used for the detection of new vehicle surface defects.

[0019] Furthermore, the SEnet module includes a global pool, two fully connected layers, a Sigmoid activation function, and a Hadamard product.

[0020] Furthermore, in Step 1, set up a gantry in the vehicle production workshop, install two cameras with opposite directions on the crossbeam of the gantry, and install one camera on each of the left and right brackets of the gantry. The cameras on the crossbeam form a 45° angle with the ground, and the cameras on the left and right brackets are level with the ground; use the cameras to obtain the front, rear, left, right, and top images of the vehicle, and use the LabelImg tool for annotation to obtain the annotated new vehicle images for constructing the image dataset.

[0021] Furthermore, the comprehensive loss function of the FamsYOLO network is the sum of the confidence loss and 3 / 2 times the bounding box localization loss. Specifically:

[0022]

[0023]

[0024]

[0025] Among them, represents the comprehensive loss function, is the bounding box localization loss, is the confidence loss, represents the number of samples, represents the true label of sample i, represents the predicted class probability, represents the predicted bounding box, represents the true bounding box, represents containing and the diagonal length of the smallest closed region, represents the Euclidean distance between the center points of the predicted box and the true box, and IOU is the intersection over union.

[0026] Furthermore, the size of the images in the image dataset is 416×416; when starting training, the parameters initialized for the network are set as follows: the epoch for training the network is set to 850, the optimizer uses the adaptive learning rate optimization algorithm Adam optimizer, and the initial learning rate of Adam is set to 0.01; stop training when the loss change error is within ±1e-4.

[0027] Furthermore, the resolution of the surface defects of the vehicle is less than 32 pixels × 32 pixels.

[0028] Furthermore, the average pixel accuracy mAP of the trained FamsYOLO network for detecting surface defects of new vehicles is greater than 60%, the number of parameters is controlled between 4 - 5 million, and the frames per second FPS is greater than 70.

[0029] In a second aspect, the present invention provides a new vehicle surface defect detection system based on the FamsYOLO network. The system executes the steps of the method, including:

[0030] A gantry module for acquiring images of the new vehicle surface;

[0031] The FamsYOLO network for real-time object detection of surface defects of new vehicles;

[0032] A feedback module for feeding back the surface conditions of the new vehicle according to the detection results of the FamsYOLO network. If there are defects on the new vehicle surface, the operator is reminded to send the vehicle back to the factory.

[0033] Compared with the prior art, the beneficial effects of the present invention are:

[0034] The method of the present invention constructs the FamsYOLO network, which can achieve fast and high-precision detection of surface defects of new vehicles. After detecting surface defects of the vehicle, it can provide surface defect information of the vehicle for the operator, ensuring that the operator can process the vehicle more efficiently and guaranteeing the quality and speed of new vehicle production.

[0035] The neck of the FamsYOLO network of the present invention uses the Fasterc3k2 module multiple times and combines with the SEnet module, greatly reducing the computational load to enhance the speed of object detection, reducing the amount of calculation and memory access while ensuring the model performance, and being able to analyze input data from multiple angles, effectively capturing more comprehensive feature information. By adding weights to different channels, it enhances the model's attention to important channels. The two work together, enabling the model to have the ability to extract and recognize features of small targets, effectively extract all pixel feature information in the image when the target is small, improve the quality of feature extraction, and enhance the robustness of object detection and the accuracy of the model.

[0036] The present invention is proposed for the case where the detection target is small, and has high recognition accuracy and detection efficiency for small target detection objects, and can be used to complete various small target detection tasks.

[0037] In summary, in the detection of new car surface defects with relatively small targets, the FamsYOLO network of the present invention improves the feature extraction ability, captures more comprehensive features while reducing the computational load, and can better handle the detection task of small targets. It not only improves the accuracy of target detection, but also realizes lightweight and reduces the cost of the factory. At the same time, it takes into account accuracy, computational efficiency and minimization of the number of parameters, showing excellent performance in the field of new car quality detection and having broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic structural diagram of the FamsYOLO network according to an embodiment of the present invention.

[0039] Figure 2 It is a schematic structural diagram of the DBL-n module.

[0040] Figure 3 It is a schematic structural diagram of the FasterBlock module.

[0041] Figure 4 It is a schematic structural diagram of the FasterC3 module.

[0042] Figure 5 It is a schematic structural diagram of the Fasterc3k2 module.

[0043] Figure 6 It is a schematic flow diagram of using the FamsYOLO network to detect and process new car surface defects according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] In order to more clearly describe the technical problems, technical solutions and advantages of the present invention, the following will be described in detail with reference to the drawings and embodiments. It should be noted that these embodiments are only used to illustrate the principles and application scope of the present invention and cannot be used as a limitation on the protection scope of the present invention.

[0045] In the description of this specification, the specific features, structures or characteristics described in different embodiments can be combined in a suitable manner in any one or more embodiments.

[0046] Embodiment 1:

[0047] The method for detecting new car surface defects based on the FamsYOLO network in this embodiment uses the target detection algorithm for small targets - the FamsYOLO network, and includes the following steps:

[0048] Step 1: Obtain a data set

[0049] Use a gantry with a camera to obtain images of the vehicle surfaces of different vehicle models and different colors. In this embodiment, there are a total of 6,200 images in the image dataset (in this embodiment, mainly images of sedan cars, including compact sedans, medium-sized sedans, large medium-sized sedans, and large sedans), and use the LabelImg tool for annotation to identify the defects in the images. If there is one defect in an image, it is marked as a defective image, and the bounding boxes of all the identified defects are marked. If there is no defect in an image, it is recorded as a non-defective image. The image dataset is randomly divided into a training set and a test set according to a ratio of 7:3. When inputting into the network model, to ensure the same size of the images, all the images are set to a size of 416×416.

[0050] Step 2: Construct the FamsYOLO network

[0051] The FamsYOLO network has a structure as Figure 1 shown, and is mainly composed of three parts: the backbone, the neck, and the output.

[0052] The backbone part is serially composed of the head DBL-3 module and the middle residual structure from top to bottom. Among them, the middle residual structure is serially composed of a one-layer residual module and three two-layer residual modules; the structure of the DBL-n module can be referred to Figure 2 , the DBL-n module is composed of a convolutional kernel of n×n, BN normalization, and the activation function ReLu. The DBL-3 module is the DBL-n module where the convolutional kernel of n×n is a convolutional kernel of 3×3. The DBL-3 module is composed of a convolutional kernel of 3×3, BN normalization, and the activation function ReLu. The DBL-1 module in the following text is the DBL-n module where the convolutional kernel of n×n is a convolutional kernel of 1×1.

[0053] The input image is first processed by the backbone part. It is processed by the head DBL-3 module of the backbone part to perform initial feature extraction on the image and expand its number of channels. The input image with a size of 416×416×3 (height×width×number of channels) is extracted into an initial feature map with a size of 416×416×32 . Next, it sequentially passes through the middle residual structure of the backbone part. The initial feature map is subjected to deep feature extraction by the first one-layer residual module, and a feature map with an output size of 208×208×64 is obtained ; then the feature map is subjected to re-feature extraction by the first two-layer residual module to obtain a feature map with a size of 104×104×128 ; next, the feature map is subjected to the same feature extraction operation by the second two-layer residual module to obtain a feature map with a size of 52×52×256 ; The last two-layer residual module performs feature extraction on the feature map to obtain a feature map of size 26×26×512 , and the feature map , , is used as part of the input for the next stage.

[0054] The feature maps , , serve as the input for the neck part. The feature map undergoes a DBL-1 module and an upsampling process to obtain a feature map of size 52×52×128 . The feature map is then concatenated with the feature map to obtain a feature map of size 52×52×384 ; Then, using the feature map as the input, feature extraction is performed through a Fasterc3k2 module to obtain a feature map of size 52×52×384 .

[0055] In the Fasterc3k2 module, the feature map first undergoes extraction and feature segmentation operations with a convolution kernel of 1×1, and the feature map of size 52×52×384 is divided into two equal parts of size 52×52×192 , , which serves as the input for the first layer of the FasterC3 module. Among them, the feature map is first processed by a convolution kernel of 1×1 in the FasterC3 module to obtain a feature map of size 52×52×98 . The feature map serves as the input for the FasterBlock module. In the FasterBlock module, it sequentially undergoes the effects of a partial convolution kernel of 3×3, a convolution kernel of 1×1, BN normalization, the activation function ReLu, and a convolution kernel of 1×1, and then through a residual connection to obtain a feature map of size 52×52×98 , obtaining the output of the FasterBlock module. Then, the feature map of size 52×52×98 is concatenated with the feature map , and then convolution processing with a convolution kernel of 1×1 is performed to obtain a feature map of the same size as , which is the output of the FasterC3 module. ​

[0056] Feature map After being processed as the input of the second - layer FasterC3 module, a feature map with a size of 52×52×192 is obtained , and then the feature map 、 、 are subjected to feature concatenation to obtain a feature map with a size of 52×52×576 , the feature map is further passed through a convolution kernel of 1×1 to reduce the number of channels of the feature image, obtaining a feature map with a size of 52×52×384 . Among them, the FasterBlock module structure is as Figure 3 shown, the FasterC3 module structure is as Figure 4 shown, and the Fasterc3k2 module structure is as Figure 5 shown.

[0057] Among them, the operation formula of some 3×3 convolution kernels in the FasterBlock module is as follows:

[0058]

[0059] Among them, is the output feature map of partial convolution, represents the 3×3 convolution kernel operation; is the mask, which is a learnable parameter; is the input feature map.

[0060] The calculation formula in the FasterC3 module is as follows:

[0061]

[0062] Among them, is the output feature map of the FasterC3 module, represents the operation of the 1×1 convolution kernel, is the feature map after being processed by the 1×1 convolution kernel, is the feature map after being processed by the FasterBlock module.

[0063] Feature map First, a Reshape deformation operation is performed to obtain a feature map with a size of 52×52×512 , and then the feature map 、 are subjected to feature concatenation to obtain a feature map with a size of 52×52×896 , and then the feature map It is input into the second Fasterc3k2 module for processing to obtain a feature map with a size of 52×52×896 .

[0064] The SEnet module includes a global pooling layer, two fully connected layers, a Sigmoid activation function, and a Hadamard product. For the feature map first, global pooling is performed, changing the size of the feature map from 52×52×896 to 1×1×896. Secondly, it passes through the first fully connected layer to reduce the number of channels of the feature map to 56, and then passes through the second fully connected layer to increase the number of channels of the feature map back to the original 896. Then, a Sigmoid activation function is used to obtain the normalized weights between 0 and 1 for each channel. Next, the normalized weights between 0 and 1 for each channel are weighted to the feature map to obtain the feature map , that is Figure 1 the Hadamard product operation in

[0065] Adding the Fasterc3k2 module in the neck part can improve the detection accuracy and speed of the model for small targets. At the same time, adding the SEnet module can better extract image features, especially the channel weights in the SEnet module, which improves the accuracy of the model without a significant increase in parameters. The calculation process of the SEnet module is as follows:

[0066] First, it is the feature map through global pooling processing, and the process is:

[0067]

[0068] Among them, represents the feature map , represents the number of channels, that is, the c-th channel, and u c represents the feature map in the c-th channel; represents the result of global pooling for the -th channel.

[0069] Then, the result after global pooling processing is processed through a fully connected layer and a Sigmoid activation function, and the process is:

[0070]

[0071] Among them, represents the weight of the c-th channel, represents the Sigmoid activation function, represents the activation function ReLU, represents the fully connected layer;

[0072] Finally, the result of the Sigmoid activation function processing is subjected to a Hadamard product operation with the feature map to obtain , which is the output of the SEnet module:

[0073]

[0074] wherein represents the result of the Hadamard product of the c-th channel.

[0075] The output part from top to bottom is the DBL-1 module, the DBL-3 module, the DBL-1 module, the DBL-3 module, and a convolution kernel of 1×1, and finally a feature map of size 52×52×255 is obtained , and by processing the feature map, it is possible to identify whether there are defects on the surface of the new car in the image, and display the category and bounding box of the detection target in the image.

[0076] Step 3: Use the image dataset obtained in Step 1 to train the FamsYOLO network, and the trained FamsYOLO network is used for the detection of new car surface defects.

[0077] Ablation experiment: Model quality evaluation: The mAP (mean average precision) and FPS (frames per second) of object detection, as well as the number of parameters Parameters, are used to evaluate the object detection performance.

[0078] To detect the role of the SEnet module in identifying small targets, under the constraint of the same loss function, the original network YOLO-S is combined with the Fasterc3k2 module to form the Fc3YOLO network, and the FamsYOLO network formed by further combining the Fc3YOLO network with the SEnet module. In addition, the Fc3YOLO network is combined with multiple network modules to form different networks for ablation experiments. Among them, the Fc3YOLO network is combined with CBAM, that is, CB-Fc3YOLO; the Fc3YOLO network is combined with SKNet, that is, SK-Fc3YOLO combined with the selective kernel network; the Fc3YOLO network is combined with DANet, that is, DA-Fc3YOLO combined with the dual attention network. Different models are used for the detection of new car defect targets after training, and the comparison results obtained are shown in Table 1.

[0079]

[0080] The results shown in Table 1 indicate that for object detection on image datasets of different vehicle models and colors captured by the same gantry with a camera, while keeping the Parameters low and the FPS high, the mAP of F1 is higher than that of CB - FamsYOLO, SK - FamsYOLO, DA - FamsYOLO, and Fc3YOLO. It can be seen from this that the SEnet module in the model can significantly improve the recognition accuracy of small targets. The present application combines the SEnet module and the Fasterc3k2 module synergistically, which can significantly improve the accuracy and detection speed of small target detection.

[0081] Example 2:

[0082] The method for detecting new vehicle surface defects based on the FamsYOLO network in this embodiment realizes the detection of new vehicle surface defects through the following steps:

[0083] 1. Data collection and preparation stage

[0084] 1.1 Data collection

[0085] Operation method: The new vehicle passes through the gantry with a camera in sequence to ensure that the surface images of vehicles with different models and colors are covered.

[0086] 1.2 Data processing

[0087] Image dataset acquisition: The collected image data is preliminarily processed, including removing obvious noise and outliers in the images, and cropping and resizing the images to ensure the consistency and standardization of the image dataset. The collected images are divided into a training set and a test set in a ratio of 7:3.

[0088] Data augmentation: Data augmentation operations such as rotation, scaling, and flipping are performed on the training set to improve the robustness and generalization ability of the model.

[0089] 2. Model training stage

[0090] 2.1 Model construction

[0091] Construct the FamsYOLO network, which includes three parts: the backbone, the neck, and the output. The specific structure is the same as that in Example 1.

[0092] 2.2 FamsYOLO network training

[0093] When starting the training, the parameters for network initialization are set as follows: the epoch for training the network is set to 850, the optimizer uses the adaptive learning rate optimization algorithm Adam optimizer, and the initial learning rate of Adam is set to 0.01.

[0094] During the training process, the training set images are input into the FamsYOLO network, and these images are read according to the training set storage path;

[0095] When the overall network objective loss function no longer shows a significant decrease (error ±1e-4), the network model training is considered to tend to be stable, and the training process is completed.

[0096] When performing network testing, the prepared test set images are input, and the network weights after training on the training set are imported to obtain the test results, achieving the purpose of object detection.

[0097] In view of the characteristic that the detection target is small, the loss function of the FamsYOLO network is further optimized, and a comprehensive loss function is constructed by combining the bounding box localization loss and the confidence loss The specific comprehensive loss function is as follows:

[0098]

[0099]

[0100]

[0101] Among them, represents the comprehensive loss function, represents the number of samples, represents the true label of sample i, represents the predicted class probability, represents the predicted bounding box, represents the true bounding box, represents containing and the diagonal length of the smallest closed region, represents the Euclidean distance between the center points of the predicted box and the true box, and IOU is the intersection over union.

[0102] 3. Processing and analysis stage

[0103] 3.1 Model deployment

[0104] Deployment: In this example, four cameras with a resolution of 1920×1080, a frame rate of 30FPS, and a field of view FOV (H ×V ×D) of 69.4°× 42.5°× 77° (±3°) are used, and are respectively installed on the crossbeam of the gantry and the left and right brackets to form a gantry module. The camera on the crossbeam forms a 45° angle with the ground, and the cameras on the left and right brackets are horizontal with the ground, completing the preparation work of the gantry module.

[0105] Deploy the trained FamsYOLO network to the host computer for real-time surface defect detection of new vehicles. The speed of new vehicles on the vehicle manufacturing assembly line is 1.5 m / s. The trained FamsYOLO network performs object detection on the real-time surface images of new vehicles transmitted by the cameras on the gantry. After identifying the surface conditions of the new vehicles, the recognition results are transmitted to the feedback module. If a defect is identified, the feedback module reports the defect location and reminds the operator to return the vehicle to the factory. For example: There is a defect on the left door, please return the vehicle to the factory for repair; if no defect is identified on the vehicle surface, the feedback module reports that the vehicle surface is intact and can leave the factory (see Figure 6 )

[0106] 3.2 Result Analysis

[0107] Model quality assessment: Evaluate the object detection performance using the mAP (mean average precision) and FPS (frames per second) of object detection, as well as the number of parameters Parameters

[0108] On datasets of different vehicle models and colors (6200 datasets in total) captured by the same gantry with cameras, the FamsYOLO network of the present invention is compared with existing methods. The comparison results show that: compared with existing excellent object detection methods such as YOLO-S, YOLOv5s, and YOLOv8, the object detection method of the present invention has lower computational complexity and better performance. The comparison results obtained after different models are trained for vehicle object detection processing are shown in Table 2

[0109]

[0110] The results shown in Table 2 indicate that when using different object detection models on image datasets of different vehicle models and colors captured by the same gantry with cameras, the mAP of FamsYOLO is higher than that of the relatively excellent existing YOLOv8, YOLOv5s, and YOLO-S, and it still has a relatively high frames per second FPS while ensuring a lower number of parameters Parameters. Compared with the above-mentioned same type of object detection methods, the difference is significant, thus proving that the present invention has superior application performance for surface defect detection of new vehicles with small targets

[0111] In this application, small targets are defects with a resolution less than 32 pixels × 32 pixels

[0112] Example 3

[0113] The hardware devices used in the new vehicle surface defect detection system based on the FamsYOLO network in this embodiment include the following components

[0114] Image sensor: A gantry with a camera is used to collect images, convert optical images into electrical signals, and then form digital images.

[0115] Processor: As the core component of the present invention, the processor is responsible for controlling and managing the operation of the entire system, including functions such as data acquisition, data processing, and image recognition, and requires sufficient computing power and parallel processing capabilities to meet real-time requirements. The processor can adopt different forms such as single-chip microcontrollers, microprocessors, and computers to meet the needs of different application scenarios.

[0116] Memory: The memory can be used to store the collected data and historical data for subsequent processing and analysis, and has characteristics such as high speed, high reliability, and scalability to meet the needs of long-term stable operation of the system.

[0117] Database: A database is used to store and manage information such as the collected data, historical data, and analysis results.

[0118] Network interface: Used for data exchange and communication, with characteristics such as high speed, high stability, and high security to ensure the reliability and security of data transmission.

[0119] The said processor is configured to execute computer-executable instructions, and when the computer-executable instructions are executed by the processor, each step of the new vehicle surface defect detection method based on the FamsYOLO network is realized.

[0120] A computer program is stored on the memory, and the computer program can be executed by the processor to realize each step of the new vehicle surface defect detection method based on the FamsYOLO network.

[0121] The said database is configured to store and manage the data of computer application programs, including various data types and structures, and is applied to each step of the new vehicle surface defect detection method based on the FamsYOLO network.

[0122] The network interface realizes communication and data transmission between computers. The network interface can provide various communication protocols and data transmission methods to meet the communication and data transmission requirements of different application scenarios and different needs, and is applied to each step of the new vehicle surface defect detection method based on the FamsYOLO network.

[0123] Embodiment 4:

[0124] Apply the FamsYOLO network in a vehicle manufacturing factory, especially to improve the ability to identify new vehicle surface defects of different models and different colors.

[0125] Deployment: Deploy the trained FamsYOLO to the host computer in the vehicle manufacturing factory for real-time identification of new vehicle surface defects.

[0126] Quality assessment: Calculate the mAP (mean average precision) and FPS (frames per second) of object detection, as well as the number of parameters Parameters, to evaluate the object detection performance.

[0127] Performance improvement: According to the feedback in actual applications, optimize the network model parameters online to improve the processing efficiency of the system.

[0128] User feedback: Collect user feedback on the object detection effect and further adjust and improve the system settings.

[0129] This embodiment verifies the ability of FamsYOLO to identify new vehicle surface defects under different vehicle models and colors, providing a new solution for new vehicle surface quality inspection.

[0130] The present invention solves the problems of insufficient object detection accuracy and slow processing speed of the current new vehicle surface quality inspection technology when dealing with small targets. The quality inspection of the new vehicle surface with small detection targets is realized through the FamsYOLO network in the present invention, improving the recognition efficiency and saving computing resources.

[0131] The technical solution of the present invention improves the object detection efficiency, avoids the influence caused by the small detection target, and is of great significance for improving the object detection efficiency of new vehicle quality.

[0132] Matters not described in the present invention are applicable to the prior art.

Claims

1. A new car surface defect detection method based on the FamsYOLO network, characterized in that, The detection method includes the following steps: Step 1: Obtain images of the front, rear, left, right, and top surfaces of vehicles of different types before leaving the factory, identify vehicle surface defects in the images, and construct an image dataset; Step 2: Construct the FamsYOLO network. The FamsYOLO network includes three parts: a backbone, a neck, and an output. The backbone part is serially composed of a head DBL-3 module and an intermediate residual structure from top to bottom; the intermediate residual structure is serially composed of a one-layer residual module and three two-layer residual modules; The neck part, from bottom to top, is successively a DBL-1 module, an upsampling layer, a first feature concatenation layer, a Fasterc3k2 module, a Reshape layer, a second feature concatenation layer, a Fasterc3k2 module, and an SEnet module; The output of the third two-layer residual module in the backbone part is connected to the DBL-1 module in the neck. The output of the DBL-1 module is upsampled and then concatenated with the output of the second two-layer residual module in the backbone part. After being processed by the Fasterc3k2 module, it is concatenated with the result of the output of the first two-layer residual module in the backbone part after being processed by the Reshape layer in the neck. After being processed by another Fasterc3k2 module and an SEnet module, the output of the neck is obtained; The Fasterc3k2 module includes a 1×1 convolutional kernel, a feature segmentation operation, two FasterC3 modules, a feature concatenation layer, and a 1×1 convolutional kernel connected in sequence. Part of the result of the feature segmentation operation is concatenated with the outputs of the first FasterC3 module and the second FasterC3 module at the same time, and then processed by the 1×1 convolutional kernel to obtain the output of the Fasterc3k2 module; The FasterC3 module includes a 1×1 convolutional kernel, a FasterBlock module, a feature concatenation layer, and a 1×1 convolutional kernel connected in sequence. The output of the first 1×1 convolutional kernel is concatenated with the output of the FasterBlock module and then processed by the second 1×1 convolutional kernel to obtain the output of the FasterC3 module; The FasterBlock module includes a partial 3×3 convolutional kernel, a 1×1 convolutional kernel, a BN normalization layer, a ReLu activation function, and a 1×1 convolutional kernel connected in sequence. The input of the partial 3×3 convolutional kernel is connected to the output of the last 1×1 convolutional kernel by a residual connection to obtain the output of the FasterBlock module; The output part, from top to bottom, is successively a DBL-1 module, a DBL-3 module, a DBL-1 module, a DBL-3 module, and a 1×1 convolutional kernel. By processing the output feature map of the neck, the surface defects of a vehicle in an image can be recognized, and the category and bounding box of the detection target can be displayed in the image; Thus, the construction of the FamsYOLO network is completed; Step 3: Use the image dataset in Step 1 to train the FamsYOLO network, and the trained FamsYOLO network is used for the detection of new vehicle surface defects.

2. The method according to claim 1, wherein The SEnet module includes a global pooling layer, two fully connected layers, a Sigmoid activation function, and a Hadamard product.

3. The method according to claim 1, wherein In step 1, a gantry is set up in the vehicle production workshop. Two cameras with opposite directions are installed on the crossbeam of the gantry, and one camera is installed on each of the left and right supports of the gantry. The cameras on the crossbeam form a 45° angle with the ground, and the cameras on the left and right supports are horizontal with the ground. The cameras are used to obtain the front, rear, left, right, and top images of the vehicle, and the LabelImg tool is used for annotation to obtain the annotated new vehicle images for constructing the image dataset.

4. The method according to claim 1, wherein The comprehensive loss function of the FamsYOLO network is the sum of the confidence loss and 3 / 2 times the bounding box localization loss.

5. The method according to claim 1, characterized in that, The image size in the image dataset is 416×416. When starting training, the parameters for network initialization are set as follows: the number of epochs for training the network is set to 850, the optimizer uses the adaptive learning rate optimization algorithm Adam optimizer, and the initial learning rate of Adam is set to 0.

01. Training stops when the loss change error is within ±1e-4.

6. The method according to claim 1, characterized in that, The resolution of the surface defects of the vehicle is less than 32 pixels × 32 pixels.

7. The method according to claim 1, wherein The average pixel accuracy mAP of the trained FamsYOLO network for detecting surface defects of new vehicles is greater than 60%, the number of parameters is controlled between 4 - 5 million, and the frames per second FPS is greater than 70.

8. A new car surface defect detection system based on the FamsYOLO network, characterized in that, The system executes the steps of the method according to any one of claims 1 - 7, including: A gantry module for obtaining the surface image of a new vehicle; A FamsYOLO network for performing real-time object detection of surface defects of new vehicles; A feedback module for feeding back the surface condition of the new vehicle according to the detection result of the FamsYOLO network. If there are defects on the surface of the new vehicle, the operator is reminded to send the vehicle back to the factory.

Citation Information

Patent Citations

  • Automobile brake disc surface defect detection method based on deep learning

    CN117541568A

  • Automatic driving YOLO target detection method and system coping with light influence

    CN119478859A