New vehicle surface defect detection method and system based on FamsYOLO network

Through the new vehicle surface defect detection method based on the FamsYOLO network, the problems of insufficient accuracy and slow speed of small and medium-sized defect detection in the existing technology are solved, and the effect of high accuracy and rapid detection is achieved.

CN120070429AActive Publication Date: 2025-05-30EAST CHINA JIAOTONG UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510535095.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy, incomplete identification and slow detection speed when detecting surface defects of new vehicles, especially small defects.

Method used

The new car surface defect detection method based on the FamsYOLO network is adopted. By constructing a FamsYOLO network, the network includes the backbone, the neck and the output part, and using technologies such as multi-layer convolution kernel and feature splicing to improve the model's detection ability of small targets.

Benefits of technology

It achieves high accuracy and robustness in the case of small defects, and has faster processing capabilities, improving the detection performance of small vehicle surface defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070429A_ABST
    Figure CN120070429A_ABST
Patent Text Reader

Abstract

The invention relates to a new vehicle surface defect detection method and system based on a FamsYOLO network, and the method comprises the following steps: obtaining images of the front, rear, left, right and top surfaces of different types of new vehicles before leaving a factory, marking vehicle surface defects in the images, and constructing an image data set; constructing a FamsYOLO network, wherein a trunk part is composed of a head DBL-3 module and a middle residual structure from top to bottom in series; the neck part sequentially comprises a DBL-1 module, an up-sampling module, a first feature splicing module, a Fasterc3k2 module, a Reshape deformation module, a second feature splicing module, a Fasterc3k2 module and an SEnet module from bottom to top; and training a FamsYOLO network by using the image data set, wherein the trained network is used for detecting surface defects of a new vehicle. High accuracy and robustness can be kept under the condition that the defects are small, and the detection performance when the vehicle surface defects are small is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of small target detection, and particularly to a new vehicle surface defect detection method and system based on the FamsYOLO network. Background Art

[0002] In the technical field of the automotive manufacturing industry, it is crucial to ensure the ex-factory quality and speed of products. In automotive manufacturing, the inspection of vehicles is the last step before leaving the factory. Precise target detection to identify the surface condition of new vehicles can prevent unqualified products from entering the market, helping to eliminate potential economic losses and legal disputes; while rapid target detection to identify the surface condition of new vehicles can promptly detect quality problems and improve the overall production efficiency. With the development of technology, significant progress has been made in image-based vehicle surface defect detection technology, but it still faces various challenges. Especially when the vehicle surface defects are small, the accuracy and speed of defect detection technology need to be further improved.

[0003] In traditional surface defect detection methods, there is a certain effect in identifying larger defects in vehicles, but they often perform poorly when dealing with small defects and there are phenomena of difficult identification. In the automotive manufacturing industry, the defects of new vehicles are generally relatively small. When traditional surface defect detection methods are used for new vehicle defect detection, there are disadvantages such as insufficient accuracy, incomplete recognition, and room for improvement in detection speed.

[0004] Therefore, developing a new vehicle surface defect detection method that can effectively handle small targets and realize automatic detection of surface defects during new vehicle ex-factory has important practical significance and application value. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the purpose of the present invention is to provide a new vehicle surface defect detection method and system based on the FamsYOLO network, which can maintain high accuracy and robustness when the defects are small, and at the same time have a relatively fast processing ability to improve the detection performance when the vehicle surface defects are small.

[0006] To solve the above technical problems, the technical solution of the present invention is as follows: In the first aspect, the present invention provides a new vehicle surface defect detection method based on the FamsYOLO network, and the detection method includes the following steps: Step 1: Obtain images of the front, rear, left, right, and top surfaces of vehicles before different types of new vehicles leave the factory, and mark the vehicle surface defects in the images to construct an image dataset; Step 2: Construct the FamsYOLO network, The FamsYOLO network includes three parts: the backbone, the neck, and the output. The backbone part is serially composed of the head DBL-3 module and the middle residual structure from top to bottom. The middle residual structure is serially composed of a one-layer residual module and three two-layer residual modules. The neck part, from bottom to top, is successively the DBL-1 module, upsampling, the first feature concatenation, the Fasterc3k2 module, Reshape transformation, the second feature concatenation, the Fasterc3k2 module, and the SEnet module. The output of the third two-layer residual module in the backbone part is connected to the DBL-1 module in the neck. The output of the DBL-1 module is upsampled and then feature concatenated with the output of the second two-layer residual module in the backbone part. After being processed by the Fasterc3k2 module, it is then feature concatenated with the result of the output of the first two-layer residual module in the backbone part after being processed by the Reshape transformation in the neck. After being processed by another Fasterc3k2 module and the SEnet module, the output of the neck is obtained. The Fasterc3k2 module includes a 1×1 convolutional kernel, a feature segmentation operation, two FasterC3 modules, feature concatenation, and a 1×1 convolutional kernel connected in sequence. A part of the result of the feature segmentation operation is simultaneously feature concatenated with the outputs of the first FasterC3 module and the second FasterC3 module, and then processed by the 1×1 convolutional kernel to obtain the output of the Fasterc3k2 module. The FasterC3 module includes a 1×1 convolutional kernel, a FasterBlock module, feature concatenation, and a 1×1 convolutional kernel connected in sequence. The output of the first 1×1 convolutional kernel is feature concatenated with the output of the FasterBlock module and then processed by the second 1×1 convolutional kernel to obtain the output of the FasterC3 module. The FasterBlock module includes a partial 3×3 convolutional kernel, a 1×1 convolutional kernel, BN normalization, the activation function ReLu, and a 1×1 convolutional kernel connected in sequence. The input of the partial 3×3 convolutional kernel is connected to the output of the last 1×1 convolutional kernel for residual connection to obtain the output of the FasterBlock module. The output part, from top to bottom, is successively the DBL-1 module, the DBL-3 module, the DBL-1 module, the DBL-3 module, and a 1×1 convolutional kernel. By processing the output feature map of the neck, the surface defects of vehicles in an image can be identified, and the category and bounding box of the detection target can be displayed in the image. Thus, the construction of the FamsYOLO network is completed. Step 3: Use the image dataset in Step 1 to train the FamsYOLO network, and the trained FamsYOLO network is used for the detection of surface defects of new vehicles.

[0007] Furthermore, the SEnet module includes a global pooling layer, two fully connected layers, a Sigmoid activation function, and a Hadamard product.

[0008] Furthermore, in step 1, a gantry is set up in the vehicle production workshop. Two cameras with opposite directions are installed on the crossbeam of the gantry, and one camera is installed on each of the left and right brackets of the gantry. The cameras on the crossbeam form a 45° angle with the ground, and the cameras on the left and right brackets are horizontal with the ground; the cameras are used to obtain images of the front, rear, left, right, and top surfaces of the vehicle, and the LabelImg tool is used for annotation to obtain the annotated new vehicle images for constructing the image dataset.

[0009] Furthermore, the comprehensive loss function of the FamsYOLO network is the sum of the confidence loss and 3 / 2 times the bounding box localization loss. Specifically:

[0010]

[0011]

[0012] Among them, represents the comprehensive loss function, is the bounding box localization loss, is the confidence loss, represents the number of samples, represents the true label of sample i, represents the predicted class probability, represents the predicted bounding box, represents the true bounding box, represents including and the diagonal length of the smallest closed region containing represents the Euclidean distance between the center points of the predicted box and the true box, and IOU is the intersection over union.

[0013] Furthermore, the size of the images in the image dataset is 416×416; when starting training, the parameters for network initialization are set as follows: the epoch for training the network is set to 850, the optimizer uses the adaptive learning rate optimization algorithm Adam optimizer, and the initial learning rate of Adam is set to 0.01; training stops when the loss change error is within ±1e - 4.

[0014] Furthermore, the resolution of the surface defects of the vehicle is less than 32 pixels × 32 pixels.

[0015] Furthermore, the average pixel accuracy mAP of the trained FamsYOLO network for detecting defects on the surface of new cars is greater than 60%, the number of parameters is controlled between 4 - 5 million, and the frames per second FPS is greater than 70.

[0016] In a second aspect, the present invention provides a new car surface defect detection system based on the FamsYOLO network. The system executes the steps of the method, including: A gantry module for acquiring images of the surface of new cars; The FamsYOLO network for real-time target detection of defects on the surface of new cars; A feedback module for feeding back the surface condition of the new car according to the detection results of the FamsYOLO network. If there are defects on the surface of the new car, the operator is reminded to return the vehicle to the factory.

[0017] Compared with the prior art, the beneficial effects of the present invention are: The method of the present invention constructs the FamsYOLO network, which can achieve fast and high-precision detection of defects on the surface of new cars. After detecting defects on the vehicle surface, it can provide vehicle surface defect information for the operator, ensuring that the operator can process the vehicle more efficiently and guaranteeing the quality and speed of new cars leaving the factory.

[0018] In the neck of the FamsYOLO network of the present invention, the Fasterc3k2 module is used multiple times and combined with the SEnet module, greatly reducing the computational load to enhance the speed of target detection, reducing the amount of calculation and memory access while ensuring the performance of the model. At the same time, it can analyze the input data from multiple angles, effectively capture more comprehensive feature information, and enhance the model's attention to important channels by adding weights to different channels. The two work together, enabling the model to have the ability to extract and identify features of small targets, effectively extract all pixel feature information in the image when the target is small, improve the quality of feature extraction, and enhance the robustness of target detection and the accuracy of the model.

[0019] The present invention is proposed for the case where the detection target is small, and has high recognition accuracy and detection efficiency for small target detection objects, and can be used to complete various small target detection tasks.

[0020] In summary, the FamsYOLO network of the present invention improves the ability of feature extraction in the detection of defects on the surface of new cars with small targets, captures more comprehensive features while reducing the amount of calculation, can better handle small target detection tasks, not only improves the accuracy of target detection, but also realizes lightweight and reduces the cost of the factory. It also takes into account accuracy, computational effect, and minimization of the number of parameters, showing excellent performance in the field of new car quality detection and having broad application prospects. Description of the Drawings

[0021] Figure 1 It is a schematic structural diagram of the FamsYOLO network according to an embodiment of the present invention.

[0022] Figure 2 It is a schematic structural diagram of the DBL-n module.

[0023] Figure 3 It is a schematic structural diagram of the FasterBlock module.

[0024] Figure 4 It is a schematic structural diagram of the FasterC3 module.

[0025] Figure 5 It is a schematic structural diagram of the Fasterc3k2 module.

[0026] Figure 6 It is a schematic flowchart of using the FamsYOLO network to detect and process new vehicle surface defects according to an embodiment of the present invention. Detailed implementation manners

[0027] In order to more clearly describe the technical problems, technical solutions and advantages of the present invention, the following will be described in detail with reference to the drawings and embodiments. It should be noted that these embodiments are only used to illustrate the principles and application scope of the present invention, and cannot be used to limit the protection scope of the present invention.

[0028] In the description of this specification, the specific features, structures or characteristics described in different embodiments can be combined in any one or more embodiments in a suitable manner.

[0029] Embodiment 1: The new vehicle surface defect detection method based on the FamsYOLO network in this embodiment uses the target detection algorithm for small targets - the FamsYOLO network, and includes the following steps: Step 1: Obtain the data set Use a gantry with a camera to obtain vehicle surface images of different models and colors. In this embodiment, there are a total of 6200 images in the image data set (mainly sedan vehicle images in this embodiment, including compact sedans, medium-sized sedans, large medium-sized sedans, and large sedans), and use the LabelImg tool for annotation to identify the defects in the images. If there is one defect in an image, it is marked as a defective image, and the bounding boxes of all identified defects are marked. If there is no defect in an image, it is recorded as a non-defective image. The image data set is randomly divided into a training set and a test set according to a ratio of 7:3. When inputting into the network model, in order to ensure that the image sizes are consistent, all images are set to a size of 416×416.

[0030] Step 2: Construct the FamsYOLO network The FamsYOLO network has a structure as shown Figure 1 and mainly consists of three parts: the backbone, the neck, and the output.

[0031] The backbone part is serially composed of the head DBL-3 module and the intermediate residual structure from top to bottom. Among them, the intermediate residual structure is serially composed of a one-layer residual module and three two-layer residual modules; the structure of the DBL-n module can be seen in Figure 2 . The DBL-n module consists of a convolutional kernel of n×n, BN normalization, and the activation function ReLu. The DBL-3 module is the DBL-n module with the convolutional kernel of n×n being 3×3. The DBL-3 module consists of a convolutional kernel of 3×3, BN normalization, and the activation function ReLu. The DBL-1 module in the following text is the DBL-n module with the convolutional kernel of n×n being 1×1.

[0032] The input image is first processed by the backbone part. It is processed by the head DBL-3 module of the backbone part to perform initial feature extraction on the image and expand its number of channels. The input image with a size of 416×416×3 (height×width×number of channels) is extracted into an initial feature map with a size of 416×416×32 . Next, it passes through the intermediate residual structure of the backbone part in sequence. The first one-layer residual module performs deep feature extraction on the initial feature map to obtain a feature map with an output size of 208×208×64 ; then the first two-layer residual module performs feature extraction on the feature map again to obtain a feature map with a size of 104×104×128 ; next, the second two-layer residual module performs the same feature extraction operation on the feature map to obtain a feature map with a size of 52×52×256 ; the last two-layer residual module performs feature extraction on the feature map to obtain a feature map with a size of 26×26×512 . The feature maps , , are used as part of the input for the next stage.

[0033] The feature maps , , are used as the input for the neck part. The feature map passes through a DBL-1 module and an upsampling process to obtain a feature map with a size of 52×52×128 . The feature map Then, it is concatenated with the feature map to obtain a feature map with a size of 52×52×384 ; Then, the feature map is used as the input, and after passing through a Fasterc3k2 module for feature extraction, a feature map with a size of 52×52×384 is obtained .

[0034] In the Fasterc3k2 module, the feature map first undergoes extraction and feature segmentation operations with a convolution kernel of 1×1, and the feature map with a size of 52×52×384 is divided into two equal parts with a size of 52×52×192 、 , which serves as the input to the first-layer FasterC3 module. Among them, the feature map is first processed by a convolution kernel of 1×1 in the FasterC3 module to obtain a feature map with a size of 52×52×98 . The feature map serves as the input to the FasterBlock module. In the FasterBlock module, it successively undergoes the effects of a partial convolution kernel of 3×3, a convolution kernel of 1×1, BN normalization, an activation function ReLu, and a convolution kernel of 1×1, and then through a residual connection, a feature map with a size of 52×52×98 is obtained , obtaining the output of the FasterBlock module . Then, the feature map with a size of 52×52×98 is concatenated with the feature map , and then convolution processing with a convolution kernel of 1×1 is performed to obtain a feature map with the same size as , which is the output of the FasterC3 module .

[0035] The feature map serves as the input to the second-layer FasterC3 module and after processing, a feature map with a size of 52×52×192 is obtained . Then, the feature maps 、 、 are concatenated to obtain a feature map with a size of 52×52×576 . The feature map is further passed through a convolution kernel of 1×1 to reduce the number of channels of the feature image and obtain a feature map with a size of 52×52×384 . Among them, the structure of the FasterBlock module is as shown in Figure 3 ​As shown, the structure of the FasterC3 module is as Figure 4 As shown, the structure of the Fasterc3k2 module is as Figure 5 shown.

[0036] Among them, the operation formula of some 3×3 convolutional kernels in the FasterBlock module is as follows:

[0037] Among them, is the output feature map of partial convolution, represents the 3×3 operation of the convolutional kernel; is the mask, which is a learnable parameter; is the input feature map.

[0038] The calculation formula in the FasterC3 module is as follows:

[0039] Among them, is the output feature map of the FasterC3 module, represents the 1×1 operation of the convolutional kernel, is the feature map after being processed by the 1×1 convolutional kernel, is the feature map after being processed by the FasterBlock module.

[0040] Feature map first performs a Reshape deformation operation to obtain a feature map with a size of 52×52×512 , and then the feature map , are feature concatenated to obtain a feature map with a size of 52×52×896 , and then the feature map is input into the second Fasterc3k2 module for processing to obtain a feature map with a size of 52×52×896 .

[0041] The described SEnet module includes global pooling, two fully connected layers, a Sigmoid activation function, and a Hadamard product. The feature map first performs global pooling, changing the feature map size from 52×52×896 to 1×1×896. Secondly, it passes through the first fully connected layer to reduce the number of channels of the feature map to 56, then passes through the second fully connected layer to increase the number of channels of the feature map back to the original 896, and then obtains the normalized weights between 0 and 1 for each channel through a Sigmoid activation function. Then, the normalized weights between 0 and 1 for each channel are weighted into the feature map to obtain the feature map , that is,Figure 1 The Hadamard product operation in

[0042] Add the Fasterc3k2 module in the neck part to improve the detection accuracy and speed of the model for small targets. At the same time, add the SEnet module, which can better extract image features. Especially the channel weights in the SEnet module improve the accuracy of the model without a significant increase in parameters. The calculation process of the SEnet module is as follows: First is the feature map After global pooling processing, the process is:

[0043] Among them, represents the feature map , represents the number of channels, that is, the c-th channel, and u c represents the feature map The feature of the c-th channel in represents the The result of global pooling of the c-th channel.

[0044] Then, perform a fully connected layer and Sigmoid activation function processing on the result of global pooling processing. The process is:

[0045] Among them, represents the weight of the c-th channel, represents the Sigmoid activation function, represents the activation function ReLU, represents the fully connected layer; Finally, perform the Hadamard product operation on the result of the Sigmoid activation function processing and the feature map to obtain , which is the output of the SEnet module:

[0046] Among them, represents the result of the Hadamard product of the c-th channel.

[0047] The output part from top to bottom is the DBL-1 module, DBL-3 module, DBL-1 module, DBL-3 module, and a convolution kernel of 1×1, and finally a feature map of size 52×52×255 is obtained. By processing the feature map, it is possible to identify whether there are defects on the surface of the new car in the image and display the category and bounding box of the detection target in the image.

[0048] Step 3: Train the FamsYOLO network using the image dataset obtained in Step 1. The trained FamsYOLO network is used for detecting defects on the surface of new cars.

[0049] Ablation experiment: Model quality assessment: Evaluate the object detection performance using the mAP (mean average precision) and FPS (frames per second) of object detection, as well as the number of parameters "Parameters".

[0050] To detect the role of the SEnet module in identifying small targets, under the constraint of the same loss function, the original network YOLO-S is combined with the Fasterc3k2 module to form the Fc3YOLO network. The Fc3YOLO network is further combined with the SEnet module to form the FamsYOLO network. In addition, the Fc3YOLO network is combined with multiple network modules to form different networks for ablation experiments. Among them, the Fc3YOLO network is combined with CBAM, namely CB-Fc3YOLO; the Fc3YOLO network is combined with SKNet, namely SK-Fc3YOLO combined with the selective kernel network; the Fc3YOLO network is combined with DANet, namely DA-Fc3YOLO combined with the dual attention network. Different models are used for detecting defect targets on new cars after training, and the comparison results obtained are shown in Table 1.

[0051]

[0052] The results shown in Table 1 indicate that when performing object detection on image datasets of different vehicle models and colors taken by the same gantry with a camera, while maintaining a low "Parameters" and a high FPS, the mAP of F1 is higher than that of CB-FamsYOLO, SK-FamsYOLO, DA-FamsYOLO, and Fc3YOLO. From this, it can be seen that the SEnet module can significantly improve the recognition accuracy of small targets in the model. The present application can significantly improve the accuracy and detection speed of small target detection by synergistically combining the SEnet module and the Fasterc3k2 module.

[0053] Example 2: The method for detecting defects on the surface of new cars based on the FamsYOLO network in this example realizes the detection of defects on the surface of new cars through the following steps: 1. Data collection and preparation stage 1.1 Data collection Operation method: New cars pass through the gantry with a camera in sequence to ensure that the surface images of different vehicle models and colors are covered.

[0054] 1.2 Data processing Image dataset acquisition: The collected image data is preliminarily processed, including removing obvious noise and outliers in the images, as well as cropping and resizing the images to ensure the consistency and standardization of the image dataset. The collected images are divided into a training set and a test set in a 7:3 ratio.

[0055] Data augmentation: Data augmentation operations such as rotation, scaling, and flipping are performed on the training set to improve the robustness and generalization ability of the model.

[0056] 2. Model training stage 2.1 Model construction Construct the FamsYOLO network, which includes three parts: the backbone, the neck, and the output. The specific structure is the same as that in Embodiment 1.

[0057] 2.2 FamsYOLO network training When starting the training, the parameters for network initialization are set as follows: the number of epochs for training the network is set to 850, the optimizer uses the adaptive learning rate optimization algorithm Adam optimizer, and the initial learning rate of Adam is set to 0.01.

[0058] During the training process, the training set images are input into the FamsYOLO network, and these images are read according to the storage path of the training set; When the overall network objective loss function no longer shows a significant decrease (error ±1e-4), it is considered that the training of the network model tends to be stable, and the training process is completed.

[0059] When conducting network testing, the prepared test set images are input, and the network weights after training on the training set are imported to obtain the test results, achieving the purpose of object detection.

[0060] In view of the characteristic that the detection target is small, the loss function of the FamsYOLO network is further optimized, and a comprehensive loss function is constructed by combining the bounding box localization loss and the confidence loss The specific comprehensive loss function is as follows:

[0061]

[0062]

[0063] Among them, represents the comprehensive loss function, represents the number of samples, represents the true label of sample i, represents the predicted class probability, represents the predicted bounding box, Represents the true bounding box, Represents the inclusion of and The diagonal length of the smallest closed region containing Represents the Euclidean distance between the center points of the predicted box and the true box, and IOU is the intersection over union.

[0064] 3. Processing and analysis stage 3.1 Model deployment Deployment: In this example, four cameras with a resolution of 1920×1080, a frame rate of 30 FPS, and a field of view FOV (H × V × D) of 69.4°× 42.5°× 77° (±3°) are used and installed on the crossbeam of the gantry and the left and right brackets respectively to form a gantry module. The cameras on the crossbeam form a 45° angle with the ground, and the cameras on the left and right brackets are horizontal with the ground to complete the preparation work of the gantry module.

[0065] The trained FamsYOLO network is deployed to the host computer for real-time surface defect detection of new vehicles. The new vehicle moves at a speed of 1.5 m / s on the vehicle manufacturing assembly line. The trained FamsYOLO network performs object detection on the real-time surface images of new vehicles transmitted by the cameras on the gantry. After identifying the surface conditions of the new vehicles, the recognition results are transmitted to the feedback module. If a defect is identified, the feedback module reports the defect location and reminds the operator to return the vehicle to the factory. For example: There is a defect on the left door, please return the vehicle to the factory for repair; if no defect is identified on the vehicle surface, the feedback module reports that the vehicle surface is intact and can leave the factory (see Figure 6 ).

[0066] 3.2 Result analysis Model quality assessment: The object detection performance is evaluated by the mAP (mean average precision) and FPS (frames per second) of object detection, as well as the number of parameters Parameters.

[0067] On datasets of different vehicle models and different colors taken by the same gantry with cameras (the number of datasets is 6200), the FamsYOLO network of the present invention is compared with existing methods. The comparison results show that: compared with existing excellent object detection methods such as YOLO-S, YOLOv5s, and YOLOv8, the object detection method of the present invention has lower computational complexity and better performance. The comparison results obtained by different models after training for vehicle object detection processing are shown in Table 2.

[0068]

[0069] The results shown in Table 2 indicate that when using different object detection models on image datasets of different vehicle models and colors captured by the same gantry with a camera, the mAP of FamsYOLO is higher than that of the relatively excellent existing YOLOv8, YOLOv5s, and YOLO-S. Moreover, while ensuring a lower number of parameters (Parameters), it still has a relatively high frames per second (FPS). Compared with the above-mentioned same type of object detection methods, the difference is significant, thus proving that the present invention has superior application performance for detecting surface defects of new vehicles with small targets.

[0070] In this application, small targets are defects with a resolution less than 32 pixels × 32 pixels.

[0071] Example 3: The hardware devices used in the new vehicle surface defect detection system based on the FamsYOLO network in this embodiment include the following components: Image sensor: The gantry with a camera is used to collect images, convert the optical images into electrical signals, and then form digital images.

[0072] Processor: As the core component of the present invention, the processor is responsible for controlling and managing the operation of the entire system, including functions such as data acquisition, data processing, and image recognition. And it needs to have sufficient computing power and parallel processing capabilities to meet the real-time requirements. The processor can adopt different forms such as single-chip microcontrollers, microprocessors, computers, etc. to meet the needs of different application scenarios.

[0073] Memory: The memory can be used to store the collected data and historical data for subsequent processing and analysis, and has characteristics such as high speed, high reliability, and scalability to meet the needs of the long-term stable operation of the system.

[0074] Database: A database is used to store and manage information such as the collected data, historical data, and analysis results.

[0075] Network interface: It is used for data exchange and communication, and has characteristics such as high speed, high stability, and high security to ensure the reliability and security of data transmission.

[0076] The described processor is configured to execute computer-executable instructions, and when the computer-executable instructions are executed by the processor, each step of the new vehicle surface defect detection method based on the FamsYOLO network is implemented.

[0077] A computer program is stored on the memory, and the computer program can be executed by the processor to implement each step of the new vehicle surface defect detection method based on the FamsYOLO network.

[0078] The database is configured to store and manage data of computer applications, including various data types and structures, and is applied to each step of the new vehicle surface defect detection method based on the FamsYOLO network.

[0079] The network interface realizes communication and data transmission between computers. This network interface can provide various communication protocols and data transmission methods to meet the communication and data transmission requirements of different application scenarios and different needs, and is applied to each step of the new vehicle surface defect detection method based on the FamsYOLO network.

[0080] Example 4: Apply the FamsYOLO network in vehicle manufacturing factories, especially to improve the ability to identify new vehicle surface defects of different vehicle models and different colors.

[0081] Deployment: Deploy the trained FamsYOLO to the upper computer in the vehicle manufacturing factory for real-time identification of new vehicle surface defects.

[0082] Quality assessment: Calculate the mAP (mean average precision) and FPS (frames per second) of object detection, as well as the number of parameters Parameters, to evaluate the object detection performance.

[0083] Performance improvement: According to the feedback in actual applications, online optimize the network model parameters to improve the processing efficiency of the system.

[0084] User feedback: Collect user feedback on the object detection effect, and further adjust and improve the system settings.

[0085] This embodiment verifies the ability of FamsYOLO to identify new vehicle surface defects under different vehicle models and different colors, and provides a new solution for new vehicle surface quality detection.

[0086] The present invention solves the problems of insufficient object detection accuracy and slow processing speed when dealing with small objects in the current new vehicle surface quality detection technology. Through the FamsYOLO network in the present invention, the quality detection of the new vehicle surface with smaller detection objects is realized, the recognition efficiency is improved, and the computing resources are saved.

[0087] The technical solution of the present invention improves the object detection efficiency, avoids the influence caused by the small detection object, and is of great significance for improving the object detection efficiency of new vehicle quality.

[0088] Matters not described in the present invention are applicable to the prior art.

Claims

1. A new car surface defect detection method based on FamsYOLO network, characterized in that: The detection method comprises the following steps: Step 1: Obtain images of the front, rear, left, right, and top surfaces of different types of new cars before they leave the factory, identify the surface defects of the vehicles in the images, and construct an image dataset; Step 2: Build the FamsYOLO network. The FamsYOLO network includes three parts: a trunk, a neck, and an output. The trunk is composed of a head DBL-3 module and an intermediate residual structure in series from top to bottom; the intermediate residual structure is composed of a first-layer residual module and three second-layer residual modules in series; The neck part is composed of DBL-1 module, upsampling, first feature splicing, Fasterc3k2 module, Reshape deformation, second feature splicing, Fasterc3k2 module and SEnet module from bottom to top; The output of the third two-layer residual module of the trunk is connected to the DBL-1 module of the neck. The output of the DBL-1 module is upsampled and concatenated with the output of the second two-layer residual module of the trunk. Then, it is processed by the Fasterc3k2 module and concatenated with the output of the first two-layer residual module of the trunk after the Reshape deformation of the neck. The output of the neck is obtained after processing by a Fasterc3k2 module and a SEnet module. The Fasterc3k2 module includes a sequentially connected convolution kernel 1×1, a feature segmentation operation, two FasterC3 modules, feature splicing and a convolution kernel 1×1. A part of the result of the feature segmentation operation is simultaneously feature spliced ​​with the output of the first FasterC3 module and the output of the second FasterC3 module, and then processed by the convolution kernel 1×1 to obtain the output of the Fasterc3k2 module; The FasterC3 module includes a convolution kernel 1×1, a FasterBlock module, feature splicing and a convolution kernel 1×1 connected in sequence. The output of the first convolution kernel 1×1 is feature spliced ​​with the output of the FasterBlock module, and the output of the FasterC3 module is obtained by processing with the second convolution kernel 1×1. The FasterBlock module includes a partial convolution kernel 3×3, a convolution kernel 1×1, BN normalization, an activation function ReLu and a convolution kernel 1×1 connected in sequence, and the input of the partial convolution kernel 3×3 is residually connected with the output of the last convolution kernel 1×1 to obtain the output of the FasterBlock module; The output part is DBL-1 module, DBL-3 module, DBL-1 module, DBL-3 module and convolution kernel 1×1 from top to bottom. By processing the output feature map of the neck, it can identify the surface defects of the vehicle in an image and display the category and bounding box of the detection target in the image; This completes the construction of the FamsYOLO network; Step 3: Use the image dataset in step 1 to train the FamsYOLO network. The trained FamsYOLO network is used to detect surface defects of new cars.

2. The method according to claim 1, characterized in that The SEnet module includes a global pool, two fully connected layers, a Sigmoid activation function, and a Hadamard product.

3. The method according to claim 1, characterized in that In step 1, a gantry is set up in the vehicle production workshop, two cameras in opposite directions are installed on the crossbeam of the gantry, and a camera is installed on each of the left and right brackets of the gantry. The camera on the crossbeam forms a 45° angle with the ground, and the cameras on the left and right brackets are kept horizontal to the ground; the camera is used to obtain images of the front, rear, left, right and top surfaces of the vehicle, and the LabelImg tool is used to annotate them to obtain the annotated new car images for constructing the image dataset.

4. The method according to claim 1, characterized in that: The comprehensive loss function of the FamsYOLO network is the sum of the confidence loss and 3 / 2 times the bounding box localization loss.

5. The method according to claim 1, characterized in that The image size in the image dataset is 416×416; when starting training, the parameters of network initialization are set as follows: the epoch of the training network is set to 850, the optimizer adopts the adaptive learning rate optimization algorithm Adam optimizer, and the initial learning rate of Adam is set to 0.01; the training is stopped when the loss change error is within ±1e-4.

6. The method according to claim 1, characterized in that The resolution of the surface defects of the vehicle is less than 32 pixels × 32 pixels.

7. The method according to claim 1, characterized in that After training, the average pixel accuracy mAP of the FamsYOLO network for new car surface defect detection is greater than 60%, the number of parameters is controlled between 4-5 million, and the frame rate per second FPS is greater than 70.

8. A new car surface defect detection system based on FamsYOLO network, characterized in that: The system performs the steps of any one of the methods of claims 1 to 7, including: Gantry module, used to obtain surface images of new vehicles; FamsYOLO network, used for real-time object detection of surface defects on new cars; The feedback module is used to provide feedback on the surface condition of the new car based on the detection results of the FamsYOLO network. If there are defects on the surface of the new car, the operator will be reminded to return the vehicle to the factory.

Citation Information

Patent Citations

  • Insulator defect detection method based on adaptive feature fusion and lightweight YOLOv5s

    CN117173526A

  • Automobile brake disc surface defect detection method based on deep learning

    CN117541568A

  • Automatic driving YOLO target detection method and system coping with light influence

    CN119478859A

  • Intelligent detection method for multiple types of diseases of bridge near water, and unmanned surface vessel device

    WO2022193420A1