Clip and traditional image processing fusion method for automobile fuse box assembly detection

Through the integration of Clip and traditional image processing methods, combined with CLIP model and traditional image processing technology, the robustness and interpretability problems in the assembly detection of automobile fuse box are solved, and better detection performance and adaptability are achieved.

CN120259265AActive Publication Date: 2025-07-04HEFEI UNIV OF TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510406250.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing automobile fuse box assembly detection technology has problems such as artificial error, high time consumption, high cost, poor robustness, and deep learning methods require a large amount of labeling data, high calculation costs, and poor interpretability.

Method used

Using Clip and traditional image processing fusion method, the CLIP model migration training combines the similarity of image and text features, combined with histogram similarity and cosine similarity judgment, fuse box assembly detection is achieved.

Benefits of technology

It improves the robustness and interpretability of detection, reduces the sensitivity to light conditions, improves the recognition performance in the case of poor light, fading, staining, etc., has strong adaptability, and reduces the need for labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259265A_ABST
    Figure CN120259265A_ABST
Patent Text Reader

Abstract

The invention discloses a Clip and traditional image processing fusion method for automobile fuse box assembly detection, which relates to the technical field of image processing, and comprises the following steps: step 1, collecting a fuse box image, cutting m types of fuses from the image and n images for each type, and then carrying out image scale normalization operation; and step 2, carrying out CLIP model migration training. And step 3, making a detection rule. And step 4, cutting out a detected target and a corresponding standard image. And 5, judging the type of the cut image by the CLIP migration model. And step 6, judging the type of the cut image according to histogram similarity. And 7, performing cosine similarity judgment to cut out the type of the detection image. According to the method, a deep learning method and a traditional image processing method are combined, the advantages of the deep learning method and the traditional image processing method are integrated, the problem of poor recognition caused by poor light, color fading, fading, stains and the like is well solved, and better performance and higher robustness are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to a method for fusing Clip for automotive fuse box assembly detection with traditional image processing. Background Art

[0002] The technology for automotive fuse box assembly detection is a key technology related to automotive safety. It can effectively detect and verify the correct assembly of the fuse box and its internal components to ensure the reliability and safety of the automotive electrical system.

[0003] Traditional methods for automotive fuse box assembly detection include manual visual inspection, which has human errors: manual inspection is easily affected by the subjective judgment and fatigue of the operator, resulting in inconsistencies and errors. Time consumption: Manual inspection takes a lot of time, especially in large-scale automotive production; it is costly; in addition, it is difficult to achieve product traceability and other problems. With the development of computers, video surveillance technology has become more and more popular, and automotive fuse box assembly detection by means of image processing has become a solution. Currently, common processing methods include color contrast, character recognition, Sift, Surf feature point matching, deep learning and other methods.

[0004] Selecting important features from images by these traditional computer vision methods is a necessary step, and feature extraction depends on manual judgment and long-term trial and error. This way of extracting features is not universal, resulting in the defect of poor robustness. Specifically, the color contrast method is very sensitive to good lighting conditions. Insufficient lighting or strong light reflection will affect the detection performance and limit the scope of use; while feature point methods such as Sift and Surf are more sensitive to transformations such as lighting, rotation, and scaling, and the matching accuracy is low; the character recognition method cannot distinguish fuse types without text, and the recognition rate for positive and negative is not high. At the same time, in the face of problems such as dirt, fading, and color loss on the fuse piece, the detection effect drops significantly. Deep learning performs well in terms of performance, but still faces problems such as the need for a large amount of labeled data, high computational costs, and poor interpretability. Summary of the Invention

[0005] The technical solution of the present invention aims at the technical problem that the existing technical solutions are too single, and provides a solution significantly different from the existing technology. Specifically, the purpose of the present invention is to provide a method for fusing Clip for automotive fuse box assembly detection with traditional image processing, so as to solve the problems that the existing technology detection is very sensitive to light conditions, the selection of color thresholds depends on manual experience, the recognition effect is not good for fading and pollution, the robustness is poor, and deep learning performs well in terms of performance, but still faces the problems of the need for a large amount of labeled data, high computational costs, and poor interpretability.

[0006] To achieve the above object, the present invention provides the following technical solution: A method for fusing Clip with traditional image processing for the assembly detection of an automotive fuse box, comprising the following steps:

[0007] Step 1, Collect fuse box images, crop m types of fuses from the images, with n images for each type, and then perform image scale normalization operation;

[0008] Step 2, CLIP model transfer training: Load the pre-trained CLIP model, adopt the contrastive loss function, select the optimizer Adam and set the learning rate to 1e-6. On the cropped fuse image dataset, input the image and the corresponding fuse label text into the CLIP model, calculate the similarity between the image and text features output by the model, and optimize the model parameters through the contrastive loss function until the model converges on the training set to determine the corresponding network model;

[0009] Step 3, Formulate detection rules: Take the correctly installed fuse box image as a reference, mark the positions of all fuses, and then give the fuse image category label for each position;

[0010] Step 4, Crop the detected object and the corresponding standard image: Take a photo of the fuse box assembled on the production line, and crop the image of each position from the photo according to the positions marked in the reference image;

[0011] Step 5, The CLIP transfer model determines the type of the cropped image: Input the cropped detected image into the image encoder of the CLIP transfer model to obtain the image embedding vector. At the same time, input the fuse type text label into the text encoder of the transfer CLIP model to obtain the text embedding vector; Calculate the cosine similarity between the detected image embedding vector and the embedding vectors of each type of text label, and classify the detected image into the type text label with the highest cosine similarity; When the highest cosine similarity is greater than a certain threshold, if the type text label with the highest cosine similarity is consistent with the fuse type at the same position of the standard image, it is correctly installed, otherwise, it is incorrectly installed and an error prompt is given; When the highest cosine similarity is less than a certain threshold, go to Step 6;

[0012] Step 6, Histogram similarity determines the type of the cropped image: First, divide the cropped image into left and right half images, and at the same time divide the corresponding image of the reference image into left and right half images; Calculate the histogram similarity between the left half image of the cropped image and the left half image of the corresponding image of the reference image, and at the same time calculate the histogram similarity between the right half image of the cropped image and the right half image of the corresponding image of the reference image; When the left half histogram similarity and the right half histogram similarity cannot be greater than a certain threshold at the same time, the type is inconsistent with the target type, it is incorrectly installed and an error prompt is given; Otherwise, go to Step 7;

[0013] Step 7, Cosine similarity judgment to crop the type of the detected image: Use resnet18 and remove the last fully connected layer, and obtain the feature vectors of the cropped detected image, the standard image at the corresponding position, and the standard image flipped 180 degrees respectively; Calculate the cosine similarity between the feature vectors of the cropped detected image and the standard image at the corresponding position, and at the same time calculate the cosine similarity between the cropped detected image and the feature vector of the standard image flipped 180 degrees at the corresponding position; If the cosine similarity with the feature vector of the standard image is greater than the cosine similarity with the feature vector of the standard image flipped 180 degrees, then the type of the fuse at the same position as the reference image is the same, and it is correctly installed; If the cosine similarity with the feature vector of the standard image is less than the cosine similarity with the feature vector of the standard image flipped 180 degrees, then the type of the fuse at the same position as the reference image is inconsistent, and it is incorrectly installed and an error prompt is given at the same time.

[0014] Preferably, in step 1, the specific steps of data acquisition are:

[0015] Step 1.1, First, collect the fuse box image, crop out various types of fuse images from the fuse box image, assume there are m types of fuses, and n images for each type of fuse;

[0016] Step 1.2, Normalize the image scale, and divide the image dataset into a training set and a test set for testing.

[0017] Preferably, in step 2, the specific steps of CLIP model transfer training are:

[0018] Step 2.1, Load the pre-trained CLIP model (using the ViT-B / 32 model architecture), adopt the contrastive loss function, select the optimizer Adam and set the learning rate to 1e-6;

[0019] Step 2.2, On the cropped fuse image dataset, input the image and the corresponding fuse label text into the CLIP model, calculate the similarity between the image and text features output by the model, and optimize the model parameters through the contrastive loss function until the model converges on the training set to determine the corresponding network model.

[0020] Preferably, in step 3, the specific steps of formulating the detection rules are:

[0021] Step 3.1, Put the correctly installed fuse box into the system for taking pictures, and use this photo as a reference image;

[0022] Step 3.2, In the reference image, manually mark the positions and types of all fuses.

[0023] Preferably, in step 4, the specific steps for cropping the detected image and the corresponding standard image are as follows:

[0024] Step 4.1, take a photo of the fuse box on the production line;

[0025] Step 4.2, according to the positions marked in the reference image, crop out the detected images at each position from the photo and the standard images at the same positions from the reference image respectively.

[0026] Preferably, in step 5, the specific steps for the CLIP transfer model to determine the type of the cropped detected image are as follows:

[0027] Step 5.1, input the cropped detected image into the image encoder of the CLIP transfer model to obtain an image embedding vector, and at the same time input the fuse type text label into the text encoder of the transfer CLIP model to obtain a text embedding vector;

[0028] Step 5.2, calculate the cosine similarity between the detected image embedding vector and each type of text label embedding vector, and classify the detected image into the type of text label with the highest cosine similarity;

[0029] Step 5.3, when the highest cosine similarity is greater than a certain threshold, if the type of text label with the highest cosine similarity is consistent with the fuse type at the same position in the reference image, it is correctly installed, otherwise, it is incorrectly installed and an error prompt is given; when the highest cosine similarity is less than a certain threshold, go to step 6.

[0030] Preferably, in step 6, the specific steps for the histogram similarity to determine the type of the cropped detected image are as follows:

[0031] Step 6.1, first divide the cropped detected image into left and right half images, and at the same time divide the standard image at the corresponding position into left and right half images;

[0032] Step 6.2, calculate the histogram similarity between the left half image of the cropped detected image and the left half image of the corresponding standard image, and at the same time calculate the histogram similarity between the right half image of the cropped detected image and the right half image of the corresponding standard image;

[0033] Step 6.3, when the left half histogram similarity and the right half histogram similarity cannot both be greater than a certain threshold, the type is inconsistent with the fuse type at the same position in the reference image, it is incorrectly installed and an error prompt is given; otherwise, go to step 7.

[0034] Preferably, in step 7, the specific steps for the cosine similarity to determine the type of the cropped detected image are as follows:

[0035] Step 7.1: Use resnet18 and remove the last fully connected layer to obtain the feature vectors of the cropped detection images, the standard images at the corresponding positions, and the standard images flipped 180 degrees respectively.

[0036] Step 7.2: Calculate the cosine similarity between the feature vectors of the cropped detection images and the standard images at the corresponding positions, and at the same time calculate the cosine similarity between the feature vectors of the cropped detection images and the standard images flipped 180 degrees at the corresponding positions.

[0037] Step 7.3: Calculate the cosine similarity between the feature vectors of the cropped detection images and the standard images at the corresponding positions, and at the same time calculate the cosine similarity between the feature vectors of the cropped detection images and the standard images flipped 180 degrees at the corresponding positions. If the cosine similarity with the feature vectors of the standard images is greater than the cosine similarity with the feature vectors of the standard images flipped 180 degrees, then the types of fuses at the same positions as the reference images are the same, and the installation is correct; if the cosine similarity with the feature vectors of the standard images is less than the cosine similarity with the feature vectors of the standard images flipped 180 degrees, then the types of fuses at the same positions as the reference images are different, the installation is incorrect, and an error message is given.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] 1. The multi-modal deep learning model CLIP uses contrastive learning to achieve image-text matching. By combining the characteristics of fuse box assembly images and descriptive texts, CLIP can determine whether the fuses are assembled in the specified manner. The zero-shot ability of CLIP enables it to have good adaptability in some new and unknown classification tasks.

[0040] 2. For the types with relatively low classification similarity achieved in CLIP, traditional image processing methods are used as supplementary processing. First, histogram similarity discrimination is used to filter out the detection images that are inconsistent with the standard images, and then the similarities between the detection images and the standard images, and between the detection images and the standard images rotated 180 degrees are calculated based on the cosine similarity method, so as to determine whether the installation of the detection images and the standard images is correct or reversed, and complete the detection of fuse assembly. This method can well enhance the adaptability of the deep learning model and improve the interpretability of the deep learning model.

[0041] 3. The present invention realizes the combination of deep learning methods and traditional image processing methods, combines the advantages of these two methods, well solves the problems of poor recognition caused by bad light, fading, discoloration, stains, etc., and achieves better performance, stronger robustness, and improved interpretability of the deep learning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flowchart;

[0043] Figure 2 The maximum value of the cosine similarity between the detected image calculated using the CLIP migration model and the type text label is greater than a certain threshold, and the corresponding type text label is consistent with the fuse type at the same position of the reference image. Correct installation example diagram;

[0044] Figure 3 The maximum value of the cosine similarity between the detected image calculated using the CLIP migration model and the type text label is greater than a certain threshold, and the corresponding type text label is inconsistent with the fuse type at the same position of the reference image. Incorrect installation example diagram;

[0045] Figure 4 The left - half histogram similarity and the right - half histogram similarity cannot be greater than a certain threshold simultaneously. Incorrect installation example diagram;

[0046] Figure 5 The cosine similarity with the standard image feature vector is greater than the cosine similarity with the feature vector of the standard image flipped 180 degrees. Correct installation example diagram;

[0047] Figure 6 The cosine similarity with the standard image feature vector is greater than the cosine similarity with the feature vector of the standard image flipped 180 degrees. Correct installation example diagram;

[0048] Figure 7 The cosine similarity with the standard image feature vector is less than the cosine similarity with the feature vector of the standard image flipped 180 degrees. Incorrect installation example diagram;

[0049] Figure 8 The cosine similarity with the standard image feature vector is less than the cosine similarity with the feature vector of the standard image flipped 180 degrees. Incorrect installation example diagram. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] As Figure 1 shown, a method for fusing Clip and traditional image processing for detecting the assembly of an automotive fuse box disclosed by the present invention includes the following steps:

[0052] Step 1, collect fuse box images, crop m types of fuses from the images, with n images for each type, and then perform image scale normalization operations.

[0053] Step 2, CLIP model transfer training: Load the pre-trained CLIP model (using the ViT-B / 32 model architecture), adopt the contrastive loss function, select the optimizer Adam and set the learning rate to 1e-6. On the cropped fuse image dataset, input the image and the corresponding fuse label text into the CLIP model, calculate the similarity between the image and text features output by the model, and optimize the model parameters through the contrastive loss function until the model converges on the training set to determine the corresponding network model.

[0054] Step 3, Formulate detection rules: Using the image of the correctly installed fuse box as a reference, mark the positions of all fuses, and then give the fuse image category labels for each position.

[0055] Step 4, Crop the detected targets and corresponding standard images: Take pictures of the fuse boxes assembled on the production line, and crop the images of each position from the pictures according to the positions marked in the reference image.

[0056] Step 5, The CLIP transfer model determines the type of the cropped image: Input the cropped image into the image encoder of the CLIP transfer model to obtain the image embedding vector. At the same time, input the fuse type text label into the text encoder of the transfer CLIP model to obtain the text embedding vector; calculate the cosine similarity between the image embedding vector and the embedding vectors of each type of text label, and classify the image into the type text label with the highest cosine similarity; when the highest cosine similarity is greater than a certain threshold, if the type text label with the highest cosine similarity is consistent with the fuse type at the same position in the reference image, it is correctly installed, otherwise it is incorrectly installed and an error prompt is given; when the highest cosine similarity is less than a certain threshold, go to Step 6.

[0057] Step 6, Histogram similarity determines the type of the cropped image: First, divide the cropped image into left and right half-images, and at the same time divide the corresponding image of the reference image into left and right half-images; calculate the histogram similarity between the left half-image of the cropped image and the left half-image of the corresponding image of the reference image, and at the same time calculate the histogram similarity between the right half-image of the cropped image and the right half-image of the corresponding image of the reference image; when the left half histogram similarity and the right half histogram similarity cannot both be greater than a certain threshold, the type is inconsistent with the target type, it is incorrectly installed and an error prompt is given; otherwise, go to Step 7.

[0058] Step 7, Cosine Similarity Judgment to Determine the Type of the Cropped Detection Image: Use resnet18 and remove the last fully connected layer. Obtain the feature vectors of the cropped detection image, the standard image at the corresponding position, and the standard image flipped 180 degrees respectively. Calculate the cosine similarity between the feature vector of the cropped detection image and the feature vector of the standard image at the corresponding position, and at the same time calculate the cosine similarity between the feature vector of the cropped detection image and the feature vector of the standard image flipped 180 degrees at the corresponding position. If the cosine similarity with the feature vector of the standard image is greater than the cosine similarity with the feature vector of the standard image flipped 180 degrees, it means that the type of the fuse at the same position as the reference image is the same, and it is correctly installed. If the cosine similarity with the feature vector of the standard image is less than the cosine similarity with the feature vector of the standard image flipped 180 degrees, it means that the type of the fuse at the same position as the reference image is inconsistent, and it is incorrectly installed and an error message is given.

[0059] In this embodiment, in step 1, the specific steps of data acquisition are as follows:

[0060] Step 1.1, First, collect the fuse box image, and crop out various types of fuse images from the fuse box image. Assume that there are m types of fuses, and there are n images for each type of fuse.

[0061] Step 1.2, Normalize the image scale, and divide the image dataset into a training set and a test set.

[0062] In this embodiment, in step 2, the specific steps of the CLIP model transfer training are as follows.

[0063] Step 2.1, Load the pre-trained CLIP model (using the ViT-B / 32 model architecture), adopt the contrastive loss function, select the optimizer Adam and set the learning rate to 1e-6.

[0064] Step 2.2, On the cropped fuse image dataset, input the image and the corresponding fuse label text into the CLIP model, calculate the similarity between the image and text features output by the model, and optimize the model parameters through the contrastive loss function until the model converges on the training set to determine the corresponding network model.

[0065] In this embodiment, in step 3, the specific steps of formulating the detection rules are as follows:

[0066] Step 3.1, Put the correctly installed fuse box into the system for taking pictures, and use this picture as a reference image.

[0067] Step 3.2, In the reference image, manually mark the positions and types of all fuses.

[0068] In this embodiment, in step 4, the specific steps for cropping the detection image and the corresponding standard image are as follows:

[0069] Step 4.1, take a photo of the fuse box on the production line.

[0070] Step 4.2, according to the positions marked in the reference image, crop the detection image at each position from the photo and the standard image at the same position from the reference image respectively.

[0071] In this embodiment, in step 5, the specific steps for the CLIP transfer model to determine the type of the cropped detection image are as follows:

[0072] Step 5.1, input the cropped detection image into the image encoder of the CLIP transfer model to obtain an image embedding vector, and at the same time input the fuse type text label into the text encoder of the transfer CLIP model to obtain a text embedding vector.

[0073] Step 5.2, calculate the cosine similarity between the detection image embedding vector and each type of text embedding vector, and classify the detection image into the type text label with the highest cosine similarity.

[0074] Step 5.3, when the highest cosine similarity is greater than a certain threshold, if the type text label with the highest cosine similarity is consistent with the fuse type at the same position in the reference image, it is correctly installed, as Figure 2 shown, otherwise it is incorrectly installed, as Figure 3 shown; when the highest cosine similarity is less than a certain threshold, go to step 6.

[0075] In this embodiment, in step 6, the specific steps for the histogram similarity to determine the type of the cropped detection image are as follows:

[0076] Step 6.1, first divide the cropped detection image into left and right half-images, and at the same time divide the standard image at the corresponding position into left and right half-images.

[0077] Step 6.2, calculate the histogram similarity between the left half-image of the cropped detection image and the left half-image of the corresponding standard image, and at the same time calculate the histogram similarity between the right half-image of the cropped detection image and the right half-image of the corresponding standard image.

[0078] Step 6.3, when the left half histogram similarity and the right half histogram similarity cannot be greater than a certain threshold at the same time, the type is inconsistent with the fuse type at the same position in the reference image, and it is incorrectly installed and an error prompt is given, as Figure 4 shown; otherwise, go to step 7.

[0079] In this embodiment, in step 7, the specific steps for the image cosine similarity judgment to crop out the type of the detected image are as follows:

[0080] Step 7.1: Calculate the cosine similarity between the cropped detected image and the standard image at the corresponding position, and at the same time calculate the cosine similarity between the cropped detected image and the image obtained by flipping the standard image at the corresponding position by 180 degrees.

[0081] Step 7.2: If the cosine similarity with the standard image is greater than the cosine similarity with the image obtained by flipping the standard image by 180 degrees, then the types of the fuses at the same position as the reference image are the same, and it is correctly installed, as Figure 5 , Figure 6 shown; if the cosine similarity with the standard image is less than the cosine similarity with the image obtained by flipping the standard image by 180 degrees, then the types of the fuses at the same position as the reference image are inconsistent, and it is incorrectly installed and an error prompt is given at the same time, as Figure 7 , Figure 8 shown.

[0082] In summary, the present invention realizes the combination of the multi-modal deep learning model CLIP and the traditional image processing method, combines the advantages of these two methods, and well solves the problem of poor recognition caused by bad light, fading, color loss, stains, etc. At the same time, the CLIP multi-modal model can show good classification performance with less annotation. The zero-shot ability of CLIP enables it to have good adaptability in some new and unknown classification tasks, and obtain better interpretability and stronger robustness.

[0083] The CLIP multi-modal model realizes the matching of images and texts by mapping images and texts to the same semantic space and using contrastive learning. In the present invention, CLIP is used to compare an image with a set of text labels (fuse image categories), and the most matching label is selected as the classification result of the fuse image, so as to judge whether the fuse is of the same category as the reference fuse. CLIP can show good classification performance with less annotation; the zero-shot ability of CLIP makes it have good adaptability in some new and unknown classification tasks, and does not require retraining from scratch like traditional image classification models, providing a possibility for future detection of new fuse types; compared with traditional deep learning networks that require a large amount of labeled data for training, CLIP can perform task transfer with less annotation data in many cases. At the same time, for types with low classification confidence of CLIP, traditional image processing methods, namely histogram combined with cosine similarity method of images, are used for supplementary processing. The traditional method does not require a large amount of labeled data, and can still maintain good results when the data is limited after combination; the traditional method is used as a post-processing step to improve the adaptability of the deep learning model in different scenarios; the traditional method is usually easier to interpret, and the interpretability of the deep learning model can be improved after combination.

[0084] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacement on some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A fusion method of Clip for automotive fuse box assembly detection and traditional image processing, characterized in that It includes the following steps: Step 1: Collect the images of the fuse box, crop m types of fuses from the images, with n images for each type, and then perform image scale normalization operation; Step 2: CLIP model transfer training: Load the pre-trained CLIP model, adopt the contrastive loss function, select the optimizer Adam and set the learning rate to 1e-6. On the cropped fuse image dataset, input the image and the corresponding fuse label text into the CLIP model, calculate the similarity between the image and text features output by the model, and optimize the model parameters through the contrastive loss function until the model converges on the training set to determine the corresponding network model; Step 3: Formulate detection rules: Based on the image of the correctly installed fuse box, mark the positions of all fuses, and then give the fuse image category labels for each position; Step 4: Crop the detected targets and corresponding standard images: Take pictures of the fuse boxes assembled on the production line, and crop the images of each position from the photos according to the positions marked in the reference image; Step 5: The CLIP transfer model determines the type of the cropped image: Input the cropped detection image into the image encoder of the CLIP transfer model to obtain the image embedding vector. At the same time, input the fuse type text label into the text encoder of the transferred CLIP model to obtain the text embedding vector; Calculate the cosine similarity between the detection image embedding vector and the embedding vectors of each type of text label, and classify the detection image into the type text label with the highest cosine similarity; When the highest cosine similarity is greater than a certain threshold, if the type text label with the highest cosine similarity is consistent with the fuse type at the same position of the standard image, it is correctly installed, otherwise, it is incorrectly installed and an error prompt is given; When the highest cosine similarity is less than a certain threshold, go to Step 6; Step 6: Histogram similarity determines the type of the cropped image: First, divide the cropped image into left and right half images, and at the same time divide the corresponding image of the reference image into left and right half images; Calculate the histogram similarity between the left half image of the cropped image and the left half image of the corresponding image of the reference image, and at the same time calculate the histogram similarity between the right half image of the cropped image and the right half image of the corresponding image of the reference image; When the left half histogram similarity and the right half histogram similarity cannot be greater than a certain threshold at the same time, the type is inconsistent with the target type, it is incorrectly installed and an error prompt is given; Otherwise, go to Step 7; Step 8: Cosine similarity determines the type of the cropped detection image: Use resnet18 and remove the last fully connected layer to obtain the feature vectors of the cropped detection image, the standard image at the corresponding position, and the standard image flipped 180 degrees respectively; Calculate the cosine similarity between the feature vectors of the cropped detection image and the standard image at the corresponding position, and at the same time calculate the cosine similarity between the feature vectors of the cropped detection image and the standard image flipped 180 degrees at the corresponding position; If the cosine similarity with the standard image feature vector is greater than the cosine similarity with the feature vector of the standard image flipped 180 degrees, the types of fuses at the same positions as the reference image are the same, and it is correctly installed; if the cosine similarity with the standard image feature vector is less than the cosine similarity with the feature vector of the standard image flipped 180 degrees, the types of fuses at the same positions as the reference image are different, and it is incorrectly installed and an error prompt is given simultaneously.

2. A method for fusing Clip in the assembly inspection of an automotive fuse box with traditional image processing according to claim 1, characterized in that: In step 1, the specific steps of data acquisition are as follows: Step 1.1, first collect the fuse box image, crop out various types of fuse images from the fuse box image. Assume there are m types of fuses, and n images for each type of fuse. Step 1.2, normalize the image scale, and divide the image dataset into a training set and a test set.

3. A method for fusing Clip in the assembly inspection of an automotive fuse box with traditional image processing according to claim 1, characterized in that: In step 2, the specific steps of the CLIP model transfer training are as follows: Step 2.1, load the pre-trained CLIP model (using the ViT-B / 32 model architecture), adopt the contrastive loss function, select the optimizer Adam and set the learning rate to 1e-6. Step 2.2, on the cropped fuse image dataset, input the image and the corresponding fuse label text into the CLIP model, calculate the similarity between the image and text features output by the model, and optimize the model parameters through the contrastive loss function until the model converges on the training set to determine the corresponding network model.

4. A method for fusing Clip in the assembly detection of an automotive fuse box with traditional image processing according to claim 1, characterized in that: In step 3, the specific steps of formulating the detection rules are as follows: Step 3.1, put the correctly installed fuse box into the system for taking a photo, and use this photo as a reference image. Step 3.2, in the reference image, manually label the positions and types of all fuses.

5. A method for fusing Clip in the assembly detection of an automotive fuse box with traditional image processing according to claim 1, characterized in that: In step 4, the specific steps of cropping the detection image and the corresponding standard image are as follows: Step 4.1, take a photo of the fuse box on the production line. Step 4.2, according to the positions marked in the reference image, crop out the detection image at each position from the photo and the standard image at the same position from the reference image respectively.

6. A method for fusing Clip in the assembly inspection of an automotive fuse box with traditional image processing according to claim 1, characterized in that: In step 5, the specific steps of the CLIP transfer model for judging the type of the cropped detection image are as follows: Step 5.1, input the cropped detection image into the image encoder of the CLIP transfer model to obtain the image embedding vector, and at the same time input the fuse type text label into the text encoder of the transfer CLIP model to obtain the text embedding vector. Step 5.2, calculate the cosine similarity between the detection image embedding vector and the embedding vectors of each type of text label, and classify the detection image into the type text label with the highest cosine similarity. Step 5.3, when the highest cosine similarity is greater than a certain threshold, if the type text label with the highest cosine similarity is the same as the type of the fuse at the same position in the reference image, it is correctly installed; otherwise, it is incorrectly installed and an error prompt is given. When the highest cosine similarity is less than a certain threshold, go to step 6.

7. A method for fusing Clip in the assembly inspection of an automotive fuse box shown in claim 1 with traditional image processing, characterized in that: In step 6, the specific steps of the histogram similarity for judging the type of the cropped detection image are as follows: Step 6.1, first divide the cropped detection image into left and right half images, and at the same time divide the standard image at the corresponding position into left and right half images. Step 6.2: Calculate the histogram similarity between the left - hand image of the cropped detection image and the left - hand image of the corresponding standard image, and at the same time, calculate the histogram similarity between the right - hand image of the cropped detection image and the right - hand image of the corresponding standard image; Step 6.3: When the left - hand histogram similarity and the right - hand histogram similarity cannot both be greater than a certain threshold, it means that the fuse type at the same position as the reference image is inconsistent, and it is misinstalled, and an error prompt is given at the same time; otherwise, go to Step 7.

8. A fusion method of Clip for automotive fuse box assembly detection shown in claim 1 and traditional image processing, characterized in that: In Step 7, the specific steps for judging the type of the cropped detection image by cosine similarity are as follows: Step 7.1: Use resnet18 and remove the last fully - connected layer to obtain the feature vectors of the cropped detection image, the standard image at the corresponding position, and the standard image flipped 180 degrees respectively; Step 7.2: Calculate the cosine similarity between the feature vectors of the cropped detection image and the standard image at the corresponding position, and at the same time, calculate the cosine similarity between the feature vectors of the cropped detection image and the standard image flipped 180 degrees at the corresponding position; Step 7.3: Calculate the cosine similarity between the feature vectors of the cropped detection image and the standard image at the corresponding position, and at the same time, calculate the cosine similarity between the feature vectors of the cropped detection image and the standard image flipped 180 degrees at the corresponding position; If the cosine similarity with the feature vector of the standard image is greater than the cosine similarity with the feature vector of the standard image flipped 180 degrees, it means that the fuse type at the same position as the reference image is consistent, and it is correctly installed; if the cosine similarity with the feature vector of the standard image is less than the cosine similarity with the feature vector of the standard image flipped 180 degrees, it means that the fuse type at the same position as the reference image is inconsistent, and it is misinstalled, and an error prompt is given at the same time.

Citation Information

Patent Citations

  • Vehicle-mounted-fuse-box relay-installing error-proof recognition device special for production line

    CN107132476A

  • Cigarette defect detection method based on deep transfer learning

    CN113034483A

  • Visual implementation method for measuring assembly height and inclination degree of device in automobile fuse box

    CN113624145A

  • CLIP model-based geological disaster image recognition method and system

    CN117079048A

  • Detection method for automobile fuse box assembly based on machine vision

    CN117893495A