An Optimization Method and System for Visual Inspection of Mobile Phone Glass Covers for Edge Devices

By optimizing the student model structure and distillation temperature on edge devices, the deployment problem of deep learning models on edge devices is solved, and efficient and accurate detection of mobile phone glass cover defects is achieved, reducing the complexity and cost of the production line.

CN119559122BActive Publication Date: 2025-07-08GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411494556.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-07-08
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deploy deep learning machine vision detection models on edge devices, resulting in high complexity and cost of production line layout, and insufficient computing power of edge devices cannot run the server model.

Method used

By obtaining edge device performance information, establishing student models, and using knowledge distillation technology to train student models, determine distillation temperature based on target defect differences, optimize the model structure to adapt to edge devices, including adjusting the number of fully connected layers and the number of neurons, and optimizing the distillation process in combination with historical application models.

Benefits of technology

It realizes efficient and accurate deep learning machine vision detection on edge devices, reduces the complexity and cost of production line layout, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559122B_ABST
    Figure CN119559122B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of visual inspection of defects in mobile phone glass covers, and particularly to a visual inspection optimization method and system for mobile phone glass covers for edge devices. First, a student model is established according to the performance information of the target edge device, and then the target distillation temperature is determined according to the differences of various target defects. The knowledge of a preset teacher model is distilled based on the target distillation temperature to train the student model, and the trained student model is obtained. Finally, the trained student model is deployed in the target edge performance device. Compared with the prior art, the present invention establishes a student model that conforms to its computing power according to the performance information of the target edge device, then trains the student model through knowledge distillation technology to make it close to the performance of the preset teacher model, and controls the temperature during knowledge distillation according to the differences of various target defects during the distillation process, solving the problem of how to deploy the deep learning machine vision detection model of mobile phone glass covers on edge devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual inspection of mobile phone glass covers, and particularly to a visual inspection optimization method and system for mobile phone glass covers for edge devices. Background Art

[0002] With the development of deep learning technology, machine vision inspection technology has been widely applied in the field of defect inspection of mobile phone glass covers. Deep learning machine vision inspection technology mainly relies on the image recognition ability of neural networks. By training a large number of defective and non-defective samples, the model learns to automatically distinguish and identify defect types. Currently, most defect detection systems will send the collected image data to the server for analysis and processing. Although this approach can achieve high detection accuracy, it also increases the layout difficulty of the mobile phone glass cover production line because it requires a stable network connection, sufficient data transmission bandwidth, and the maintenance of expensive servers, resulting in high costs.

[0003] Edge devices refer to devices or servers that perform data processing near the data source, which can be industrial gateways, intelligent cameras, or other devices with computing capabilities. Deploying the deep learning model on edge devices can significantly reduce the latency of data transmission, reduce the dependence on network bandwidth, and improve data security and privacy protection. In this way, defect detection can be directly carried out on the production line without sending data to a remote server, thereby reducing the complexity and cost of the production line layout.

[0004] However, the computing power of edge devices is often inferior to that of servers, and the defect recognition model running on the server may not be able to run on edge devices. Therefore, people need a visual inspection optimization method for mobile phone glass covers for edge devices to achieve deep learning machine vision inspection on edge devices. Summary of the Invention

[0005] Therefore, the present invention provides a visual inspection optimization method and system for mobile phone glass covers for edge devices to solve the problem in the prior art of how to deploy the deep learning machine vision inspection model of mobile phone glass covers on edge devices.

[0006] The present invention provides a visual inspection optimization method for mobile phone glass covers for edge devices, including:

[0007] Obtaining the performance information of the target edge device and the defect information of multiple target defects;

[0008] Establishing a student model according to the performance information;

[0009] Obtain a preset teacher model and preset training data, where both the preset teacher model and the student model are neural network models for detecting defects in mobile phone glass covers based on images;

[0010] Determine the target distillation temperature according to the differences between multiple target defects characterized by multiple defect information;

[0011] Based on the target distillation temperature, perform knowledge distillation on the preset teacher model according to the preset training data and train the student model to obtain the trained student model;

[0012] Deploy the trained student model to the target edge performance device.

[0013] The present invention also provides a preferred solution: the defect information includes defect sample pictures, and each defect sample picture corresponds to one target defect; determining the target distillation temperature according to the differences between multiple target defects characterized by multiple defect information includes:

[0014] Analyze the similarity between every two defect sample pictures pairwise to obtain multiple similarity feature values;

[0015] Obtain the target distillation temperature according to the distribution of multiple similarity feature values, where the more similar target defects characterized by multiple similarity feature values, the smaller the target distillation temperature.

[0016] The present invention also provides a preferred solution: analyzing the similarity between every two defect sample pictures pairwise to obtain multiple similarity feature values includes:

[0017] Obtain two target defect sample pictures, where the target defect sample pictures are two defect sample pictures for which the similarity is to be analyzed currently;

[0018] Convert the two target defect sample pictures into two picture vectors respectively, where each element in the picture vector is a pixel value in the target defect sample picture, and the pixel positions corresponding to the elements at the same position in the two picture vectors are the same;

[0019] Calculate the cosine similarity of the two picture vectors and obtain the similarity feature value of the two target defect sample pictures according to the cosine similarity.

[0020] The present invention also provides a preferred solution: the larger the value of the similarity feature value, the more similar the two defect sample pictures corresponding to it; obtaining the target distillation temperature according to the distribution of multiple similarity feature values includes:

[0021] Count the total number of types of multiple target defects;

[0022] Calculate the sum of multiple similarity feature values;

[0023] Calculate the statistical eigenvalue of multiple similar eigenvalues, where the statistical eigenvalue is used to represent the degree of dispersion of multiple similar eigenvalues;

[0024] Based on the total number of types of multiple target defects, the sum of multiple similar eigenvalues, and the statistical eigenvalue of multiple similar eigenvalues, obtain the target distillation temperature, where the target distillation temperature is inversely proportional to the total number of types of multiple target defects, inversely proportional to the sum of multiple similar eigenvalues, and directly proportional to the degree of dispersion characterized by the statistical eigenvalue of multiple similar eigenvalues.

[0025] The present invention also provides a preferred solution: Based on the target distillation temperature, perform knowledge distillation on a preset teacher model according to preset training data and train a student model to obtain a trained student model, including:

[0026] Obtain a historical application model, where the historical application model is a neural network model for detecting defects in mobile phone glass covers based on images that has been actually applied in the past;

[0027] Based on the target distillation temperature, establish a first reconstruction loss function between the preset teacher model and the student model;

[0028] Based on the target distillation temperature, establish a second reconstruction loss function between the historical application model and the student model;

[0029] Establish a supervised loss function for the student model;

[0030] According to the first reconstruction loss function, the second reconstruction loss function, and the supervised loss function, establish a target loss function;

[0031] Based on the target loss function, train the student model according to the preset training data to obtain a trained student model.

[0032] The present invention also provides a preferred solution: According to the first reconstruction loss function, the second reconstruction loss function, and the supervised loss function, establish a target loss function, including:

[0033] Statistically calculate the type overlap degree of the defects to be recognized by the historical application model and the student model;

[0034] According to the type overlap degree, determine a weight adjustment parameter, where the weight adjustment coefficient is used to adjust the influence proportion of the second reconstruction loss function on the target loss function;

[0035] According to the preset weight and the weight adjustment coefficient, establish a target loss function.

[0036] The present invention also provides a preferred solution: The target loss function is:

[0037]

[0038] Among them, L is the target loss function, is the first reconstruction loss function, is the second reconstruction loss function, L hard is the supervision loss function, α and β are respectively different preset weights, ε is the weight adjustment parameter, the coincidence degree characterized by the species coincidence degree is inversely proportional to the weight adjustment parameter, and when the species coincidence degree takes the maximum value, the weight adjustment parameter is zero.

[0039] The present invention also provides a preferred solution: according to the performance information, establish a student model, including:

[0040] According to the performance information, obtain the computing power level characteristic value of the target edge device;

[0041] According to the computing power level characteristic value, obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer;

[0042] According to the number of fully connected layers and the number of neurons in each fully connected layer, establish a student model.

[0043] The present invention also provides a preferred solution: according to the computing power level characteristic value, obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer, including:

[0044] Obtain a preset operation amount characteristic function, which is used to obtain an operation amount characteristic value according to the number of fully connected layers and the number of neurons in each fully connected layer, and the operation amount characteristic value is used to characterize the operation complexity of the student model established based on the number of fully connected layers and the number of neurons;

[0045] Establish a fitness function based on the difference between the computing power level characteristic value and the preset operation amount characteristic function;

[0046] Taking the minimum difference between the computing power level characteristic value and the operation amount characteristic value as the optimization goal, based on the preset optimization algorithm, obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer.

[0047] The present invention also provides a vision detection optimization system for mobile phone glass covers for edge devices, including:

[0048] A data input module, which is used to obtain the performance information of the target edge device and the defect information of various target defects;

[0049] A model establishment module, which is used to establish a student model according to the performance information;

[0050] A data preparation module, which is used to obtain a preset teacher model and preset training data, wherein the preset teacher model and the student model are both neural network models for detecting defects of mobile phone glass covers based on images;

[0051] A distillation preparation module, configured to determine a target distillation temperature according to the differences between multiple target defects characterized by multiple defect information;

[0052] A model optimization module, configured to perform knowledge distillation on a preset teacher model based on the target distillation temperature according to preset training data and train a student model to obtain a trained student model;

[0053] A model deployment module, configured to deploy the trained student model to a target edge performance device.

[0054] The beneficial effects of adopting the above embodiments are as follows:

[0055] The present invention provides a method and system for optimizing the visual inspection of mobile phone glass covers for edge devices. First, it obtains the performance information of the target edge device and the defect information of multiple target defects, then establishes a student model according to the performance information, then obtains a preset teacher model and preset training data, and then determines the target distillation temperature according to the differences between multiple target defects characterized by multiple defect information. Based on the target distillation temperature, knowledge distillation is performed on the preset teacher model according to the preset training data and the student model is trained to obtain a trained student model. Finally, the trained student model is deployed to the target edge performance device. Compared with the prior art, the present invention establishes a student model that conforms to its computing power according to the performance information of the target edge device, then trains the student model through knowledge distillation technology to make it close to the performance of the preset teacher model, and controls the temperature during knowledge distillation according to the differences between multiple target defects during the distillation process, realizing the efficient and accurate optimization of the preset teacher model for the target edge device and the actual production situation, and solving the problem of how to deploy the deep learning machine vision detection model of the mobile phone glass cover to the edge device. Description of the Drawings

[0056] Figure 1 It is a flowchart of a method according to an embodiment of the method for optimizing the visual inspection of mobile phone glass covers for edge devices provided by the present invention;

[0057] Figure 2 is Figure 1 a specific step diagram of step S104 in

[0058] Figure 3 is Figure 1 a specific step diagram of step S105 in

[0059] Figure 4 It is a schematic diagram of knowledge distillation in an embodiment of the method for optimizing the visual inspection of mobile phone glass covers for edge devices provided by the present invention;

[0060] Figure 5This is the system structure diagram of an embodiment of the mobile phone glass cover visual detection optimization system for edge devices provided by the present invention. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0062] Combined with Figure 1 As shown, a specific embodiment of the present invention discloses a method for optimizing visual detection of mobile phone glass covers for edge devices, including:

[0063] S101. Obtain the performance information of the target edge device and the defect information of multiple target defects;

[0064] S102. Establish a student model according to the performance information;

[0065] S103. Obtain a preset teacher model and preset training data, where both the preset teacher model and the student model are neural network models for detecting defects of mobile phone glass covers based on images;

[0066] S104. Determine the target distillation temperature according to the differences of multiple target defects characterized by multiple defect information;

[0067] S105. Based on the target distillation temperature, perform knowledge distillation on the preset teacher model according to the preset training data and train the student model to obtain a trained student model;

[0068] S106. Deploy the trained student model in the target edge performance device.

[0069] Compared with the prior art, the present invention establishes a student model that conforms to its computing power according to the performance information of the target edge device, and then trains the student model through knowledge distillation technology to make it close to the performance of the preset teacher model, and controls the temperature during knowledge distillation according to the differences of multiple target defects during the distillation process, realizing the efficient and accurate optimization of the preset teacher model for the target edge device and the actual production situation, and solving the problem of how to deploy the deep learning machine vision detection model of the mobile phone glass cover on the edge device.

[0070] The target edge device in the above process refers to the device on which the deep learning machine vision detection model is to be deployed, which can be any device with computing power such as a newly configured computer or intelligent camera in the mobile phone glass cover plate production line. The target defect is the defect of the mobile phone glass cover plate that needs to be detected by the target edge device, such as scratches, pits, stains, chipped edges, etc. Obviously, according to the actual situation, the defects to be detected by different edge devices may be different.

[0071] Knowledge Distillation is a model compression technique that allows a small and efficient model (student model) to learn the knowledge of a large and complex model (teacher model). This method can improve the performance of the student model, making it close to the performance of the teacher model while maintaining computational efficiency.

[0072] The teacher model is a well-trained and high-performance large neural network. This model is usually trained on a large amount of data to achieve a high accuracy rate. The role of the teacher model is to provide a source of "knowledge", that is, its output (including the predicted probability distribution) contains rich information, and this information can help the student model learn the complex features of the data. In the present invention, the preset teacher model is a model that has been pre-trained to accurately and comprehensively identify the defects of mobile phone glass cover plates.

[0073] The student model is a neural network with a simpler structure and fewer parameters. The purpose of the student model is to learn the knowledge of the teacher model so as to achieve the performance of the teacher model as much as possible while maintaining a small model size. The student model is usually trained on the same data set, but uses the output of the teacher model as additional guiding information.

[0074] Obviously, in the present invention, the specific structure of the student model needs to be determined according to the hardware capabilities of the target edge device and the actual working requirements. In practice, any existing method can be used to determine the structure of the student model as needed. For example, the structure of the student model can be determined manually according to experience based on factors such as the hardware parameters of the target edge device, the technological position where it is located, and the number of tasks during actual operation, or a more complex method can be used to perform pruning and other processing on the preset teacher model according to the above factors to obtain the structure of the student model.

[0075] The present invention provides a preferred method. In a preferred embodiment, the above step S102, establishing a student model according to the performance information, specifically includes:

[0076] Obtaining the computing power level characteristic value of the target edge device according to the performance information;

[0077] Based on the eigenvalue of computing power level, obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer;

[0078] Based on the number of fully connected layers and the number of neurons in each fully connected layer, establish the student model.

[0079] In practice, the YOLO model is generally used for defect detection. When optimizing the YOLO model to adapt to devices with different computing powers, adjusting the fully connected layer is the most effective method. This is because the fully connected layer contains a large number of parameters, accounting for most of the model parameters. By reducing the number of neurons in these layers, the computational complexity and the number of parameters of the model can be significantly reduced, thus reducing the demand for computing power. In addition, the adjustment of the fully connected layer has little impact on the feature extraction ability of the model, which means that optimization can be carried out without sacrificing accuracy and with the same model input size. The optimization of the fully connected layer can also speed up the inference speed of the model on resource-constrained devices, and it is relatively simple to implement without making complex modifications to the convolutional layer. Therefore, adjusting the fully connected layer is the preferred strategy for optimizing the preset teacher model (especially the YOLO model) on different devices.

[0080] In this embodiment, the computing power of the target edge device can be quantitatively represented by the eigenvalue of computing power level (for example, a performance score value is calculated by any existing method according to parameters such as the processor frequency and memory size of the target edge device and used as the eigenvalue of computing power level). In this way, the scale of the fully connected layer can be adjusted based on the eigenvalue of computing power level. For example, arbitrarily take a value less than the number of fully connected layers in the preset teacher model as the number of fully connected layers, perform a linear transformation on the eigenvalue of computing power level to obtain the maximum value of the total number of neurons that the target edge device can theoretically bear, and then evenly distribute this maximum value to these fully connected layers to obtain the number of neurons in each fully connected layer.

[0081] Obviously, it can be seen that the above process needs to determine the number of fully connected layers and multiple numbers of neurons at the same time. In fact, it is a multi-objective optimization problem. Therefore, using optimization algorithms (such as genetic algorithms, particle swarm optimization algorithms, etc.) can more reasonably solve the problem of determining the structure of the student model. The number of neurons in each fully connected layer obtained by the method described above can only be equal, while through the optimization algorithm, a more flexible structural solution with unequal numbers of neurons in each fully connected layer can be obtained.

[0082] Specifically, in a preferred embodiment, the above step: based on the eigenvalue of computing power level, obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer, specifically includes:

[0083] Obtain a preset computational complexity feature function, which is used to obtain a computational complexity feature value according to the number of fully connected layers and the number of neurons in each fully connected layer. The computational complexity feature value is used to characterize the computational complexity of the student model established based on the number of fully connected layers and the number of neurons;

[0084] Establish a fitness function based on the difference between the computing power level feature value and the preset computational complexity feature function;

[0085] Taking the minimum difference between the computing power level feature value and the computational complexity feature value as the optimization goal, obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer based on a preset optimization algorithm.

[0086] The preset computational complexity feature function in the above process can be flexibly determined according to specific situations. For example, the total number of parameters in the fully connected layer part of the current learning model can be calculated according to the number of fully connected layers and the number of neurons in each fully connected layer, and then a relatively simple linear transformation is performed on it as the computational complexity of the current learning model, that is, the computational complexity feature value. In this way, using the difference between the computing power level feature value and the preset computational complexity feature function as the fitness function, the structure of the student model that best matches the computing power of the target edge device can be obtained, maximizing the performance of the target edge device. It can be understood that the optimization algorithm is a prior art. Based on the content described above, how to encode the number of fully connected layers and the number of neurons in each fully connected layer, as well as the specific type of the preset optimization algorithm and the specific optimization process of the preset optimization algorithm can be understood and conceived by those skilled in the art, so no more details will be described in this article.

[0087] Furthermore, the distillation temperature is an important hyperparameter in the knowledge distillation process, which is used to adjust the "softness" of the soft targets output by the teacher model. In knowledge distillation, the output of the teacher model (usually a probability distribution) is used as the target for training the student model, and these targets are called soft targets. The temperature parameter affects the fineness of the knowledge learned by the student model by adjusting the smoothness of the probability distribution of the soft targets. Generally speaking, a lower temperature will make the soft targets closer to the hard targets (that is, the theoretically true and accurate recognition results), which may cause the student model to overly focus on the prediction details of the teacher model, including noise. While a higher temperature will make the soft targets smoother, which helps the student model learn more general features rather than overly relying on the specific predictions of the teacher model.

[0088] It should be emphasized that different from conventional knowledge distillation techniques, in the present invention, the distillation temperature is determined according to the difference degrees of various target defects, which can make the learning of the student model in the present invention more targeted, be able to more precisely extract the useful data in the preset teacher model, and make the training more efficient and accurate.

[0089] For example, if the defects that need to be recognized by the actual student model are black dots, pit points, and stain points, and the binary images of these three types of defects are relatively similar, the probability values corresponding to the three types of defects in the soft targets output by the preset teacher model may have a small difference. In this case, a lower distillation temperature can be used for distillation to amplify the difference in the probability values corresponding to the three types of defects in the soft targets, so that the student model is more inclined to learn the ability to analyze details in the preset teacher model. On the contrary, if the defects that need to be recognized by the actual student model are black dots, orange streaks, and scratches, and the binary images of these three types of defects have a large difference, the probability value differences corresponding to the three types of defects in the soft targets output by the preset teacher model may be already obvious enough. In this case, a higher distillation temperature can be used for distillation, so that the student model is more inclined to comprehensively learn the majority of the general abilities in the teacher model

[0090] Specifically, as shown in Figure 2 In a preferred embodiment, the defect information includes defect sample pictures, and each defect sample picture corresponds to a target defect. On this basis, the above step S104 of determining the target distillation temperature according to the differences between multiple target defects characterized by multiple defect information specifically includes:

[0091] S201. Analyze the similarity of every two defect sample pictures pairwise to obtain multiple similarity feature values;

[0092] S202. Obtain the target distillation temperature according to the distribution of multiple similarity feature values, where the more similar target defects represented by multiple similarity feature values, the smaller the target distillation temperature.

[0093] In the above process, the similarity between two defect sample pictures is represented by similarity feature values to quantify the difference between two target defects. At the same time, the similarity feature values of every two target defects are calculated pairwise, so that the difference between all target defects can be quantified as a whole according to multiple similarity feature values, enabling this embodiment to adapt to more target defect situations.

[0094] In the above process, any existing method can be used to analyze the similarity of every two defect sample pictures. For example, the two defect sample pictures can be subjected to color reduction and superposition, and then the result image is analyzed. If the values of most pixels in the result image are very small, it can be considered that the two images are very similar. Or other existing methods such as the structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), and histogram comparison can be used.

[0095] The present invention provides a preferred method to analyze the similarity of every two defect sample pictures. Specifically, in a preferred embodiment, the above step S201 of analyzing the similarity of every two defect sample pictures pairwise to obtain multiple similarity feature values specifically includes:

[0096] Obtain two target defect sample images, where the target defect sample images are two defect sample images for which the similarity is to be analyzed currently.

[0097] Convert the two target defect sample images into two image vectors respectively. Each element in the image vector is a pixel value in the target defect sample image, and the pixel positions corresponding to the elements at the same position in the two image vectors are the same.

[0098] Calculate the cosine similarity of the two image vectors, and obtain the similarity eigenvalue of the two target defect sample images according to the cosine similarity.

[0099] In this embodiment, the cosine similarity method is used to measure the difference between two target defect sample images. Its advantages are as follows: By converting the image into a vector, the cosine similarity can retain the spatial position information of the pixels, which helps the computer understand the local patterns and structures in the image. At the same time, the calculation of the cosine similarity is relatively simple. For large-scale image datasets, its calculation efficiency is high because it mainly involves dot product and norm calculations. In addition, the cosine similarity is particularly suitable for high-dimensional data and sparse data. For example, when facing the feature vectors of images, it will not be affected by the curse of dimensionality. When facing defect sample images with most pixel values being 0, the cosine similarity can provide better performance than distance-based methods. All these advantages make the cosine similarity particularly suitable for the application scenario of the present invention.

[0100] Similarly, in addition to analyzing the similarity of every two defect sample images, any existing method can be used to analyze the distribution of multiple similarity eigenvalues. For example, assuming that the larger the similarity eigenvalue, the more similar the two images are. At this time, the proportion of similarity eigenvalues exceeding a certain threshold among multiple similarity eigenvalues can be directly counted as the distribution of multiple similarity eigenvalues. The larger this proportion value is, the more similar target defects there are.

[0101] The present invention also provides a preferred method to analyze the distribution of multiple similarity eigenvalues. Specifically, in a preferred embodiment, the larger the value of the similarity eigenvalue, the more similar the two corresponding defect sample images are. On this basis, step S202 above, obtaining the target distillation temperature according to the distribution of multiple similarity eigenvalues, specifically includes:

[0102] Count the total number of types of multiple target defects;

[0103] Calculate the sum of multiple similarity eigenvalues;

[0104] Calculate the statistical eigenvalue of multiple similarity eigenvalues, and the statistical eigenvalue is used to represent the degree of dispersion of multiple similarity eigenvalues;

[0105] Based on the total number of types of multiple target defects, the sum of multiple similarity feature values, and the statistical feature value of multiple similarity feature values, the target distillation temperature is obtained. Among them, the target distillation temperature is inversely proportional to the total number of types of multiple target defects, inversely proportional to the sum of multiple similarity feature values, and directly proportional to the degree of dispersion characterized by the statistical feature value of multiple similarity feature values.

[0106] The significance of the above process lies in jointly analyzing the distribution of multiple similarity feature values through three dimensions: the total number of types of multiple target defects, the sum of multiple similarity feature values, and the statistical feature value of multiple similarity feature values, in order to obtain a more scientific and reasonable target distillation temperature. Among them, the greater the total number of types of target defects, the more content the student model learns. At this time, the target distillation temperature should be reduced to increase the discrimination degree of the probability values corresponding to different defects in the soft target. The sum of multiple similarity feature values is used to represent the similarity degree of multiple target defects as a whole, and the statistical feature value of multiple similarity feature values is used to further analyze the similarity between multiple similarity feature values, that is, to analyze the similarity between the similarities of different defects. Among them, the statistical feature value can be represented in any way according to specific circumstances. For example, the variance of multiple similarity feature values can be directly used as the statistical feature value, or multiple statistical features such as the mean, median, and standard deviation of multiple similarity feature values can be calculated, and these statistical features can be weighted and summed to obtain the statistical feature value.

[0107] A simple example of calculating the target distillation temperature is as follows:

[0108]

[0109] Among them, S represents the target distillation temperature, n represents the total number of types of multiple target defects, sum represents the sum of multiple similarity feature values, sta represents the statistical feature value of multiple similarity feature values, and w1, w2, and w3 are different preset influence weights respectively.

[0110] Furthermore, as shown in Figure 3 In a preferred embodiment, the above step S105, based on the target distillation temperature, performs knowledge distillation on the preset teacher model according to the preset training data and trains the student model to obtain the trained student model, specifically including:

[0111] S301: Obtain a historical application model, which is a neural network model for detecting defects in mobile phone glass covers based on images that has been actually applied in the past;

[0112] S302: Based on the target distillation temperature, establish a first reconstruction loss function between the preset teacher model and the student model;

[0113] S303. Based on the target distillation temperature, establish a second reconstruction loss function between the historical application model and the student model;

[0114] S304. Establish a supervision loss function for the student model;

[0115] S305. Based on the first reconstruction loss function, the second reconstruction loss function, and the supervision loss function, establish a target loss function;

[0116] S306. Based on the target loss function, train the student model according to the preset training data to obtain a trained student model.

[0117] This embodiment is mainly applied to the situation where the products on a production line change, or only the target edge device is updated in the entire production line, and the production line remains unchanged as a whole. In the above process, the historical application model is a deep learning machine vision detection model that has been put into actual use before this re-optimization of the preset teacher model and its deployment on the target edge device. For example, if the factory equipment is upgraded in practice and a new target edge device is purchased to replace the original edge device, the defect detection model running in the original edge device is the historical application model.

[0118] Based on the existing knowledge distillation technology, this embodiment further increases the learning of the historical application model during the distillation process, enabling the student model to not only learn the knowledge of the preset teacher model but also learn the relevant knowledge that conforms to the actual situation in the historical application model (such as the impact of the lighting conditions in the factory on the input image and the output results, and the knowledge that certain specific defects are more likely to occur due to the process design, factory environment, or material characteristics of the raw materials), making the learning model more accurate and practical.

[0119] In the prior art, the reconstruction loss function is used to measure the difference between the output of the student model and the soft target of the teacher model at the set distillation temperature, and the supervision loss function is used to measure the difference between the output of the student model and the true label at the normal temperature (i.e., the temperature value is 1). Among them, the true label, needless to say, has the same definition as any existing deep learning model, and is a vector or value representing the correct output result. The soft target is the probability distribution obtained according to the preset teacher model, which is generally obtained at a certain distillation temperature. The prior art uses the reconstruction loss function and the supervision loss function to jointly train the student model to achieve the effect of learning from the teacher model.

[0120] In this embodiment, the first reconstruction loss function is the same as the reconstruction loss function in the prior art. The definitions of the first reconstruction loss function and the supervision loss function in the present invention are the same as those in the knowledge distillation technology in the prior art. Therefore, no further elaboration will be provided herein. The second reconstruction loss function is the improvement of the present invention and is used to incorporate the historical application model into the learning process. Its concept is actually similar to that of the first reconstruction loss function. For example, an example of the second reconstruction loss function using the cross-entropy function as the loss measurement method is as follows:

[0121]

[0122] where T is the target distillation temperature, represents the second reconstruction function, which can be understood as the loss of the output result of the learning model compared to the soft target of the output of the historical application model under the condition of the target distillation temperature. i represents the serial number corresponding to the target defect, represents the value corresponding to the i-th target defect in the soft target output by the historical application model under the target distillation temperature (since the historical application model and the preset teacher model play the same teaching role, the probability distribution obtained by the historical application model at a certain distillation temperature is also called the soft target in this embodiment), represents the value corresponding to the i-th target defect in the output result of the student model under the target distillation temperature, and log is the logarithm symbol.

[0123] Furthermore, the soft label output by the historical application model is specifically:

[0124]

[0125] where k also represents the serial number corresponding to the target defect, and v k is the logical value corresponding to the k-th target defect output by the historical application model before normalization (generally the output of the layer before the last softmax layer). exp() represents the natural exponential function. The distillation based on the target distillation temperature is reflected in the above formula. When the target distillation temperature is 1, the above formula is the conventional softmax function, which plays a role in normalization. Adding the target distillation temperature is equivalent to further amplifying the difference between different logical values during normalization, increasing the information content contained in the soft target, similar to "distillation" in chemistry. The principle of adding the historical application model to train the student model in this embodiment is as Figure 4 shown.

[0126] Similarly, there may also be a situation where the target defects to be detected in the historical application model are inconsistent with those in the student model. Therefore, further, in a preferred embodiment, the above step S305 of establishing the target loss function according to the first reconstruction loss function, the second reconstruction loss function, and the supervision loss function specifically includes:

[0127] Statistically analyze the overlap degree of the types of defects to be recognized by the historical application model and the student model;

[0128] Determine a weight adjustment parameter based on the overlap degree. The weight adjustment coefficient is used to adjust the influence proportion of the second reconstruction loss function on the target loss function;

[0129] Establish a target loss function according to the preset weight and the weight adjustment coefficient.

[0130] The significance of this embodiment lies in adjusting the influence proportion of the second reconstruction loss function on the target loss function according to the overlap degree of the types of defects to be recognized by the historical application model and the student model, so as to obtain a more accurate recognition result. For example, to give an extreme example, if the target defects to be detected by the historical application model and the student model are completely inconsistent, that is, the overlap degree is 0. Then it can be considered that the historical application model has no reference value at this time, and the preset weight corresponding to the second reconstruction loss function should be adjusted to 0 to avoid meaningless calculations. The weight adjustment coefficient in this embodiment is used to adjust the preset weight corresponding to the second reconstruction loss function, and it can have a linear relationship with the overlap degree.

[0131] Specifically, in a preferred embodiment, the target loss function is:

[0132]

[0133] where L is the target loss function, is the first reconstruction loss function, is the second reconstruction loss function, L hard is the supervision loss function, α and β are different preset weights respectively, ε is the weight adjustment parameter, the degree of overlap characterized by the overlap degree is inversely proportional to the weight adjustment parameter, and when the overlap degree takes the maximum value, the weight adjustment parameter is zero.

[0134] The meaning of the above formula is that when the target defects to be detected by the historical application model and the student model are exactly the same, the influence of the first reconstruction loss function and the second reconstruction loss function on the target loss function should be the same, that is, the student model should learn the same proportion of content from the historical application model and the preset teacher model, and at this time the weight adjustment parameter can be 0. As the overlap degree decreases, the weight adjustment parameter increases. At this time, the influence of the first reconstruction loss function on the target loss function will gradually be higher than that of the second reconstruction loss function, and the student model will focus more on learning from the preset teacher model.

[0135] It can be understood that except for the content mentioned above, other steps in the knowledge distillation process of the present invention are the same as those in the prior art, and will not be elaborated herein.

[0136] Combined Figure 5 As shown, the present invention also provides a vision detection optimization system for mobile phone glass covers for edge devices, including:

[0137] A data input module 510, configured to obtain the performance information of a target edge device and the defect information of multiple target defects;

[0138] A model establishment module 520, configured to establish a student model according to the performance information;

[0139] A data preparation module 530, configured to obtain a preset teacher model and preset training data, wherein both the preset teacher model and the student model are neural network models for detecting defects in mobile phone glass covers based on images;

[0140] A distillation preparation module 540, configured to determine a target distillation temperature according to the differences of multiple target defects characterized by multiple defect information;

[0141] A model optimization module 550, configured to perform knowledge distillation on the preset teacher model based on the target distillation temperature according to the preset training data and train the student model to obtain a trained student model;

[0142] A model deployment module 560, configured to deploy the trained student model to the target edge performance device.

[0143] It should be noted here that: the corresponding system provided in the above embodiment can implement the technical solutions described in the above method embodiments. The specific implementation principles of the above modules or units can refer to the corresponding content in the above method embodiments, and will not be elaborated here.

[0144] The present invention provides a vision detection optimization method and system for mobile phone glass covers for edge devices. It first obtains the performance information of a target edge device and the defect information of multiple target defects, then establishes a student model according to the performance information, then obtains a preset teacher model and preset training data, and then determines a target distillation temperature according to the differences of multiple target defects characterized by multiple defect information. Based on the target distillation temperature, it performs knowledge distillation on the preset teacher model according to the preset training data and trains the student model to obtain a trained student model. Finally, it deploys the trained student model to the target edge performance device. Compared with the prior art, the present invention establishes a student model that conforms to its computing power according to the performance information of the target edge device, then trains the student model through knowledge distillation technology to make it close to the performance of the preset teacher model, and controls the temperature during knowledge distillation according to the differences of multiple target defects during the distillation process, realizing the efficient and accurate optimization of the preset teacher model for the target edge device and the actual production situation, and solving the problem of how to deploy the deep learning machine vision detection model of mobile phone glass covers to edge devices.

[0145] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other.

[0146] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. An optimized method for visual inspection of mobile phone glass covers for edge devices, characterized in that Including: Obtain the performance information of the target edge device and the defect information of multiple target defects; Establish a student model according to the performance information; Obtain a preset teacher model and preset training data, where both the preset teacher model and the student model are neural network models for detecting defects in mobile phone glass covers based on images; Determine the target distillation temperature according to the differences of multiple target defects characterized by multiple defect information; Based on the target distillation temperature, perform knowledge distillation on the preset teacher model according to the preset training data and train the student model to obtain the trained student model; Deploy the trained student model to the target edge performance device; Among them, based on the target distillation temperature, performing knowledge distillation on the preset teacher model according to the preset training data and training the student model to obtain the trained student model includes: Obtain a historical application model, which is a neural network model for detecting defects in mobile phone glass covers that has been actually applied in the past; Based on the target distillation temperature, establish a first reconstruction loss function between the preset teacher model and the student model; Based on the target distillation temperature, establish a second reconstruction loss function between the historical application model and the student model; Establish a supervision loss function for the student model; According to the first reconstruction loss function, the second reconstruction loss function and the supervision loss function, establish a target loss function; Based on the target loss function, train the student model according to the preset training data to obtain the trained student model; Among them, according to the first reconstruction loss function, the second reconstruction loss function and the supervision loss function, establishing the target loss function includes: Statistically analyze the type coincidence degree of the defects to be recognized by the historical application model and the student model; Determine a weight adjustment parameter according to the type coincidence degree, and the weight adjustment coefficient is used to adjust the influence proportion of the second reconstruction loss function on the target loss function; Establish a target loss function according to the preset weight and the weight adjustment coefficient; Among them, the target loss function is: ; Among them, is the target loss function, is the first reconstruction loss function, is the second reconstruction loss function, is the supervised loss function, and are different preset weights respectively, is the weight adjustment parameter. The coincidence degree characterized by the type coincidence degree is inversely proportional to the weight adjustment parameter. When the type coincidence degree takes the maximum value, the weight adjustment parameter is zero.

2. The visual inspection optimization method for mobile phone glass covers facing edge devices according to claim 1, wherein, The defect information includes defect sample pictures, and each defect sample picture corresponds to a target defect; Determine the target distillation temperature according to the differences of multiple target defects characterized by multiple defect information, including: Analyze the similarity of every two defect sample pictures pairwise to obtain multiple similarity eigenvalues; Obtain the target distillation temperature according to the distribution of multiple similarity eigenvalues, where the more similar target defects characterized by multiple similarity eigenvalues, the smaller the target distillation temperature.

3. The optimized method for visual inspection of mobile phone glass covers for edge devices according to claim 2, wherein Analyze the similarity of every two defect sample pictures pairwise to obtain multiple similarity eigenvalues, including: Obtain two target defect sample pictures, which are two defect sample pictures to be analyzed for similarity currently; Convert the two target defect sample pictures into two picture vectors respectively, where each element in the picture vector is a pixel value in the target defect sample picture, and the elements at the same position in the two picture vectors correspond to the same pixel position; Calculate the cosine similarity of the two picture vectors and obtain the similarity eigenvalue of the two target defect sample pictures according to the cosine similarity.

4. The optimized method for visual inspection of mobile phone glass covers for edge devices according to claim 2, wherein The larger the value of the similarity eigenvalue, the more similar the two corresponding defect sample pictures are; Obtain the target distillation temperature according to the distribution of multiple similarity eigenvalues, including: Count the total number of types of multiple target defects; Calculate the sum of multiple similar eigenvalue; Calculate the statistical eigenvalue of multiple similar eigenvalue, where the statistical eigenvalue is used to represent the degree of dispersion of multiple similar eigenvalue; Obtain the target distillation temperature according to the total number of types of multiple target defects, the sum of multiple similar eigenvalue, and the statistical eigenvalue of multiple similar eigenvalue, where the target distillation temperature is inversely proportional to the total number of types of multiple target defects, inversely proportional to the sum of multiple similar eigenvalue, and directly proportional to the degree of dispersion characterized by the statistical eigenvalue of multiple similar eigenvalue.

5. The optimized method for visual inspection of mobile phone glass covers for edge devices according to claim 1, characterized in that, Build a student model according to the performance information, including: Obtain the computing power level eigenvalue of the target edge device according to the performance information; Obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer according to the computing power level eigenvalue; Build a student model according to the number of fully connected layers and the number of neurons in each fully connected layer.

6. The optimized method for visual inspection of mobile phone glass covers for edge devices according to claim 5, characterized in that, Obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer according to the computing power level eigenvalue, including: Obtain a preset operation amount characteristic function, where the operation amount characteristic function is used to obtain an operation amount eigenvalue according to the number of fully connected layers and the number of neurons in each fully connected layer, and the operation amount eigenvalue is used to characterize the operation complexity of the student model established based on the number of fully connected layers and the number of neurons; Build a fitness function based on the difference between the computing power level eigenvalue and the preset operation amount characteristic function; Taking the minimum difference between the computing power level eigenvalue and the operation amount eigenvalue as the optimization goal, obtain the number of fully connected layers in the student model and the number of neurons in each fully connected layer based on a preset optimization algorithm.

7. An optimized visual inspection system for mobile phone glass covers for edge devices, characterized in that, Including: A data input module for obtaining the performance information of the target edge device and the defect information of multiple target defects; A model building module for building a student model according to the performance information; A data preparation module for obtaining a preset teacher model and preset training data, where the preset teacher model and the student model are both neural network models for detecting defects in mobile phone glass covers based on images; A distillation preparation module for determining the target distillation temperature according to the differences of multiple target defects characterized by multiple defect information; A model optimization module for performing knowledge distillation on the preset teacher model based on the target distillation temperature and training the student model according to the preset training data to obtain a trained student model; A model deployment module for deploying the trained student model to the target edge performance device; Among them, performing knowledge distillation on the preset teacher model based on the target distillation temperature and training the student model according to the preset training data to obtain a trained student model, including: Obtain a historical application model, where the historical application model is a neural network model for detecting defects in mobile phone glass covers based on images that has been actually applied in the past; Build a first reconstruction loss function between the preset teacher model and the student model based on the target distillation temperature; Build a second reconstruction loss function between the historical application model and the student model based on the target distillation temperature; Build a supervision loss function for the student model; Build a target loss function according to the first reconstruction loss function, the second reconstruction loss function, and the supervision loss function; Based on the target loss function, the student model is trained according to the preset training data to obtain the trained student model; Among them, the target loss function is established according to the first reconstruction loss function, the second reconstruction loss function and the supervision loss function, including: Statistically calculate the species overlap degree of the defects to be recognized by the historical application model and the student model; According to the species overlap degree, determine the weight adjustment parameter, and the weight adjustment coefficient is used to adjust the influence proportion of the second reconstruction loss function on the target loss function; According to the preset weight and the weight adjustment coefficient, establish the target loss function; Among them, the target loss function is: ; Among them, is the target loss function, is the first reconstruction loss function, is the second reconstruction loss function, is the supervision loss function, and are different preset weights respectively, is the weight adjustment parameter. The coincidence degree characterized by the type coincidence degree is inversely proportional to the weight adjustment parameter. When the type coincidence degree takes the maximum value, the weight adjustment parameter is zero.

Citation Information

Patent Citations

  • Knowledge distillation and quantification technology for power scene edge calculation large model compression

    CN115223049A

  • Industrial defect detection model compression method based on knowledge distillation

    CN116152240A