Target object grade determination method, apparatus, device and storage medium
By training the model to extract product image features and perform linear fitting, the problem that it is difficult for the human eye to accurately determine the target object level in the product is solved, and the accurate measurement and screening of product level is achieved.
Patent Information
- Application Number
- PCT/CN2023/142821
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the rating of target objects in a product mainly relies on human eye resolution, making it difficult to achieve accurate measurement, resulting in inaccurate product screening.
Through the training, the image features of the sample product are extracted using the EfficientNet network and the feature pyramid FPN network, regression analysis is performed in combination with the full connection layer, and a hierarchical value sequence is generated, and linearly fitted through Gaussian smoothing or scoring to obtain the target hierarchical value.
It realizes accurate measurement of the target object level in the product, and supports accurate product screening and capacity adjustment.
Smart Images

Figure CN2023142821_03072025_PF_FP_ABST
Abstract
Description
Method, device, equipment and storage medium for determining target object level Technical Field
[0001] The present application belongs to the field of industrial detection technology, and in particular relates to a method, device, equipment and storage medium for determining the level of a target object. Background Art
[0002] Product screening is a key process in industrial inspection. Before screening, it is usually necessary to determine the grade of target objects (e.g., target defects, target objects, etc.) within the product. During product screening, a certain percentage of products are inspected based on the grade of the target objects.
[0003] Currently, the grade of target objects in products is usually distinguished by the human eye, but it is difficult for the human eye to accurately measure the grade of target objects. Therefore, how to accurately determine the grade of target objects in products is an urgent problem that needs to be solved.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for determining the level of a target object, thereby accurately determining the level of a target object in a product, at least to a certain extent.
[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0007] According to a first aspect of an embodiment of the present application, a method for determining a grade of a target object is provided, comprising: determining a grade value of each sample product image in a sample product image set on a target object through a preset model to obtain a first grade value sequence, wherein the grade value is used to characterize the grade determination priority of the target object, and the preset model is trained by the sample product image set; determining a grade value of the target product image on the target object through the preset model as an initial grade value; and determining a target grade value of the target product image on the target object based on the first grade value sequence and the initial grade value.
[0008] In some embodiments, the individual level values in the first level value sequence are arranged in ascending or descending order, and determining the target level value of the target product image on the target object based on the first level value sequence and the initial level value includes: performing linear fitting on the individual level values in the first level value sequence and the initial level value, or performing linear fitting on the individual level values in the first level value sequence to obtain a fitted level value; and determining the target level value of the target product image on the target object based on the fitted level value.
[0009] In some embodiments, the linear fitting of each level value in the first level value sequence and the initial level value to obtain a fitted level value includes: using a Gaussian function to linearly fit each level value in the first level value sequence and the initial level value to obtain a fitted level value corresponding to the initial level value.
[0010] In some embodiments, the use of a Gaussian function to perform linear fitting on each level value in the first level value sequence and the initial level value to obtain a fitted level value corresponding to the initial level value includes: adding the initial level value to the first level value sequence to obtain a second level value sequence; determining a convolution kernel length based on the number of each level value in the second level value sequence; and fitting each level value in the second level value sequence using the Gaussian function and the convolution kernel length to obtain a fitted level value corresponding to the initial level value.
[0011] In some embodiments, determining the target grade value of the target product image on the target object based on the fitting grade value includes: determining the fitting grade value corresponding to the initial grade value as the target grade value of the target product image on the target object.
[0012] In some embodiments, linearly fitting each level value in the first level value sequence to obtain a fitted level value includes: linearly assigning points to each level value in the first level value sequence using the target level value sequence of the target object to obtain the fitted level value.
[0013] In some embodiments, the target grade value sequence of the target object is used to linearly assign scores to each grade value in the first grade value sequence to obtain the fitting grade value, including: segmenting the first grade value sequence according to the number of target grade values in the target grade value sequence to obtain multiple grade value subsequences; using the target grade value in the target grade value sequence to assign scores to each grade value subsequence, and using the target grade value corresponding to each grade value subsequence as the fitting grade value corresponding to each grade value in each grade value subsequence.
[0014] In some embodiments, determining the target level value of the target product image on the target object based on the fitting level value includes: determining the fitting level value corresponding to the initial level value based on a level value mapping table, wherein the level value mapping table is used to characterize the correspondence between each level value and the fitting level value in the first level value sequence; and determining the fitting level value corresponding to the initial level value as the target level value of the target product image on the target object.
[0015] In some embodiments, the target object grade determination method further includes: training an initial model using each sample product image in the sample product image set, wherein the initial model includes a convolutional neural network and a fully connected layer, the convolutional neural network is used to extract features from each sample product image by uniformly scaling the network depth, network width and resolution of each sample product image, and fusing features of different scales, and the fully connected layer is used to perform regression analysis on the results output by the convolutional neural network; when the loss function of the initial model reaches a preset convergence condition, the preset model is obtained.
[0016] In some embodiments, the convolutional neural network includes an EfficientNet network and a feature pyramid FPN network in sequence.
[0017] In some embodiments, the loss function is a mean square error loss function.
[0018] In some embodiments, the training of the initial model using each sample product image in the sample product image set includes: labeling a reference grade value for each sample product image in the sample product image set, wherein the granularity of the reference grade value is greater than the granularity of the grade value in the first grade value sequence; and training the initial model based on the sample product images labeled with the reference grade values.
[0019] In some embodiments, the target object is a target defect, and the grade value is used to represent the severity of the target defect.
[0020] According to a second aspect of an embodiment of the present application, a device for determining a level of a target object is provided, comprising: a first prediction unit, configured to determine a level value of each sample product image in a set of sample product images on a target object through a preset model, and obtain a first level value sequence, wherein the level value is used to characterize a level determination priority of the target object, and the preset model is obtained by training the set of sample product images; a second prediction unit, configured to determine a level value of a target product image on the target object through the preset model as an initial level value; and a target level value determination unit, configured to determine a target level value of the target product image on the target object based on the first level value sequence and the initial level value.
[0021] According to a third aspect of an embodiment of the present application, a target object level determination device is provided, comprising a processor and a memory, wherein the memory stores computer program instructions that can be executed by the processor, and when the processor executes the computer program instructions, the steps of the method described in any one of the first aspects above are implemented.
[0022] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor is prompted to implement the steps of the method described in any one of the first aspects above.
[0023] In the present application, a preset model is used to determine the grade value of each sample product image on a target object in a sample product image set, thereby obtaining a first grade value sequence, wherein the grade value is used to characterize the grade determination priority of the target object, and the preset model is trained by the sample product image set; the grade value of the target product image on the target object is determined by the preset model as an initial grade value; and based on the first grade value sequence and the initial grade value, a target grade value of the target product image on the target object is determined. The technical solution provided by the present application can accurately measure the grade of the target object in the product and obtain a target grade value, so as to accurately screen the products according to the target grade value.
[0024] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0026] FIG1 is a schematic flow chart showing a method for determining the level of a target object in one embodiment;
[0027] FIG2 shows a schematic structural diagram of an initial model in one embodiment;
[0028] FIG3 shows a schematic diagram of the structure of the EfficientNet network in FIG2 ;
[0029] FIG4 shows a schematic structural diagram of the MBConv module in FIG3 ;
[0030] FIG5 shows a schematic structural diagram of the FPN network in FIG2 ;
[0031] FIG6 shows a schematic structural diagram of the fully connected layer in FIG2 ;
[0032] FIG7 shows a distribution diagram of the first level value sequence in FIG1 ;
[0033] FIG8 shows another distribution diagram of the first level value sequence in FIG1 ;
[0034] FIG9 shows a schematic diagram of the distribution of the first level value sequence in FIG1 after linear fitting;
[0035] FIG10 is a schematic flow chart showing a method for determining the level of a target object in another embodiment;
[0036] FIG11 shows a schematic diagram of data flow in FIG10 ;
[0037] FIG12 shows a schematic diagram of the sample product image set in FIG10 ;
[0038] FIG13 is a schematic diagram showing some sample product images in FIG12 ;
[0039] FIG14 is a schematic flow chart showing a method for determining the level of a target object in another embodiment;
[0040] FIG15 shows a schematic diagram of data flow in FIG14 ;
[0041] FIG16 shows a schematic diagram of the sample product image set in FIG14 ;
[0042] FIG17 is a schematic diagram showing some sample product images in FIG14 ;
[0043] FIG18 shows a block diagram of a target object level determination device according to one embodiment;
[0044] FIG19 shows a schematic structural diagram of a target object level determination device in one embodiment. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0046] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0047] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0048] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0049] It should also be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described.
[0050] In order to enable those skilled in the art to better understand the present application, the application scenario involved in the present application is first briefly described by taking the target object being the target defect existing in the product as an example.
[0051] The defect levels currently used in factories are mainly divided into four categories: the first is minor defects, which generally refer to defects that may have a slight impact on the product appearance and the next process; the second is general defects, which generally refer to defects that do not affect the operation and function of the product and will not become the cause of failure, but have a greater impact on the product appearance and the next process; the third is serious defects, which generally refer to defects that are likely to cause irreparable failures or have unacceptable effects on the product appearance; the fourth is fatal defects, which generally refer to defects that may cause safety problems.
[0052] During the production process, the factory has certain requirements for production capacity, and the daily shipment volume varies. Under this requirement, it is necessary to appropriately adjust the detection volume of defective products according to the defect level. For example, industrial products with defect levels higher than the defect level threshold are identified as defective products and intercepted, thereby achieving production capacity adjustment.
[0053] Currently, defect levels are typically discerned by the human eye, which is difficult to accurately measure, making it impossible to accurately screen industrial products based on defect levels. Based on this, embodiments of the present application propose a target object level determination method that can accurately measure the level of a target object (e.g., the level of a target defect) to obtain a target level value, thereby facilitating accurate screening of industrial products based on the target level value and achieving flexible adjustment of production capacity.
[0054] Fig. 1 shows a schematic flow chart of a method for determining the level of a target object in one embodiment. As shown in Fig. 1 , the method for determining the level of a target object may include the following steps 101 to 103 .
[0055] In step 101, the grade value of each sample product image in the sample product image set on the target object is determined by a preset model to obtain a first grade value sequence, wherein the grade value is used to characterize the grade determination priority of the target object, and the preset model is trained by the sample product image set.
[0056] Among them, the target object can be a target defect, a target object that is poorly integrated with the background of the product image, etc., and the embodiments of the present application are not limited to this.
[0057] When the target object is a target defect, the priority of the target object's grade determination is the severity of the target defect, that is, the grade value is used to represent the severity of the target defect.
[0058] When the target object is the above-mentioned target object, the priority of the target object's level determination is the fusion degree of the target object, that is, the level value is used to represent the fusion degree of the target object.
[0059] For ease of understanding, the following detailed description is given by taking the target object being a target defect as an example.
[0060] It can be understood that the preset model is a model obtained by training the initial model using a set of sample product images. The initial model can be a Bayesian Ridge regression model or other regression model, and the embodiments of the present application are not limited to this.
[0061] In some embodiments, the initial model may include a convolutional neural network and a fully connected layer. The convolutional neural network is used to extract features from each sample product image by uniformly scaling the network depth, network width, and resolution of each sample product image, and to fuse features of different scales. The fully connected layer is used to perform regression analysis on the results output by the convolutional neural network.
[0062] A convolutional neural network can be implemented using a single network or a combination of multiple networks. When implemented using a combination of multiple networks, the convolutional neural network can sequentially include an EfficientNet network and a Feature Pyramid (FPN) network. The EfficientNet network is used to extract features from each sample product image by uniformly scaling the network depth, network width, and resolution of each sample product image, while the FPN network is used to fuse features at different scales.
[0063] It is understandable that EfficientNet is a highly scalable and efficient convolutional neural network that uniformly scales the network depth, network width, and image resolution through a set of fixed scaling factors, balancing the network depth, network width, and resolution, and improving model efficiency and accuracy.
[0064] FPN (Feature Pyramid Network) is a convolutional neural network that integrates features of different scales and maintains information richness. It can solve the scale change problem in target detection and has good robustness for small target detection.
[0065] Figure 2 shows a schematic diagram of the structure of an initial model in one embodiment. As shown in Figure 2, in some embodiments, the initial model may sequentially include an EfficientNet network, an FPN network, and a fully connected layer. The initial model is trained using sample product images from a sample product image collection. When the loss function of the initial model meets a preset convergence condition, a preset model is obtained.
[0066] Before training the initial model, a reference grade value can be labeled for each sample product image in the sample product image set, wherein the granularity of the reference grade value is greater than the granularity of the grade value in the first grade value sequence; and then the initial model is trained based on the sample product images labeled with the reference grade values.
[0067] Granularity refers to the level of detail in the data. Compared to the values in the first-level value sequence, the reference level values have a higher granularity. For example, the reference level values could be 0, 1, and 2, corresponding to minor defects, general defects, and severe defects, respectively. The level values in the first-level value sequence are even more refined, such as 0.1, 0.2, 1.2, and 2.3.
[0068] It should be noted that, limited by experience and subjective initiative, when people label the reference grade values of target defects in sample product images, they tend to quickly identify several major categories and then label them, but are unable to quantify the details of the values. For example, when the target defect is a minor defect, the labeled reference grade value is 0; when the target defect is a general defect, the labeled reference grade value is 1; when the target defect is a serious defect, the labeled reference grade value is 2. However, it is difficult to quantify the details of minor defects, general defects or serious defects of different degrees of severity. Using the above initial model, it is only necessary to label 3 to 5 major categories of reference grade values for each sample product image in the sample product image set, that is, a more refined grade value (for example, 10 to 100 categories) corresponding to each sample product image can be obtained, achieving the purpose of "simple" labeling and "rich" results, thereby solving the problem of difficult labeling of target objects in products and the inability to quantify the grades of target objects.
[0069] Figure 3 shows a schematic diagram of the structure of the EfficientNet network in Figure 2. In some embodiments, the EfficientNet network may be an EfficientNet B0 network. As shown in Figure 3, the EfficientNet B0 network includes 16 MBConv (mobile inverted bottleneck convolution) modules, 2 convolutional layers, 1 global average pooling layer, and 1 classification layer.
[0070] Referring to Figure 4, the MBConv module mainly performs dimensionality increase through a common convolution Conv unit, and then further extracts features through a depth-wise separable convolution Depwise Conv unit and an attention mechanism SE unit, and then reduces the dimensionality through a common convolution Conv unit and outputs it through a Dropout layer.
[0071] The core structure of the EfficientNet B0 network is the MBConv module, which introduces the attention idea of SENet (Squeeze-and-Excitation Network). SENet has achieved the highest accuracy on public data sets. At the same time, the MBConv module uses the jump connection of the input to give the model a random depth, shorten the time required for model training, and improve the performance of the model. Therefore, the EfficientNet B0 network has natural advantages in efficiency and accuracy. The embodiment of the present application requires that the sample product image marked with a reference grade value with a larger granularity be used as the input of the initial model, so that the model outputs a grade value with a smaller granularity, that is, the reference grade value of the cluster needs to be refined, which places extremely high requirements on the precision and accuracy of the model output results. Therefore, the EfficientNet B0 network is selected as the main network.
[0072] Figure 5 shows a schematic diagram of the structure of the FPN network in Figure 2. As shown in Figure 5, the FPN network mainly includes a bottom-up unit, a top-down unit, a lateral connection unit (not labeled), and a convolutional fusion unit. The advantage of the FPN network lies in the detection of small targets. The embodiment of the present application uses the FPN network to focus on feature extraction and accurate reasoning of minor defects, and then fine-tune the feature values output by the EfficientNet B0 network, improving the reliability and accuracy of the ranking of the output results of the entire model.
[0073] Figure 6 shows a schematic diagram of the fully connected layer in Figure 2. As shown in Figure 6, the fully connected layer can be set to three hidden layers, with the number of neurons in each layer being 512, 64, and 8, respectively. The fully connected layer connects the FPN network to the output layer (also called the regression head), thereby obtaining the regression results of the entire model.
[0074] In some embodiments, considering that the initial model is a regression model, a mean square error loss function may be used as a loss function during the training process to improve the prediction accuracy of the model.
[0075] Among them, the mean square error loss function can refer to the following formula:
[0076] in, is the reference level value, Y i is the predicted grade value, and n is the number of sample product images.
[0077] During the model training process, if the value of the mean square error loss function is less than 0.08, it can be determined that the loss function meets the preset convergence condition and the training is terminated.
[0078] After obtaining the preset model through model training, each sample product image in the sample product image set can be input into the preset model to obtain the grade value of each sample product image output by the preset model on the target object, and then obtain a first grade value sequence composed of each grade value.
[0079] In step 102, a grade value of the target product image on the target object is determined by a preset model as an initial grade value.
[0080] The target product image may be an image of an industrial product to be screened.
[0081] During the implementation process, the target product image is input into the preset model, and the grade value of the target product image output by the preset model on the target object can be obtained.
[0082] In step 103, a target grade value of the target product image on the target object is determined based on the first grade value sequence and the initial grade value.
[0083] Figure 7 shows a distribution diagram of the first grade value sequence in Figure 1. The horizontal axis in Figure 7 represents the serial number of the sample product image, and the vertical axis represents the grade value of the sample product image on the target defect output by the preset model. If the preset model is trained using sample product images labeled with three reference grade values, 0, 1, and 2, then when the sample product images are input into the preset model for inference, the grade values output by the preset model are also arranged from small to large, showing a monotonically increasing trend. However, some grade values are concentrated near the three reference grade values, indicating clustering.
[0084] FIG8 shows another distribution diagram of the first grade value sequence in FIG1 . This figure is a histogram generated after dividing each grade value in the first grade value sequence into 100 categories. As shown in FIG8 , the number of sample product images corresponding to the 10th, 50th, and 90th categories is relatively concentrated, which verifies that the grade value output by the preset model is reasonable. However, the grade value output by the preset model is not convenient for subsequent product screening. Because when performing product screening, the preset grade threshold needs to be in a first-order linear relationship with the output of industrial products as much as possible, so that the output of industrial products can be uniformly controlled by adjusting the preset grade threshold. However, if the grade value output by the preset model is directly used for product screening, then due to the clustering phenomenon of the grade value output by the preset model, it is uneven, and the output of industrial products cannot be uniformly controlled by adjusting the preset grade threshold.
[0085] Referring back to Figure 7, assuming that the preset level threshold is adjusted from 0.9 to 1.1, a large number of industrial products corresponding to target product images with level values near 1 cannot be detected; and assuming that the preset level threshold is adjusted from 1.4 to 1.8, since there are fewer level values within this range, although the preset level threshold has been adjusted to a large extent, it has little effect on the detection volume of industrial products.
[0086] Therefore, for the target product image, after obtaining the initial grade value of the target product image on the target defect through the preset model, the initial grade value needs to be further processed to obtain the target grade value that is convenient for product screening.
[0087] In some embodiments, linear fitting can be performed on each level value in the first level value sequence and the initial level value, or linear fitting can be performed on each level value in the first level value sequence to obtain a fitted level value; based on the fitted level value, the target level value of the target product image on the target object is determined.
[0088] It should be noted that the various level values in the first level value sequence output by the preset model are arranged in ascending or descending order. This can avoid a large difference between the fitted level value and the actual level value during the linear fitting process, which in turn leads to inaccurate screening results.
[0089] During the implementation process, linear fitting can be achieved through Gaussian smoothing, scoring, etc. The specific method of performing linear fitting will be explained in the following by way of examples, and the embodiments of the present application are not limited thereto.
[0090] Figure 9 shows the distribution of the first grade value sequence in Figure 1 after linear fitting. As shown in Figure 9, after linear fitting, the grade values increase evenly across the board, with no clustering observed. By performing linear fitting on the grade values, we can obtain the fitted grade value corresponding to each grade value. This fitted grade value is then used to determine the target grade value for the target product image at the target defect. This ensures that the target grade values for different target product images at the target defect are also linearly arranged, facilitating subsequent product screening.
[0091] After the target grade value is determined, the industrial products corresponding to the target product image may be screened according to the target grade value.
[0092] In some embodiments, when screening industrial products, it is necessary to obtain the target product image of each industrial product, determine the target level value of the target product image on the target defect, and then determine whether the corresponding industrial product is a defective product based on the target level value. If the industrial product is a defective product, it is intercepted; if the industrial product is not a defective product, it is transported to the next process.
[0093] In some embodiments, the target grade value may be directly compared with a preset grade threshold; when the target grade value is greater than the preset grade threshold, the industrial product corresponding to the target product image is determined to be a defective product.
[0094] It is understandable that, since the target grade values of different target product images on the target defects are arranged linearly, the number of defective products can be evenly controlled by adjusting the preset grade threshold, thereby achieving the purpose of adjusting production capacity.
[0095] In other embodiments, the target level value may be divided into a preset number of levels, for example, ten levels, and then the divided levels are compared with a preset level threshold.
[0096] In an embodiment of the present application, a first grade value sequence is obtained by determining the grade value of each sample product image in a sample product image set on a target object through a preset model, wherein the grade value is used to characterize the grade determination priority of the target object, and the preset model is trained by the sample product image set; the grade value of the target product image on the target object is determined through the preset model as an initial grade value; and the target grade value of the target product image on the target object is determined based on the first grade value sequence and the initial grade value. The technical solution provided by the present application can accurately measure the grade of the target object in the product and obtain the target grade value, so as to accurately screen the products according to the target grade value.
[0097] FIG10 is a flow chart showing a method for determining the level of a target object in another embodiment. This embodiment of the present application uses a Gaussian smoothing method to perform linear fitting to obtain a target level value. As shown in FIG10 , the method for determining the level of a target object may include the following steps:
[0098] Step 1001: Determine the rank value of each sample product image in the sample product image set on the target object using a preset model to obtain a first rank value sequence, wherein the rank value is used to represent the rank determination priority of the target object. The preset model is trained using the sample product image set.
[0099] Step 1002: Determine the grade value of the target product image on the target object using a preset model as an initial grade value;
[0100] Step 1003: Perform linear fitting on each level value in the first level value sequence and the initial level value using a Gaussian function to obtain a fitted level value corresponding to the initial level value;
[0101] Step 1004 : Determine the fitting grade value corresponding to the initial grade value as the target grade value of the target product image on the target object.
[0102] Figure 11 is a schematic diagram of the data flow in Figure 10. As shown in Figure 11, a set of sample product images is labeled with reference grade values, and the sample product image set is used to train an initial model to obtain a preset model. The preset model is used to predict the set of sample product images to obtain a first grade value sequence, which is arranged in ascending or descending order. The preset model is used to predict the target product image to obtain an initial grade value. The initial grade value is added to the first grade value sequence and Gaussian smoothed to obtain a fitted grade value (i.e., a target grade value) corresponding to the initial grade value.
[0103] In some embodiments, the initial level value can be added to the first level value sequence to obtain a second level value sequence; the convolution kernel length is determined based on the number of each level value in the second level value sequence; and the Gaussian function and the convolution kernel length are used to fit each level value in the second level value sequence to obtain a fitting level value corresponding to the initial level value.
[0104] During the implementation process, the level value closest to the initial level value can be found in the first level value sequence, and based on the size of the initial level value and the closest level value, it can be determined whether to add the initial level value to the front or back of the closest level value to obtain the second level value sequence.
[0105] The Gaussian function can refer to the following formula:
[0106] Where x is the level value in the second level value sequence, f(x) is the fitted level value, with one convolution kernel length as the object, μ is the mean of the level values within the convolution kernel length, and σ is the standard deviation of the level values within the convolution kernel length.
[0107] It's understandable that Gaussian smoothing replaces each data value with the average of its surrounding data. The length of the Gaussian convolution kernel is positively correlated with the length of the data. As long as the convolution kernel length is chosen appropriately, the desired smoothed data can be obtained. In practice, the convolution kernel length can be set to one-tenth the number of levels in the second-level sequence.
[0108] Since Gaussian smoothing does not change the order of data, the fitted grade value of the initial grade value after Gaussian smoothing can be obtained according to the position of the initial grade value in the second grade value sequence, and the fitted grade value can be used as the target grade value.
[0109] The following experiment verifies the target object level determination method of the embodiment of the present application. Figure 12 shows a schematic diagram of the sample product image set in Figure 10. As shown in Figure 12, the sample product image set comes from the appearance of a certain industrial product, and the target object is a target defect, specifically a foreign body under the membrane. The severity of the target defect in the sample product images in columns 1 to 5 is the strongest, and the reference level value 2 is marked; the severity of the target defect in the sample product images in columns 6 to 10 is second, and the reference level value 1 is marked; the severity of the target defect in the sample product images in columns 11 to 15 is the weakest, and the reference level value 0 is marked. 80% of the data in the sample product image set is used to train the initial model, and 20% of the data is used for testing to obtain the preset model. The sample product image set is input into the preset model to obtain a first level value sequence consisting of multiple level values. After Gaussian smoothing of the first level value sequence, multiple fitting level values are obtained. These fitting level values are evenly divided into ten levels (1 to 10). A sample product image is randomly selected for each level. The result is shown in Figure 13. The severity of the target defect in the sample product image with a level of 10 is the strongest. As the level decreases, the severity of the target defect in the sample product image gradually decreases. It can be seen that the scheme of the embodiment of the present application can achieve a mapping from three types of defect levels to ten types of defect levels, and the size of these ten types of defect levels is linearly related to the severity of their corresponding target defects.
[0110] The embodiment of the present application performs linear fitting by using Gaussian smoothing to obtain the target grade value. Compared with manual scoring of the grade of the target object, it is less dependent on human experience and improves the accuracy of the grade determination of the target object.
[0111] FIG14 shows a flow chart of a method for determining the level of a target object in another embodiment. The embodiment of the present application uses a scoring method to perform linear fitting to obtain a target level value. As shown in FIG14 , the method for determining the level of a target object may include the following steps:
[0112] Step 1401: Determine the rank value of each sample product image in the sample product image set on the target object using a preset model to obtain a first rank value sequence, wherein the rank value is used to represent the rank determination priority of the target object. The preset model is trained using the sample product image set.
[0113] Step 1402: Determine the grade value of the target product image on the target object using a preset model as an initial grade value;
[0114] Step 1403: linearly assign scores to the grade values in the first grade value sequence using the target grade value sequence of the target object to obtain fitting grade values.
[0115] Step 1404: Determine the fitting grade value corresponding to the initial grade value according to the grade value mapping table, wherein the grade value mapping table is used to represent the correspondence between each grade value in the first grade value sequence and the fitting grade value;
[0116] Step 1405 : Determine the fitting grade value corresponding to the initial grade value as the target grade value of the target product image on the target object.
[0117] It should be noted that the preset model in the embodiment shown in FIG14 may be different from or the same as the preset model in the embodiment shown in FIG10 . During implementation, the preset model in the embodiment shown in FIG10 may be trained based on an initial model including an EfficientNet network, an FPN network, and a fully connected layer, and a set of sample product images, while the preset model in the embodiment shown in FIG14 may be trained based on a Bayesian Ridge regression model and a set of sample product images; or both the preset model in the embodiment shown in FIG10 and the preset model in the embodiment shown in FIG14 may be trained based on an initial model including an EfficientNet network, an FPN network, and a fully connected layer, and a set of sample product images.
[0118] Figure 15 shows the data flow diagram in Figure 14. As shown in Figure 15, after obtaining the first grade value sequence, the fitting grade value corresponding to each grade value in the first grade value sequence is determined by assigning points, and a grade value mapping table is generated. During the inference phase, the fitting grade value corresponding to the initial grade value of the target product image on the target object is determined by looking up the table, and this is then determined as the target grade value corresponding to the initial grade value.
[0119] In some embodiments, the first level value sequence can be segmented according to the number of target level values in the target level value sequence to obtain multiple level value subsequences; each level value subsequence is scored using the target level value in the target level value sequence, and the target level value corresponding to each level value subsequence is used as the fitting level value corresponding to each level value in each level value subsequence.
[0120] Among them, the target level value sequence is composed of multiple target level values. Taking the target level value sequence as an example, which has 100 target level values of 0.01, 0.02......., and 1, the first level value sequence can be segmented to obtain 100 level value subsequences, and the first level value subsequence is assigned to 0.01, the second level value subsequence is assigned to 0.02,......., and the 100th level value subsequence is assigned to 1. Then, each level value in the first level value subsequence is assigned to 0.01, each level value in the second level value subsequence is assigned to 0.02, and so on. Each level value in the 100th level value subsequence is assigned to 1. In this way, each level value in the first level value sequence can be evenly distributed in the range of 0 to 1.
[0121] To facilitate subsequent reasoning, a grade value mapping table can be generated based on the correspondence between each grade value in the first grade value sequence and the fitted grade value (i.e., the assigned value). During the reasoning phase, the grade value mapping table is searched based on the initial grade value of the target product image on the target object, and the fitted grade value corresponding to the grade value closest to the initial grade value is used as the fitted grade value corresponding to the initial grade value.
[0122] The following experiment verifies the target object level determination method of the embodiment of the present application. Figure 16 shows a schematic diagram of the sample product image set in Figure 14. As shown in Figure 16, the sample product image set comes from the appearance of the display screen, and the target object is a target defect, specifically a bulge on the screen surface. The severity of the target defect in the sample product images in columns 1 to 5 is the weakest, and is marked with a reference level value of 0; the severity of the target defect in the sample product images in columns 6 to 10 is relatively strong, and is marked with a reference level value of 1; the severity of the target defect in the sample product images in columns 11 to 15 is the strongest, and is marked with a reference level value of 2. 80% of the data in the sample product image set is used to train the initial model, and 20% of the data is used for testing to obtain a preset model. The sample product image set is input into the preset model to obtain a first level value sequence consisting of multiple level values. After assigning points to each level value in the first level value sequence, multiple fitting level values are obtained. These fitting level values are evenly divided into ten levels (1 to 10). Five sample product images are randomly selected for each level. The results are shown in Figure 17. The severity of the target defect in the sample product image with a level of 1 is the weakest. As the level increases, the severity of the target defect in the sample product image gradually increases. This shows that the solution of this embodiment can also achieve a mapping from three defect levels to ten defect levels, and the magnitude of these ten defect levels is linearly related to the severity of the corresponding target defects.
[0123] The embodiment of the present application performs linear fitting by using a scoring method to obtain a target grade value, thereby achieving uniform quantization of the grade value of the target object and accurate measurement of the grade of the target object.
[0124] The following describes an embodiment of the device of the present application, which can be used to execute the target object level determination method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the target object level determination method in the above embodiment of the present application.
[0125] Referring to Figure 18 , a block diagram of a target object rank determination apparatus in an embodiment of the present application is shown. As shown in Figure 18 , the target object rank determination apparatus in an embodiment of the present application may include: a first prediction unit 1801, a second prediction unit 1802, and a target rank value determination unit 1803, wherein the first prediction unit 1801 is configured to determine the rank value of each sample product image in a sample product image set on the target object through a preset model to obtain a first rank value sequence, wherein the rank value is used to characterize the rank determination priority of the target object, and the preset model is obtained by training the sample product image set; the second prediction unit 1802 is configured to determine the rank value of the target product image on the target object through the preset model as an initial rank value; and the target rank value determination unit 1803 is configured to determine a target rank value of the target product image on the target object based on the first rank value sequence and the initial rank value.
[0126] In some embodiments, the level values in the first level value sequence are arranged in ascending or descending order, and the target level value determination unit 1803 can also be used to perform linear fitting on the level values and initial level values in the first level value sequence, or to perform linear fitting on the level values in the first level value sequence to obtain a fitting level value; based on the fitting level value, determine the target level value of the target product image on the target object.
[0127] In some embodiments, the target level value determination unit 1803 can also be used to use a Gaussian function to perform linear fitting on each level value in the first level value sequence and the initial level value to obtain a fitting level value corresponding to the initial level value.
[0128] In some embodiments, the target level value determination unit 1803 can also be used to add the initial level value to the first level value sequence to obtain a second level value sequence; determine the convolution kernel length based on the number of each level value in the second level value sequence; and use the Gaussian function and the convolution kernel length to fit each level value in the second level value sequence to obtain a fitting level value corresponding to the initial level value.
[0129] In some embodiments, the target grade value determining unit 1803 may also be configured to determine the fitting grade value corresponding to the initial grade value as the target grade value of the target product image on the target object.
[0130] In some embodiments, the target grade value determination unit 1803 may also be used to linearly assign scores to each grade value in the first grade value sequence using the target grade value sequence of the target object to obtain the fitting grade value.
[0131] In some embodiments, the target grade value determination unit 1803 can also be used to segment the first grade value sequence according to the number of target grade values in the target grade value sequence to obtain multiple grade value subsequences; use the target grade value in the target grade value sequence to assign points to each grade value subsequence, and use the target grade value corresponding to each grade value subsequence as the fitting grade value corresponding to each grade value in each grade value subsequence.
[0132] In some embodiments, the target level value determination unit 1803 can also be used to determine the fitting level value corresponding to the initial level value based on the level value mapping table, wherein the level value mapping table is used to characterize the correspondence between each level value in the first level value sequence and the fitting level value; the fitting level value corresponding to the initial level value is determined as the target level value of the target product image on the target object.
[0133] In some embodiments, the target object grade determination device may further include a model training unit (not shown) for training an initial model using each sample product image in a sample product image set, wherein the initial model includes a convolutional neural network and a fully connected layer. The convolutional neural network is used to extract features from each sample product image by uniformly scaling the network depth, network width, and resolution of each sample product image, and to fuse features of different scales. The fully connected layer is used to perform regression analysis on the output results of the convolutional neural network.
[0134] In some embodiments, the convolutional neural network includes an EfficientNet network and a feature pyramid FPN network in sequence.
[0135] In some embodiments, the loss function is a mean square error loss function.
[0136] In some embodiments, the model training unit can also be used to label reference grade values for each sample product image in the sample product image set, wherein the granularity of the reference grade values is greater than the granularity of the grade values in the first grade value sequence; and the initial model is trained based on the sample product images labeled with reference grade values.
[0137] In some embodiments, the target object is a target defect, and the grade value is used to represent the severity of the target defect.
[0138] Based on the same inventive concept, an embodiment of the present application also provides a target object level determination device. Referring to Figure 19, a structural schematic diagram of the target object level determination device in an embodiment of the present application is shown. The target object level determination device includes one or more memories 1904, one or more processors 1902, and at least one computer program (computer program instruction) stored on the memory 1904 and executable on the processor 1902. When the processor 1902 executes the computer program, the method described above is implemented.
[0139] In FIG. 19 , a bus architecture (represented by bus 1900) is shown. Bus 1900 may include any number of interconnected buses and bridges. Bus 1900 links various circuits, including one or more processors represented by processor 1902 and memory represented by memory 1904. Bus 1900 may also link various other circuits, such as peripherals, voltage regulators, and power management circuits, all of which are well known in the art and, therefore, will not be described further herein. Bus interface 1905 provides an interface between bus 1900 and receiver 1901 and transmitter 1903. Receiver 1901 and transmitter 1903 may be the same component, namely a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 1902 is responsible for managing bus 1900 and general processing, while memory 1904 may be used to store data used by processor 1902 when performing operations.
[0140] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor is prompted to implement the steps of the method as described above.
[0141] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of this application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or a combination of any of these. Furthermore, the functional units may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0142] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0143] The units described as separate components may or may not be physically separate, and the components of the control device may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0144] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store computer program instructions.
[0145] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for determining the level of a target object, characterized in that, Including: Determine the rank values of each sample product image in the sample product image set on the target object through a preset model to obtain a first rank value sequence, where the rank value is used to represent the rank determination priority of the target object, and the preset model is trained by the sample product image set; Determine the rank value of the target product image on the target object through the preset model as the initial rank value; Determine the target rank value of the target product image on the target object according to the first rank value sequence and the initial rank value.
2. The method according to claim 1, wherein The rank values in the first rank value sequence are arranged in ascending or descending order. The determining the target rank value of the target product image on the target object according to the first rank value sequence and the initial rank value includes: Perform linear fitting on each rank value in the first rank value sequence and the initial rank value, or perform linear fitting on each rank value in the first rank value sequence to obtain a fitted rank value; Determine the target rank value of the target product image on the target object according to the fitted rank value.
3. The method according to claim 2, characterized in that, The performing linear fitting on each rank value in the first rank value sequence and the initial rank value to obtain a fitted rank value includes: Perform linear fitting on each rank value in the first rank value sequence and the initial rank value by using a Gaussian function to obtain the fitted rank value corresponding to the initial rank value.
4. The method according to claim 3, characterized in that The performing linear fitting on each rank value in the first rank value sequence and the initial rank value by using a Gaussian function to obtain the fitted rank value corresponding to the initial rank value includes: Add the initial rank value to the first rank value sequence to obtain a second rank value sequence; Determine the convolution kernel length according to the number of rank values in the second rank value sequence; Use the Gaussian function and the convolution kernel length to fit each rank value in the second rank value sequence to obtain the fitted rank value corresponding to the initial rank value.
5. The method according to claim 3, wherein The determining the target rank value of the target product image on the target object according to the fitted rank value includes: Determine the fitted rank value corresponding to the initial rank value as the target rank value of the target product image on the target object.
6. The method according to claim 2, wherein The performing linear fitting on each rank value in the first rank value sequence to obtain a fitted rank value includes: Perform linear scoring on each rank value in the first rank value sequence by using the target rank value sequence of the target object to obtain the fitted rank value.
7. The method according to claim 6, characterized in that, The performing linear scoring on each rank value in the first rank value sequence by using the target rank value sequence of the target object to obtain the fitted rank value includes: Segment the first rank value sequence according to the number of target rank values in the target rank value sequence to obtain a plurality of rank value subsequences; Score each rank value subsequence by using the target rank value in the target rank value sequence, and use the target rank value corresponding to each rank value subsequence as the fitted rank value corresponding to each rank value in the rank value subsequence.
8. The method according to claim 7, wherein Determining the target level value of the target product image on the target object according to the fitting level value includes: Determining the fitting level value corresponding to the initial level value according to a level value mapping table, where the level value mapping table is used to represent the corresponding relationship between each level value in the first level value sequence and the fitting level value; Determining the fitting level value corresponding to the initial level value as the target level value of the target product image on the target object.
9. The method according to claim 1, wherein It further includes: Training an initial model using each sample product image in the sample product image set, where the initial model includes a convolutional neural network and a fully connected layer. The convolutional neural network is used to extract features in each sample product image by uniformly scaling the network depth, network width, and the resolution of each sample product image, and fuse features of different scales. The fully connected layer is used to perform regression analysis on the result output by the convolutional neural network; Obtaining the preset model when the loss function of the initial model reaches a preset convergence condition.
10. The method according to claim 9, wherein The convolutional neural network sequentially includes an EfficientNet network and a Feature Pyramid Network (FPN).
11. The method according to claim 9, characterized in that, The loss function is a mean squared error loss function.
12. The method according to any one of claims 9 to 11, characterized in that, The training of the initial model using each sample product image in the sample product image set includes: Annotating a reference level value for each sample product image in the sample product image set, where the granularity of the reference level value is greater than the granularity of the level values in the first level value sequence; Training the initial model based on the sample product images annotated with reference level values.
13. The method according to claim 1, characterized in that, The target object is a target defect, and the level value is used to represent the severity of the target defect.
14. A grading device for a target object, characterized in that, It includes: A first prediction unit for determining the level value of each sample product image in the sample product image set on the target object through a preset model to obtain a first level value sequence, where the level value is used to represent the level determination priority of the target object, and the preset model is trained by the sample product image set; A second prediction unit for determining the level value of the target product image on the target object through the preset model as the initial level value; A target level value determination unit for determining the target level value of the target product image on the target object according to the first level value sequence and the initial level value.
15. A level determination device for a target object, comprising a processor and a memory, characterized in that, The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 13 are implemented.
16. A computer-readable storage medium, characterized in that, Computer program instructions are stored in the computer-readable storage medium. When the computer program instructions are executed by the processor, the processor is prompted to implement the steps of the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method and device for training defect grading detection model, equipment and storage medium
CN113205176A
Training method of grading model, and object grading method and device
CN116824244A
Deep-learning-based method for managing quality of target, and system using same
WO2022255566A1