Target detection method, system and electronic equipment based on digital image
The deep learning model processes the detected images, which solves the problem of poor image quality in complex light source environments, and improves the accuracy and speed of product defect detection.
Patent Information
- Application Number
- CN202111214050.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The prior art is difficult to obtain high-quality digital images in complex light source environments, resulting in low accuracy and slow speed of product defect detection.
Deep learning related models are used to process the detected images, and feature maps of multiple different scales are obtained by downsampling, and tensor calculations are performed to determine the probability value of each pixel point containing the target, and finally the target detection result is obtained.
It improves the accuracy of the product quality inspection process, improves the quality inspection speed, and can more effectively identify product defects and watermark text.
Smart Images

Figure CN113888522B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pattern recognition, and in particular to a target detection method, system and electronic equipment based on digital images. Background Art
[0002] In the quality inspection process of industrial products, how to quickly and accurately detect the defective areas and related text areas of the products to be tested, the target detection method used is the key factor to improve the detection performance. In the production process of existing technologies, a relatively complex light source environment is arranged, which is not conducive to obtaining high-quality digital images of the products to be tested, resulting in poor detection of defective areas of the products to be tested, especially when detecting defects on complex surfaces, the shortcomings of accuracy and speed are more obvious.
[0003] It can be seen that there are still technical problems such as low accuracy and slow speed in the existing product quality inspection process. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a target detection method, system and electronic device based on digital images, which utilize deep learning related models to identify product defects and watermark text, etc., and obtain the probability value of each pixel containing the target by performing tensor calculation on multiple feature maps of different scales obtained by downsampling the image to be detected, and finally obtain the target detection result, thereby improving the accuracy of the quality inspection process and increasing the quality inspection speed.
[0005] In a first aspect, an embodiment of the present invention provides a method for object detection based on a digital image, the method comprising:
[0006] Acquire the image to be detected;
[0007] Input the image to be detected into the trained target detection model to determine the probability value of each pixel in the image to be detected containing the target; wherein the target detection model is used to downsample the image to be detected into three feature maps of different scales, and obtain the probability value of the target corresponding to each pixel in the image to be detected according to the output tensor of the feature map;
[0008] The target area of the image to be detected is determined according to the probability value that each pixel in the image to be detected contains the target, and the detection result of the target is determined according to the target area.
[0009] In some embodiments, the step of inputting the image to be detected into the trained target detection model to determine the probability value of each pixel in the image to be detected containing the target includes:
[0010] The target detection model is used to downsample the image to be detected, and the first feature map, the second feature map, and the third feature map of the image to be detected are obtained through tensor connection; wherein the side length of the first feature map is half of the side length of the second feature map; the side length of the second feature map is half of the side length of the third feature map;
[0011] The priori boxes corresponding to the first feature map, the second feature map, and the third feature map are set respectively, and the feature maps in the priori boxes are predicted to obtain the probability value of the target corresponding to each pixel in the image to be detected.
[0012] In some embodiments, the above-mentioned determining the target area of the image to be detected according to the probability value that each pixel in the image to be detected contains the target includes:
[0013] Compare the probability value of each pixel in the image to be detected containing the target with a preset threshold, and obtain the pixel whose probability value of the target in the image to be detected is greater than the preset threshold;
[0014] The pixel points whose probability values of the target in the image to be detected are greater than a preset threshold are connected in sequence to determine the target area of the image to be detected.
[0015] In some embodiments, determining the detection result of the target according to the target area includes:
[0016] Get the average probability value of all pixels in the target area containing the target;
[0017] The average probability value is compared with the preset probability threshold, and the detection result of the target is determined based on the comparison result.
[0018] In some embodiments, determining the detection result of the target according to the target area includes:
[0019] Get the bounding rectangle of the target area;
[0020] The area of the circumscribed rectangle is compared with the preset area threshold, and the detection result of the target is determined based on the comparison result.
[0021] In some implementations, if the target area includes characters, determining the detection result of the target according to the target area includes:
[0022] Get the rotation angle of the characters in the target area;
[0023] Determine the rotation matrix of the target area according to the rotation angle of the character;
[0024] The target area is rotated using a rotation matrix, and the target detection result is determined based on the rotated target area.
[0025] In some embodiments, the training process of the target detection model includes:
[0026] Acquire a sample image; wherein the sample image includes: a digital image of one or more defects on the workpiece surface, such as scratches, bruises, abrasions, pits, white spots, and character watermarks;
[0027] Input the sample image into the initialized convolutional neural network for training;
[0028] The loss value of the convolutional neural network is calculated according to the preset loss function; when the loss value meets the preset expected threshold, the training is stopped to obtain the target detection model.
[0029] In some embodiments, the convolutional neural network has a total of 252 layers, including: 23 augmentation layers, 72 batch normalization layers, 2 connection layers, 75 convolutional layers, 1 input layer, 72 activation layers, 2 upsampling layers, and 5 zero-filling layers;
[0030] Among them, the convolution layer, batch normalization layer, and activation layer are connected in sequence to obtain the DBL layer; there are 72 DBL layers in total;
[0031] The added layers include: 5 groups of blocks, namely res1, res2, res4 and 2 res8; the res1 block, res2 block, res4 block and res8 block contain 1, 2, 4 and 8 added layers respectively; the added layers are connected according to the res1 block, the res2 block, the first res8 block, the res4 block and the second res8 block; the zero padding layers are connected in sequence between the blocks of the added layers;
[0032] The input end of the added layer is connected to the input layer; the output end of the added layer is connected to the DBL layer;
[0033] The connection layer is recorded as: the first connection layer and the second connection layer; the upsampling layer is recorded as the first upsampling layer and the second upsampling layer; the input end of the first connection layer is connected to the first upsampling layer and the first res8 block respectively; the output end of the first connection layer is connected to the DBL layer; the input end of the second connection layer is connected to the second upsampling layer and the second res8 block respectively; the output end of the second connection layer is connected to the DBL layer.
[0034] In a second aspect, an embodiment of the present invention provides a target detection system based on a digital image, the system comprising:
[0035] An image acquisition unit, used for acquiring an image to be detected;
[0036] The target detection unit is used to input the image to be detected into the trained target detection model to determine the probability value of each pixel in the image to be detected containing the target; wherein the target detection model is used to downsample the image to be detected into feature maps of three different scales, and obtain the probability value of the target corresponding to each pixel in the image to be detected according to the output tensor of the feature map;
[0037] The target acquisition unit is used to determine the target area of the image to be detected according to the probability value that each pixel point in the image to be detected contains the target, and determine the detection result of the target according to the target area.
[0038] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, wherein when the processor executes the computer program, the steps of the digital image-based target detection method mentioned in the first aspect are implemented.
[0039] In a fourth aspect, an embodiment of the present invention further provides a computer-readable medium having a non-volatile program code executable by a processor, wherein the program code enables the processor to execute the digital image-based target detection method mentioned in the first aspect above.
[0040] The embodiments of the present invention bring the following beneficial effects:
[0041] The present invention provides a target detection method, system and electronic device based on digital images. The method first obtains an image to be detected; then the image to be detected is input into a trained target detection model to determine the probability value of each pixel in the image to be detected containing a target; wherein the target detection model is used to downsample the image to be detected to feature maps of three different scales, and obtain the probability value of the target corresponding to each pixel in the image to be detected according to the output tensor of the feature map; finally, the target area of the image to be detected is determined according to the probability value of each pixel in the image to be detected containing a target, and the detection result of the target is determined according to the target area. The method uses a deep learning related model to identify product defects and watermark text, etc., and obtains the probability value of each pixel containing a target by performing tensor calculation on the feature maps of multiple different scales obtained by downsampling the image to be detected, and finally obtains the target detection result, thereby improving the accuracy of the quality inspection process and improving the quality inspection speed.
[0042] Other features and advantages of the present invention will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by implementing the above-mentioned technology of the present invention.
[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are specifically listed below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 A flowchart of a digital image-based target detection method provided by an embodiment of the present invention;
[0046] Figure 2 A flowchart of a training process of a target detection model in a digital image-based target detection method provided by an embodiment of the present invention;
[0047] Figure 3 A schematic diagram of the structure of a convolutional neural network used for training a target detection model in a digital image-based target detection method provided in an embodiment of the present invention;
[0048] Figure 4 A flowchart of step S102 in the digital image-based target detection method provided by an embodiment of the present invention;
[0049] Figure 5 Another schematic diagram of the structure of a convolutional neural network used for training a target detection model in a digital image-based target detection method provided by an embodiment of the present invention;
[0050] Figure 6 A flow chart of determining a target area of an image to be detected according to a probability value that each pixel in the image to be detected contains a target in a target detection method based on a digital image provided by an embodiment of the present invention;
[0051] Figure 7 A flow chart of determining a detection result of a target according to a target area in a target detection method based on a digital image provided by an embodiment of the present invention;
[0052] Figure 8 Another flow chart of determining a detection result of a target according to a target area in the target detection method based on a digital image provided by an embodiment of the present invention;
[0053] Fig. 9 A flowchart of a third method of determining a detection result of a target according to a target area in a target detection method based on a digital image provided by an embodiment of the present invention;
[0054] Fig.10 A schematic diagram of the structure of a digital image-based target detection system provided by an embodiment of the present invention;
[0055] Fig.11 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0056] icon:
[0057] 1010 - image acquisition unit; 1020 - target detection unit; 1030 - target acquisition unit; 101 - processor; 102 - memory; 103 - bus; 104 - communication interface. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] In the quality inspection process of industrial products, how to quickly and accurately detect defective areas and related text areas of the products to be tested, and the target detection method used is a key factor in improving detection performance.
[0060] The hardware currently required for industrial AI quality inspection includes:
[0061] Industrial cameras: installed on the production line or integrated into the equipment, used to photograph parts to be inspected. On the one hand, they collect data sample sets, and on the other hand, they serve as data input for quality inspection operations.
[0062] Light source: illuminates parts and provides a stable light source to assist industrial cameras in shooting.
[0063] Server: Terminal device, including processor, graphics card, hard disk, network card, etc., used to receive data input from industrial cameras, recognize and process images through the AI algorithm within its system, and transmit the results back.
[0064] The current server software for industrial AI quality inspection includes:
[0065] Algorithm software: Using deep learning algorithm, the input is the image array matrix, and the output is the coordinates of the defective pixels in the image.
[0066] Interface software: An interactive interface presented to customers, with functions including quality inspection result display, database information interaction, algorithm software control parameters, etc.
[0067] In terms of hardware, with the continuous development of technology, setting up a more complex light source environment, configuring servers and racks on a compact and highly automated production line will be more costly, difficult to maintain, and have a long acceptance cycle. In terms of software, the existing deep learning algorithms have weak detection capabilities on edge devices and slow computing speeds. They are easier to handle simpler deep learning problems such as detection of presence and number of detections, but in the field of industrial quality inspection, the accuracy and speed of detecting defects on complex surfaces are relatively low.
[0068] It can be seen that the production process of products in the prior art usually has a relatively complex light source environment, which is not conducive to obtaining high-quality digital images of the products to be tested, resulting in poor detection of defective areas of the products to be tested, especially when detecting defects on complex surfaces. The shortcomings of accuracy and speed are more obvious. Therefore, the existing product quality inspection process still has technical problems of low accuracy and slow speed.
[0069] Based on this, an embodiment of the present invention provides a target detection method, system and electronic device based on digital images, which use deep learning related models to identify product defects and watermark text, etc., and obtain the probability value of each pixel containing the target by performing tensor calculation on multiple feature maps of different scales obtained by downsampling the image to be detected, and finally obtain the target detection result, thereby improving the accuracy of the quality inspection process and increasing the quality inspection speed.
[0070] To facilitate understanding of this embodiment, a digital image-based target detection method disclosed in an embodiment of the present invention is first introduced in detail.
[0071] See also Figure 1 The flowchart of a target detection method based on digital images is shown, and the specific steps of the method include:
[0072] Step S101, obtaining an image to be detected.
[0073] The image to be detected is obtained by shooting with an industrial camera installed on the production line, or by image frames in the video stream of the relevant monitoring in the production line. The image to be detected is a digital image. Generally speaking, the shooting area of the image to be detected is fixed, and the target area to be detected is included in the shooting area. When the shooting angle of the image to be detected is large, that is, when the target area occupies a very small area in the image to be detected, the image to be detected can be cropped to place the target area to be detected in a larger range of the image to be detected as much as possible, which can be beneficial to the detection accuracy.
[0074] Specifically, when detecting certain defects in the product to be tested, the image to be tested can be cropped according to the characteristics of the product. For example, the image area to be tested with a higher probability of scratches and bumps in the product can be cropped, and the cropped images are used as the images to be tested for scratches and bumps respectively.
[0075] Step S102, input the image to be detected into the trained target detection model to determine the probability value of each pixel in the image to be detected containing the target; wherein the target detection model is used to downsample the image to be detected to feature maps of three different scales, and obtain the probability value of the target corresponding to each pixel in the image to be detected according to the output tensor of the feature map.
[0076] The target detection model is a model in the field of deep learning. Specifically, the image to be detected is used as input data and is input through the relevant input interface of the target detection model. After being processed by the target detection model, the probability value of each pixel in the image to be detected containing the target is obtained.
[0077] After the image to be detected is input into the target detection model, the target detection model downsamples the image to be detected to feature maps of three different scales, and obtains the probability value of the target corresponding to each pixel in the image to be detected based on the output tensor of the feature map. The output tensor can be packaged according to the grids, types, and scores involved in feature maps of different sizes. For each grid, it contains 5 parameters, namely: the horizontal coordinate of the grid center point, the vertical coordinate of the grid center point, the grid width, the grid length, and the accuracy. In specific use, there are 80 categories for the COCO category, so the dimension of the output tensor of the target detection model is 3×(80+5)=255 layers.
[0078] For each grid, logistic regression can be used to score the targetness of the content enclosed in the grid, and the prior box can be selected for prediction based on the targetness score result, and finally the probability value of the target corresponding to each pixel in the image to be detected is obtained.
[0079] Step S103, determining the target area of the image to be detected according to the probability value that each pixel in the image to be detected contains the target, and determining the detection result of the target according to the target area.
[0080] After obtaining the probability value of each pixel in the image to be detected containing the target, the region with a larger probability value can be obtained by setting a threshold, and the region can be connected to obtain the target region; edge detection can also be used to obtain the edge lines of the region with a larger probability value to obtain the outline of the target region; the minimum circumscribed quadrilateral of the target region can also be obtained based on the outline of the target region. The target region can be extracted in the above manner, and the detection result of the target can be obtained based on the relevant attributes of the target region.
[0081] Specifically, when product defects such as scratches, bruises, abrasions, pits, and white spots are used as target detection, the brightness value at the defect is different from that on the workpiece surface. Therefore, after obtaining the target area, the average pixel value in the target area can be judged to determine whether the workpiece is scratched, and finally obtain the corresponding detection result.
[0082] When character watermarks are used as target detection, after obtaining the target area, the OCR algorithm can be used to extract the characters to obtain relevant characters, and the information displayed by the characters can be judged to obtain the final detection result.
[0083] It can be seen from the digital image-based target detection method provided in the above embodiment that the method utilizes deep learning related models to identify product defects and watermark text, etc., and obtains the probability value of each pixel containing the target by performing tensor calculations on multiple feature maps of different scales obtained by downsampling the image to be detected, and finally obtains the target detection result, thereby improving the accuracy of the quality inspection process and increasing the speed of quality inspection.
[0084] In some embodiments, the training process of the target detection model, such as Figure 2 As shown, including:
[0085] Step S201, obtaining a sample image; wherein the sample image includes: a digital image of one or more defects on the workpiece surface, such as scratches, bruises, abrasions, pits, white spots, and character watermarks.
[0086] The acquisition of sample images is related to the final output characteristics of the target detection model. For the target detection model that takes product defects as the target, the pre-saved digital images of scratches, bruises, abrasions, pits, and white spots are used as positive sample images for the training of the target detection model; digital images that do not contain the above defects can also be used as negative sample images for the training of the target detection model to improve the overall performance of the model. Similarly, for character watermarks as the target, various watermark images containing characters can be used as positive sample images. These images can contain various fonts, various rotated fonts, and other images.
[0087] Step S202: input the sample image into the initialized convolutional neural network for training.
[0088] In the specific implementation process, the convolutional neural network is as follows Figure 3As shown, it can be set to 252 layers, including: 23 augmentation layers, 72 batch normalization layers, 2 connection layers, 75 convolutional layers, 1 input layer, 72 activation layers, 2 upsampling layers, and 5 zero-filling layers. Among them, the convolution layer, the batch normalization layer, and the activation layer are connected in sequence to obtain the DBL layer; there are 72 DBL layers in total; the added layer includes: 5 groups of blocks: res1, res2, res4 and 2 res8; the res1 block, the res2 block, the res4 block, and the res8 block contain 1, 2, 4, and 8 added layers respectively; the added layer is connected according to the res1 block, the res2 block, the first res8 block, the res4 block, and the second res8 block; the zero padding layer is connected between the blocks of the added layer in sequence; the input end of the added layer is connected to the input layer; the output end of the added layer is connected to the DBL layer; the connection layer is recorded as: the first connection layer and the second connection layer; the upsampling layer is recorded as the first upsampling layer and the second upsampling layer; the input end of the first connection layer is connected to the first upsampling layer and the first res8 block respectively; the output end of the first connection layer is connected to the DBL layer; the input end of the second connection layer is connected to the second upsampling layer and the second res8 block respectively; the output end of the second connection layer is connected to the DBL layer.
[0089] The backbone network uses 5 res structures. n represents a number, res1, res2, ..., res8, etc., which means that this res_block contains n res_units, which are large components of the network. Based on the residual structure of ResNet, using this structure can make the network structure deeper to ensure better detection accuracy.
[0090] There is a tensor concatenation operation in the prediction branch. The implementation method is to concatenate the upsampling of the middle layer and a layer after the middle layer. It is worth noting that the tensor concatenation and the add operation of the Res_unit structure are different. Tensor concatenation will expand the dimension of the tensor, while add is just a direct addition and will not cause the tensor dimension to change.
[0091] In the convolutional neural network, each batch normalization layer BN is followed by an activation layer LeakyReLU. There are 2 upsampling and tensor concatenation operations, and 5 zero paddings correspond to 5 res_blocks. There are 75 convolutional layers in total, 72 of which are followed by DBL composed of BN and LeakyReLU. The outputs of three different scales correspond to three convolutional layers, and the number of convolution kernels in the last convolutional layer is 255. For the 80 categories of the COCO dataset: 3×(80+4+1)=255, 3 means that a grid cell contains 3 bounding boxes, 4 means the 4 coordinate information of the box, and 1 means the confidence.
[0092] Step S203, calculating the loss value of the convolutional neural network according to a preset loss function; when the loss value meets the preset expected threshold, the training is stopped to obtain the target detection model.
[0093] The loss function can be set according to the structure of the model, such as using the square error addition formula. For this convolutional neural network, the loss function should involve four key information: grid center point coordinates, grid size, classification, and accuracy. The loss function is determined by the characteristics of the above four types of information, and the weighted sum is finally obtained. For example, the total square error function can be used for the loss function of the grid center point coordinates; the binary cross entropy function is used for the loss functions of the other three types, and finally they are added together to obtain the final loss function.
[0094] In some embodiments, the step S102 of inputting the image to be detected into the trained target detection model to determine the probability value of each pixel in the image to be detected containing the target is as follows: Figure 4 As shown, including:
[0095] Step S401, down-sample the image to be detected using the target detection model, and obtain the first feature map, the second feature map, and the third feature map of the image to be detected through tensor connection; wherein the side length of the first feature map is half of the side length of the second feature map; the side length of the second feature map is half of the side length of the third feature map.
[0096] Specifically, for the image to be detected, it is finally mapped into an output tensor of three scales, representing the probability of the existence of various targets at various positions in the image.
[0097] Step S402, respectively set the priori boxes corresponding to the first feature map, the second feature map, and the third feature map, and predict the feature maps in the priori boxes to obtain the probability value of the target corresponding to each pixel in the image to be detected.
[0098] For example, for a 416*416 image to be detected, 3 prior boxes are set in each grid of the feature map of each scale, and there are a total of 13*13*3+26*26*3+52*52*3=10647 predictions. Each prediction is a (4+1+80)=85-dimensional vector, which contains the coordinates of the bounding box (4 values), the confidence of the bounding box (1 value), and the probability of the object category (for example, the standard coco dataset used for defect data experiments has 80 objects).
[0099] like Figure 5As shown in the figure, it is set that each grid unit predicts 3 boxes, so each box needs to have five basic parameters (x, y, w, h, confidence); x is the horizontal coordinate of the center of the box; y is the vertical coordinate of the center of the box; w is the width of the box; h is the height of the box; confidence is the prediction accuracy. The backbone network outputs 3 feature maps of different scales, as shown in the figure above, y1, y2, y3. The depth of y1, y2 and y3 is 255, and the side length is 13:26:52.
[0100] The feature size obtained for each prediction task is N×N×[3*(4+1+80)], where N is the grid size, 3 is the number of bounding boxes obtained for each grid, 4 is the number of bounding box coordinates, 1 is the target prediction value, and 80 is the number of categories. For the standard dataset COCO category, there are 80 categories of probabilities, so each box should output a probability for each category, that is, 3×(5+80)=255.
[0101] For the entire network, the feature map sizes of the multi-scale prediction output are y1: (13×13), y2: (26×26), and y3: (52×52). The network receives a (416×416) image, and downsamples it to 416 / 2^5=13 through 5 convolutions with a step size of 2. y1 outputs (13×13). The convolution layer of the second-to-last layer of y1 is upsampled (x2, up sampling) and then connected to the last feature map tensor of 26×26, and y2 outputs (26×26). The convolution layer of the second-to-last layer of y2 is upsampled (x2, up sampling) and then connected to the last feature map tensor of 52×52, and y3 outputs (52×52).
[0102] In some embodiments, the target area of the image to be detected is determined according to the probability value of each pixel in the image to be detected containing the target, such as Figure 6 As shown, including:
[0103] Step S601 , comparing the probability value of each pixel in the image to be detected containing the target with a preset threshold, and obtaining the pixel whose probability value of the target in the image to be detected is greater than the preset threshold.
[0104] By comparing with the preset threshold, the pixel points with a higher probability of containing the target in the image to be detected can be obtained, and these pixel points can be merged or connected to obtain the target area of the image to be detected, and finally the target detection result can be obtained.
[0105] Step S602 , sequentially connect the pixel points whose probability values of the target in the image to be detected are greater than a preset threshold to determine the target area of the image to be detected.
[0106] Specifically, the pixels that are greater than the preset threshold are connected in sequence to form a closed area. This area is the target area of the image to be detected, and the corresponding target is contained in the target area. After the target area is determined, the target detection results in the area are further judged according to the actual scene. The first is threshold screening. The network outputs the target detection result box and its confidence x. The confidence x is a decimal between 0 and 1. The higher the confidence x is, the more confident the algorithm is that there is a target there. Most of the boxes with low confidence can be removed through threshold screening. The second is size screening. Similarly, a size threshold is set to compare with the confidence of the detection box. If it is larger than the size, it can be detected. The third is area screening. During the detection, the position of the object to be detected in the image is fixed, such as a certain side of the engine casing. Therefore, a mask matrix for area screening can be made according to the process benchmark of the object to be detected. The matrix is a two-dimensional matrix of 0 and 1, and the length and width are the length and width of the image. According to whether the coordinates of the detection box are 0 or 1 in the mask matrix, it can be judged whether the box is in the area to be detected.
[0107] In some embodiments, the detection result of the target is determined according to the target area, such as Figure 7 As shown, including:
[0108] Step S701, obtaining the average probability value of all pixels in the target area containing the target.
[0109] After finding the target area, the probability values corresponding to all the pixels in the target area are summed up to obtain the average probability value of all the pixels in the target area containing the target, and the average probability value is used as the judgment basis.
[0110] Step S702, comparing the average probability value with a preset probability threshold, and determining the detection result of the target according to the comparison result.
[0111] Taking the scratched area of the product as the target area as an example, the target area may not necessarily fully display whether there are scratches. In the scenario where the scratch range is very small, such small scratches are mostly ignored in actual scenarios. Therefore, the average probability value of all pixels in the target area containing the target can be compared with the preset probability threshold, and the detection result of the target can be determined based on the comparison result. When the preset probability threshold is set high, it indicates that scratches with a very small range in the above scenario can be ignored; when the preset probability threshold is set low, it indicates that scratches with a very small range in the above scenario can be retained.
[0112] In some embodiments, the detection result of the target is determined according to the target area, such as Figure 8 As shown, including:
[0113] Step S801, obtaining the bounding rectangle of the target area.
[0114] Taking the product bump area as the target area as an example, the bump area is generally a relatively large area, and can be judged according to the area of the target area. In the specific implementation process, the threshold judgment can be made by the bounding rectangle of the bump area. The process of obtaining the bounding rectangle can be directly obtained by the relevant function in the relevant digital image processing, and will not be repeated again.
[0115] Step S802: compare the area of the circumscribed rectangle with a preset area threshold, and determine the detection result of the target according to the comparison result.
[0116] In actual scenarios, collisions with a very small range are mostly ignored. Therefore, the area of the circumscribed rectangle can be compared with the preset area threshold, and the detection result of the target can be determined based on the comparison result. When the preset probability area is set high, it means that collisions with a very small range in the above scenario can be ignored; when the preset area threshold is set low, it means that collisions with a very small range in the above scenario can be retained.
[0117] In some embodiments, if the target area contains characters, the detection result of the target is determined according to the target area, such as Fig. 9 As shown, including:
[0118] Step S901, obtaining the rotation angle of the characters in the target area.
[0119] The characters contained in the target area can be directly extracted through the relevant OCR algorithm, or the relevant character extraction model can be used to extract the characters, and finally the minimum enclosing rectangle of the character area is extracted. The characters contained in the extracted character area cannot be guaranteed to be non-rotated, so the rotation angle of the characters can be further obtained through the relevant character extraction model, such as rotation angle, flip angle, etc.
[0120] Step S902, determining the rotation matrix of the target area according to the rotation angle of the character.
[0121] After obtaining the rotation angle of the character, the rotation matrix of the target area is obtained according to the minimum circumscribed rectangle of the character area. The rotation matrix contains the corresponding rotation angle data, and an inverse transformation can be performed according to the rotation angle data to rotate the character to a normal angle.
[0122] Step S903: Rotate the target area using a rotation matrix, and determine the target detection result according to the rotated target area.
[0123] The characters in the target area are rotated using the rotation matrix to make them non-rotated, and then the non-rotated target area is used to determine the target detection result, i.e. character extraction. Since the characters are non-rotated at this time, text extraction at this time will improve the output accuracy of OCR.
[0124] It can be seen from the digital image-based target detection method provided in the above embodiment that the method utilizes deep learning related models to identify product defects and watermark text, etc., and obtains the probability value of each pixel containing the target by performing tensor calculations on multiple feature maps of different scales obtained by downsampling the image to be detected, and finally obtains the target detection result, thereby improving the accuracy of the quality inspection process and increasing the speed of quality inspection.
[0125] Corresponding to the above method embodiment, the embodiment of the present invention further provides a target detection system based on digital images, and its structural schematic diagram is shown as follows: Fig.10 As shown, the system includes:
[0126] An image acquisition unit 1010 is used to acquire an image to be detected;
[0127] The target detection unit 1020 is used to input the image to be detected into the trained target detection model to determine the probability value of each pixel in the image to be detected containing the target; wherein the target detection model is used to downsample the image to be detected into feature maps of three different scales, and obtain the probability value of the target corresponding to each pixel in the image to be detected according to the output tensor of the feature map;
[0128] The target acquisition unit 1030 is used to determine the target area of the image to be detected according to the probability value that each pixel in the image to be detected contains the target, and determine the detection result of the target according to the target area.
[0129] The target detection system based on digital images provided in the embodiment of the present invention has the same technical features as the target detection method based on digital images provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects. For the sake of brief description, for matters not mentioned in the embodiment part, reference can be made to the corresponding contents in the above method embodiment.
[0130] This embodiment also provides an electronic device, which is a structural diagram of the electronic device as shown in Fig.11 As shown, the device includes a processor 101 and a memory 102; wherein the memory 102 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the above-mentioned target detection method based on digital images.
[0131] Fig.11The electronic device shown further includes a bus 103 and a communication interface 104 , and the processor 101 , the communication interface 104 and the memory 102 are connected via the bus 103 .
[0132] The memory 102 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The bus 103 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0133] The communication interface 104 is used to connect to at least one user terminal and other network units through a network interface, and send the encapsulated IPv4 message or IPv4 message to the user terminal through the network interface.
[0134] The processor 101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 101. The above processor 101 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 102, and the processor 101 reads the information in the memory 102 and completes the steps of the method of the above embodiment in combination with its hardware.
[0135] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method of the above embodiment are executed.
[0136] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0137] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0139] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present invention can essentially or in other words, the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0140] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A target detection method based on digital images, It is characterized in that The method comprises: Acquire the image to be detected; Input the image to be detected into a trained target detection model to determine the probability value of each pixel in the image to be detected containing the target; wherein the target detection model is used to downsample the image to be detected into feature maps of three different scales, and obtain the probability value of the target corresponding to each pixel in the image to be detected according to the output tensor of the feature map; Determining a target area of the image to be detected according to a probability value that each pixel in the image to be detected contains the target, and determining a detection result of the target according to the target area; Determining the target area of the image to be detected according to the probability value that each pixel in the image to be detected contains the target, comprising: Compare the probability value of each pixel in the image to be detected containing the target with a preset threshold, and obtain the pixel whose probability value of the target in the image to be detected is greater than the preset threshold; Connecting pixel points whose probability values of the target in the image to be detected are greater than a preset threshold in sequence to determine the target area of the image to be detected; Determining a detection result of the target according to the target area includes: Obtaining an average probability value that all pixels in the target area contain the target; The average probability value is compared with a preset probability threshold, and the detection result of the target is determined according to the comparison result.
2. The target detection method based on digital image according to claim 1, It is characterized in that The step of inputting the image to be detected into a trained target detection model and determining the probability value of each pixel in the image to be detected containing the target comprises: Down-sampling the image to be detected by using the target detection model, and obtaining a first feature map, a second feature map, and a third feature map of the image to be detected respectively through tensor connection; wherein the side length of the first feature map is half of the side length of the second feature map; and the side length of the second feature map is half of the side length of the third feature map; The a priori boxes corresponding to the first feature map, the second feature map, and the third feature map are respectively set, and the feature maps in the a priori boxes are predicted to obtain the probability value of the target corresponding to each pixel in the image to be detected.
3. The target detection method based on digital image according to claim 1, It is characterized in that Determining a detection result of the target according to the target area includes: Get the bounding rectangle of the target area; The area of the circumscribed rectangle is compared with a preset area threshold, and the detection result of the target is determined according to the comparison result.
4. The target detection method based on digital image according to claim 1, It is characterized in that If the target area contains characters, determining the detection result of the target according to the target area includes: Obtaining the rotation angle of the character in the target area; Determining a rotation matrix of the target area according to a rotation angle of the character; The target area is rotated using the rotation matrix, and a detection result of the target is determined according to the rotated target area.
5. The target detection method based on digital image according to claim 1, It is characterized in that The training process of the target detection model includes: Acquire a sample image; wherein the sample image includes: a digital image of one or more defects on the workpiece surface, such as scratches, bruises, abrasions, pits, white spots, and character watermarks; Inputting the sample image into an initialized convolutional neural network for training; The loss value of the convolutional neural network is calculated according to a preset loss function; when the loss value meets a preset expected threshold, the training is stopped to obtain the target detection model.
6. The target detection method based on digital image according to claim 5, It is characterized in that The convolutional neural network has a total of 252 layers; Among them, there are: 23 augmentation layers, 72 batch normalization layers, 2 connection layers, 75 convolutional layers, 1 input layer, 72 activation layers, 2 upsampling layers, and 5 zero-filling layers; The convolution layer, the batch normalization layer, and the activation layer are sequentially connected to obtain a DBL layer; there are 72 DBL layers in total; The added layers include: 5 groups of blocks, namely res1, res2, res4 and 2 res8; the res1 block, the res2 block, the res4 block and the res8 block respectively include 1, 2, 4 and 8 added layers; the added layers are connected according to the res1 block, the res2 block, the first res8 block, the res4 block and the second res8 block; the zero filling layers are sequentially connected between the blocks of the added layers; The input end of the added layer is connected to the input layer; the output end of the added layer is connected to the DBL layer; The connection layer is recorded as: the first connection layer and the second connection layer; the upsampling layer is recorded as the first upsampling layer and the second upsampling layer; the input end of the first connection layer is respectively connected to the first upsampling layer and the first res8 block; the output end of the first connection layer is connected to the DBL layer; the input end of the second connection layer is respectively connected to the second upsampling layer and the second res8 block; the output end of the second connection layer is connected to the DBL layer.
7. A target detection system based on digital images, It is characterized in that The system comprises: An image acquisition unit, used for acquiring an image to be detected; A target detection unit, used to input the image to be detected into a trained target detection model, and determine the probability value of each pixel in the image to be detected containing the target; wherein the target detection model is used to downsample the image to be detected into feature maps of three different scales, and obtain the probability value of the target corresponding to each pixel in the image to be detected according to the output tensor of the feature map; A target acquisition unit, used to determine a target area of the image to be detected according to a probability value that each pixel in the image to be detected contains the target, and determine a detection result of the target according to the target area; In the process of determining the target area of the image to be detected according to the probability value that each pixel in the image to be detected contains the target, the target acquisition unit is further used to: compare the probability value that each pixel in the image to be detected contains the target with a preset threshold, and obtain the pixel points in the image to be detected whose probability value of the target is greater than the preset threshold; connect the pixel points in the image to be detected whose probability value of the target is greater than the preset threshold in sequence, and determine the target area of the image to be detected; In the process of determining the detection result of the target based on the target area, the target acquisition unit is also used to: obtain the average probability value that all pixels in the target area contain the target; compare the average probability value with a preset probability threshold, and determine the detection result of the target based on the comparison result.
8. An electronic device, It is characterized in that include: processor and storage device; The storage device stores a computer program, which, when executed by the processor, executes the steps of the digital image-based target detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image processing method and device, electronic device, and storage medium
CN109255784A
Multi-target intelligent imaging and recognition device and method
CN110751206A
Automatic correction method and system for code spraying characters
CN112651401A