Target detection method and apparatus, electronic device, and storage medium
By calculating the grayscale contrast inside and outside the object detection box, filtering out the detection box that meets the conditions, solving the problem of over-detection in deep learning methods and achieving more accurate object detection.
Patent Information
- Application Number
- PCT/CN2024/136709
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-12-04
- Publication Date
- 2025-07-17
AI Technical Summary
Existing deep learning methods have problems with over-detection in object detection, and cannot effectively distinguish between visually obvious and inconspicuous defects, resulting in more over-detection phenomena.
By using the object detection model to identify the grayscale representative values in and around the target object detection box, calculate the grayscale contrast, and filter out the detection box that meets the preset conditions as the detection result to avoid over-detection.
Effectively distinguish visually obvious and inconspicuous defects, improve detection accuracy, avoid over-inspection, and improve production efficiency and product quality.
Smart Images

Figure CN2024136709_17072025_PF_FP_ABST
Abstract
Description
Target detection method, device, electronic device and storage medium
[0001] Related documents
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on January 9, 2024, with application number 202410029363.1 and invention name “Target Detection Method, Device, Electronic Device and Storage Medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of image processing technology. Specifically, the present application relates to a target detection method, device, electronic device and storage medium.
[0004] Background of the Invention
[0005] With the development of emerging technologies such as artificial intelligence and deep learning, using these methods to detect target objects has become a very common method. Deep learning methods have the advantages of high detection rates, strong generalization performance, and the ability to scale up at a low cost once the model is stabilized. However, current deep network methods may suffer from the problem of over-detection. For example, in a defect detection scenario, if a slight light gray stain appears on a gray interface, this defect may be detected, but this defect is not visually obvious. In other words, current deep network methods cannot effectively distinguish between visually obvious defects and less obvious defects, resulting in a high incidence of over-detection. Summary of the Invention
[0006] The purpose of the embodiments of the present application is to solve the problem of over-detection in target detection.
[0007] Each embodiment of the present application provides a target detection method, which is performed by an electronic device and includes:
[0008] Identifying a target object in the image to be detected by using a target detection model to obtain at least one target object detection frame;
[0009] For each target object detection frame, determining a first grayscale representative value for pixels within the target object detection frame, determining a second grayscale representative value for pixels within a background area of the target object detection frame, and determining a grayscale contrast ratio between the first grayscale representative value and the second grayscale representative value, wherein the background area is an area within a predetermined range surrounding the target object detection frame;
[0010] The target object detection frame whose grayscale contrast meets the preset conditions in the at least one target object detection frame is used to determine the detection result of the image to be detected.
[0011] Each embodiment of the present application provides a target detection device, which includes:
[0012] A target object recognition module is used to identify the target object in the image to be detected by using the target detection model to obtain at least one target object detection frame;
[0013] a contrast determination module, configured to determine, for each target object detection frame, a first grayscale representative value of pixels within the target object detection frame, determine a second grayscale representative value of pixels within a background area of the target object detection frame, and determine a grayscale contrast between the first grayscale representative value and the second grayscale representative value, wherein the background area is an area within a predetermined range surrounding the target object detection frame;
[0014] The detection result determination module determines the target object detection frame whose grayscale contrast meets the preset conditions in the at least one target object detection frame as the detection result of the image to be detected.
[0015] Each embodiment of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the data processing method provided in the embodiment of the present application.
[0016] Each embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the data processing method provided in the embodiment of the present application is implemented.
[0017] Each embodiment of the present application further provides a computer program product, including a computer program, which implements the data processing method provided in the embodiment of the present application when the computer program is executed by a processor.
[0018] The target detection method, device, electronic device and storage medium provided in the embodiments of the present application use the grayscale contrast of the first grayscale representative value of the pixels within the target object detection frame and the second grayscale representative value of the pixels in the background area within a predetermined range around the target object detection frame to distinguish the difference between the foreground and background of the target object detection frame, and are used to filter out detection results with unclear visual differences to avoid the problem of over-detection.
[0019] BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The following drawings are only some examples of the technical solutions of the present invention, and the present invention is not limited to the features shown in the drawings. In the following drawings, similar reference numerals represent similar elements:
[0021] FIG1 is a schematic diagram of a flow chart of a target detection method provided in an embodiment of the present application;
[0022] FIG2 is a schematic diagram of an original background area provided in an embodiment of the present application;
[0023] FIG3 is a schematic diagram of another original background area provided in an embodiment of the present application;
[0024] FIG4 is a schematic diagram of a background area around a predetermined range provided by an embodiment of the present application;
[0025] FIG5 is a schematic diagram of an object detection model provided in an embodiment of the present application;
[0026] FIG6 is a schematic diagram of a feature extraction network provided in an embodiment of the present application;
[0027] FIG7 is a schematic diagram of a feature fusion network provided in an embodiment of the present application;
[0028] FIG8 is a schematic diagram of a defect detection method process provided in an embodiment of the present application;
[0029] FIG9 is a schematic diagram of an application scenario of a target detection solution provided in an embodiment of the present application;
[0030] FIG10 is a schematic structural diagram of a target detection device provided in an embodiment of the present application;
[0031] FIG11 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.
[0032] Modes for Carrying Out the Invention
[0033] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0034] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an" and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to that the element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B".
[0035] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be described in detail in each embodiment with reference to the accompanying drawings.
[0036] First, several terms involved in this application are introduced and explained:
[0037] (1) Object Detection: Object detection focuses on specific target objects and aims to separate the target object of interest from the background.
[0038] (2) Defect Detection: Defect detection, also known as anomaly detection, is a type of target detection. Its main task is to determine whether an image has defects.
[0039] (3) RGB: Represents the colors of the three channels red, green, and blue. A variety of colors are obtained by varying and superimposing the three color channels red (R), green (G), and blue (B). This is the storage and display format for photographs taken by a camera. In the embodiments of this application, this also refers to general photographs, as distinguished from photometric stereo normal vector images.
[0040] (4) Grayscale value: This refers to the brightness value of each pixel in the image, usually expressed as an integer from 0 to 255. A larger grayscale value indicates a higher brightness of the pixel, while a smaller grayscale value indicates a lower brightness of the pixel. For example, a grayscale value of 0 represents black, and a grayscale value of 255 represents white. Grayscale values between 0 and 255 represent different grayscale levels.
[0041] (5) OK / NG (No Good): Indicates whether the workpiece has passed or failed quality inspection.
[0042] (6) Gaussian distribution: also known as normal distribution, is a common continuous probability distribution in fields such as mathematics, physics, and engineering.
[0043] Existing object detection methods can suffer from over-detection. For example, in industrial manufacturing, AI and deep learning technologies are often used to inspect product appearance to ensure consistency, yield, and safety, while fully automating production line quality inspections. However, current AI and deep learning methods are unable to meet the business needs of differentiating between obvious and less obvious defects, resulting in a high rate of over-detection. Simply raising the confidence threshold can also result in missed detections of obvious defects.
[0044] The target detection method, device, electronic device and storage medium provided in this application are intended to solve the above technical problems in the prior art.
[0045] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0046] An embodiment of the present application provides a target detection method, as shown in FIG1 , comprising:
[0047] Step S101: Acquire an image to be detected.
[0048] In the embodiment of the present application, the image to be detected refers to the image to be detected as the target object. Taking the defect detection scenario as an example, the image to be detected can refer to a product image, and the product in the image can be an industrial product to be detected for defects, such as industrial parts or components, etc. The target object refers to the defect, but is not limited to this. Among them, the image to be processed can be received from other devices, read locally, or taken in real time, and the embodiment of the present application is not limited here. As an example, a capture device (such as a camera) can be used to take a picture of a part or the whole of the product to obtain the image to be detected.
[0049] Step S102: Identify the target object in the image to be detected by using the target detection model to obtain at least one target object detection frame.
[0050] In the embodiments of the present application, there is no specific limitation on the type of target detection model. For example, the Cascade RCNN (Cascade Region-Convolutional Neural Network) model, the Faster RCNN (Faster Region-Convolutional Neural Network) model, the DETR (Detection Transformer) model, etc. may be used, but are not limited thereto.
[0051] In the embodiment of the present application, the target detection model needs to have a full understanding of the foreground and background of the image to be detected, and determine the description of the target object to obtain the category information (classification) and location information (localization) of the target object.
[0052] The identified target object can be represented by a rectangular detection box to represent the area including the target object. That is, the target object detection box represents the area in the image to be detected where the identified target object is located. The position information can be represented by the coordinates of the rectangular detection box, for example, it can be represented as the center coordinates and width and height (x, y, w, h) of the target object detection box. The category information of the target object can be represented as a confidence probability, that is, the probability that the target object belongs to various predetermined categories. The predetermined category with the highest probability is selected as the category information of the target object.
[0053] In various embodiments, if the target detection model detects that there are multiple target objects in the image to be detected, the output of the target detection model can be a list, and each item in the list uses an array to give the category information and location information of the detected target object.
[0054] Step S103: For each target object detection frame, determine the first grayscale representative value of the pixels in the target object detection frame, determine the second grayscale representative value of the pixels in the background area of the target object detection frame, and determine the grayscale contrast between the first grayscale representative value and the second grayscale representative value, wherein the background area is an area in a predetermined range around the target object detection frame.
[0055] In an embodiment of the present application, the first grayscale representative value can represent the overall grayscale condition of all pixels (also referred to as foreground) within the target object detection frame. Among them, the method for determining the first grayscale representative value (also referred to as foreground grayscale value) can be set according to actual conditions. As an example, if the grayscale values of all pixels within the target object detection frame are relatively single, the grayscale value of one or part of the pixels in the foreground can be used to represent the first grayscale representative value. As another example, if the grayscale values of all pixels within the target object detection frame are relatively large and relatively evenly distributed, the average value of the grayscale values of all pixels in the foreground can be used to represent the first grayscale representative value, etc. The embodiment of the present application does not limit this.
[0056] In the embodiments of the present application, for a target object detection frame, its original background area is the portion of the area other than the detection frame of the same category. For example, in the defect detection scenario, as shown in Figure 2, for defect area 1 22 of category A, the shaded portion is the original background area 21, including defect area 3 24 of category B, but excluding defect area 23 of category A. Alternatively, as shown in Figure 3, for defect area 3 24 of category B, the shaded portion is the original background area 21, including defect area 1 22 of category A and defect area 23 of category A.
[0057] However, taking into account the complexity of the background, for example, in actual industrial defect detection applications, the image to be processed that needs to be detected is usually a complete device or component, which may include various elements such as processing workpieces, machines, light sources, mechanical structures, etc. In this case, using the entire image as the background to calculate the second grayscale representative value will be interfered with by factors such as other structural parts. Therefore, in an embodiment of the present application, in order to reduce interference in the background, a predetermined range around the target object detection frame is set as the background area, for example, it can be an area of n = 10 pixels around the target object detection frame, as shown in the shaded part in Figure 4. The probability of this area containing other target objects is small, so it can be used as a background area. In actual applications, those skilled in the art can set the range of the area according to actual conditions, and the embodiment of the present application is not limited here.
[0058] In the embodiment of the present application, the second grayscale representative value can represent the overall grayscale of all pixels in the background area. The method for determining the second grayscale representative value (also referred to as the background grayscale value) can also be set according to actual conditions. For example, the grayscale value of one or a portion of the pixels in the background area can be used to represent the second grayscale representative value, or the average grayscale value of all pixels in the background area can be used to represent the second grayscale representative value, etc., which is not limited in the embodiment of the present application.
[0059] In various embodiments, after obtaining the first grayscale representative value and the second grayscale representative value, the grayscale contrast between the first grayscale representative value and the second grayscale representative value can be evaluated, that is, the contrast between the target object and the background can be evaluated. The greater the grayscale contrast, the more prominent the target object is relative to the background. The method for determining the grayscale contrast can also be set according to actual circumstances, such as calculating the difference between the first grayscale representative value and the second grayscale representative value, or calculating the ratio of the difference, etc., which is not limited in the embodiments of the present application.
[0060] In the embodiment of the present application, such processing is performed on all target object detection frames, and the grayscale contrast corresponding to each target object detection frame is finally obtained.
[0061] Step S104: determining a target object detection frame whose grayscale contrast meets a preset condition in at least one target object detection frame as a detection result of the image to be detected.
[0062] This step can also be understood as business logic post-processing, which executes business-side logical judgments. The post-processing inputs are the location information and category information output by the object detection model, as well as the grayscale contrast corresponding to each target object's detection box. The output is a binary judgment of whether each target object is detected or filtered. For defect detection scenarios, a binary judgment of whether the workpiece is OK or NG can also be directly output.
[0063] In an embodiment of the present application, based on the grayscale contrast corresponding to at least one target object detection frame, it can be determined whether the target object in each target object detection frame meets the detection requirements, so that it can be effectively distinguished whether each target object is over-detected. As an example, taking the defect detection scenario as an example, it can effectively distinguish between visually obvious defects and visually inconspicuous defects. For example, if a dark foreign matter appears on a white background, its grayscale contrast will be very large and it will be detected. If a slight dark gray stain appears on a gray interface, its grayscale contrast will be very small and it may be filtered out. In this way, visible defects on the workpiece can be effectively detected, and defects that are not easy to detect can be avoided, thereby improving production efficiency and product quality.
[0064] The object detection method provided in the embodiments of the present application calculates the grayscale contrast between the identified foreground and background areas to distinguish the difference between the foreground and background of the target object detection frame. This is used to filter out detection results where the visual difference is not obvious, thus avoiding the problem of over-detection. In addition, this grayscale contrast can also assist in business logic decision-making.
[0065] In various embodiments of the present application, step S103 of determining the first grayscale representative value of the pixels within the target object detection frame may specifically include:
[0066] Step SA1: Based on the position information of the target object detection frame, determine the weight corresponding to each pixel in the target object detection frame. In other words, determine the weight corresponding to each pixel according to the position of each pixel in the target object detection frame.
[0067] The weight corresponding to each pixel represents the influence of the pixel on the first grayscale representative value.
[0068] In the embodiment of the present application, it is taken into account that it is difficult for the target object in the target object detection frame to occupy the entire rectangular detection frame. As an example, in actual industrial defect detection, the actual defect shape is usually irregular, and the target detection model used in step 102 outputs a rectangular detection frame. This leads to a certain difference between the actual defect shape and the target object detection frame. If the grayscale value in the target object detection frame is calculated directly, it may cause an error between the calculated foreground grayscale value and the grayscale value on the real target object. In order to solve this problem, corresponding weights can be set for each pixel in the target object detection frame to increase the influence of the real target object part on the first grayscale representative value and reduce the influence of the difference part on the first grayscale representative value.
[0069] In each embodiment, based on the position information of the target object detection frame, the size of the target object detection frame and the number of pixels, rows, columns, etc. in the target object detection frame can be obtained, so as to determine which pixels need to be set with weights.
[0070] Among them, the weight corresponding to each pixel in the target object detection frame can be preset or calculated according to certain rules, which is not limited in the embodiment of the present application.
[0071] Step SA2: Obtain the grayscale value of each pixel in the target object detection frame.
[0072] In an embodiment of the present application, if the image to be processed is an RGB image, the image to be processed can be first grayscaled to obtain a grayscale image corresponding to the image to be processed, and then the grayscale value of each pixel within each target object detection frame can be obtained from the grayscale image. The method of grayscale conversion can be selected based on actual needs and is not limited in this embodiment of the present application. If the image to be processed is a grayscale image, the grayscale value of each pixel within each target object detection frame can be directly obtained from the image to be processed.
[0073] Step SA3: Based on the weight corresponding to each pixel in the target object detection frame, perform weighted averaging on the grayscale value of each pixel in the target object detection frame to obtain a first grayscale representative value of the pixel in the target object detection frame.
[0074] Among them, the weight corresponding to each pixel in the target object detection frame can be a two-dimensional matrix with the same size as the foreground. Assuming that each element in the matrix is represented by W1(x,y), this weight matrix can be used to adjust the contribution of each pixel in the target object detection frame to the first grayscale representative value.
[0075] In various embodiments, based on the weights corresponding to the pixels in the target object detection frame, the grayscale values of the pixels in the target object detection frame are weighted averaged. For the grayscale value P(x, y) of each pixel in the target object detection frame, the first grayscale representative value Fg1 can be calculated using the following formula:
[0076] Alternatively, if the weights corresponding to each pixel in the target object detection frame are a normalized two-dimensional matrix, assuming that each element in the matrix is represented by W2(x,y), the first grayscale representative value Fg1 can be calculated by the following formula: Fg1 = sum(P(x,y)*W2(x,y))
[0077] These formulas represent the weighted average of the foreground grayscale value and its weight matrix. In this way, the average grayscale value of the foreground can be calculated more accurately, thereby more accurately comparing the grayscale contrast.
[0078] As described above, there may be certain differences between the actual target object shape and the rectangular detection frame. This may be because the rectangular detection frame includes part of the background. Based on this, in an embodiment of the present application, the weights corresponding to the pixels in the target object detection frame can be set so that the weight corresponding to the center pixel in the target object detection frame is greater than the weight corresponding to the edge pixel in the target object detection frame. Among them, the center and edge in this application can be considered relative. For two positions, the weight of the pixel closer to the center is greater than the weight of the pixel closer to the edge.
[0079] In the embodiments of the present application, an optional implementation is provided for step SA1. In various embodiments, the weight corresponding to the center pixel within the target object detection frame and the weight corresponding to the edge pixel within the target object detection frame can be gradually reduced. The specific reduction step size can be set according to actual conditions and is not limited in the embodiments of the present application.
[0080] In the embodiments of the present application, another optional implementation is provided for step SA1. In each embodiment, the following may be included:
[0081] Step SA11: The center position of the target object detection frame is used as the origin of the two-dimensional Gaussian distribution function, and the distance between the position of each pixel in the target object detection frame and the center position is used as the variable of the two-dimensional Gaussian distribution function to calculate the Gaussian distribution two-dimensional matrix.
[0082] Step SA12: Determine each element value in the Gaussian distribution two-dimensional matrix as the weight corresponding to each pixel in the target object detection frame.
[0083] The embodiment of the present application aims to provide a weight corresponding to each pixel in the target object detection frame, so that the weight of the pixels in the center of the rectangular detection frame area (usually the center of the target object part) is larger, and the weight of the pixels at the edge (which may be the background) is smaller.
[0084] Specifically, the weights corresponding to each pixel in the target object detection frame can be set to the Gaussian distribution weight matrix gaussian_weights. For the embodiment of the present application, based on the position information of the target object detection frame, the center position of the target object detection frame can be used as the origin of the two-dimensional Gaussian distribution function, and each pixel in the target object detection frame can be used as a variable of the two-dimensional Gaussian distribution function to calculate the Gaussian distribution two-dimensional matrix.
[0085] Among them, in two-dimensional space, the formula of the Gaussian distribution function is as follows:
[0086] Where x and y are coordinates in two-dimensional space, μ x and μ yis the mean value. In the embodiment of the present application, μ x and μ y is set to 0 because we want the center pixel of the target object detection box to have the largest weight), and σ is the standard deviation.
[0087] Therefore, the Gaussian distribution two-dimensional matrix can be obtained as:
[0088] Similarly, we generate a Gaussian distribution 2D matrix, gaussian_weights, of the same size as the foreground, with the largest weight at the center and decreasing toward the periphery. By assigning the Gaussian distribution 2D matrix to the weights corresponding to each pixel within the target object detection frame, we can use this weight matrix to adjust the contribution of each pixel within the target object detection frame to the first grayscale representative value, so that the pixel at the center of the target object detection frame has a greater influence on the first grayscale representative value.
[0089] Based on the Gaussian distribution two-dimensional matrix, the grayscale values of each pixel are weighted averaged to obtain the first grayscale representative value of the pixel in the target object detection frame. The method and effect are similar to step SA3 and will not be repeated here.
[0090] In general, the purpose of the above processing of the foreground of the target object detection frame is to calculate the grayscale contrast between the foreground and background as accurately as possible while taking into account the difference between the real target object morphology and the rectangular detection frame.
[0091] In the embodiment of the present application, step S103 of determining the second grayscale representative value of the pixels in the background area may specifically include:
[0092] Step SB1: Clustering each pixel in the background area based on color to obtain at least two pixel sets. This step clusters each pixel in the background area based on color information to obtain at least two cluster centers and clustering results for each pixel.
[0093] In an embodiment of the present application, in order to avoid the problem that when the target object is located in a complex texture area, the area within a predetermined range around it may still be interfered with by the complex texture, thereby causing inaccurate calculation of the second grayscale representative value, a clustering algorithm is used to filter the grayscale value of the set background area to reduce the interference of the complex texture.
[0094] Among them, the color information uses grayscale values, that is, grayscale values are used to achieve clustering, or the color information page can use RGB values, for example, directly using RGB values as triples to achieve clustering, which is not limited in the embodiments of the present application.
[0095] In an embodiment of the present application, those skilled in the art may select a suitable clustering algorithm according to actual conditions to cluster the pixels in the background area. For example, a K-Means clustering algorithm (an unsupervised learning algorithm, mainly used for data clustering) may be used, or other algorithms such as hierarchical clustering, DBSCAN (Density-Based Spatial Clustering of Applications with Noise, a density-based clustering algorithm) or deep learning algorithms such as autoencoders, DEC (Deep Embedding Clustering), DBC (Density-Based Clustering), etc. The embodiment of the present application is not limited here, and ultimately obtains at least two cluster centers and clustering results of each pixel (i.e., at least two pixel sets, each pixel set is a set of pixels belonging to a cluster).
[0096] In practical applications, the appropriate number of clusters can be determined based on the specific situation. For example, if the background texture is very complex, the number of clusters can be increased. Increasing the number of clusters can make the division of the second grayscale representative value more detailed, thereby better processing complex background textures. If the background texture is relatively simple and there is not much interference, a smaller number of clusters can be set. For example, two clusters are sufficient to avoid the calculation offset of the second grayscale representative value that may be caused by too many clusters, while reducing the complexity of the calculation and the running time.
[0097] Step SB2: Based on the grayscale values of at least two cluster centers and the clustering results for each pixel, the grayscale value of each pixel in the background area, which serves as intermediate computational data, is updated. This modification is performed on a copy of the data used for subsequent calculations, rather than altering the original image. In various embodiments, this step can be omitted, and other equivalent computational methods can be used.
[0098] That is, in the embodiment of the present application, based on the clustering result of each pixel, the cluster to which the pixel belongs is determined, and the grayscale value of the pixel is then updated to the grayscale value of the corresponding cluster center. The updated grayscale value of each pixel is one of the grayscale values of at least two cluster centers. For example, assuming the number of clusters is set to 2, the grayscale values of the background area are divided into two categories.
[0099] Step SB3: Determine a second grayscale representative value of the pixels in the background area based on the at least two pixel sets.
[0100] In this way, a relatively accurate second grayscale representative value can be obtained, which is used to calculate the grayscale contrast between the foreground and background areas of the target object.
[0101] Specifically, this step may include at least one of the following methods:
[0102] (1) Based on the updated grayscale values of each pixel in the background area, a target grayscale value corresponding to the largest number of pixels among the grayscale values of at least two cluster centers is determined, and the target grayscale value is used as the second grayscale representative value of the pixels in the background area.
[0103] As an example, assuming that the grayscale values of the background area are divided into two categories, it can be determined which grayscale value of the two categories corresponds to a larger number of pixels, and the grayscale value category with the largest number of pixels can be selected as the second grayscale representative value of the background area.
[0104] (2) Calculate the average value of the updated grayscale values of each pixel in the background area as the second grayscale representative value of the pixels in the background area.
[0105] In various embodiments, the sum of the grayscale values of all pixels in the background area can be calculated and then divided by the number of pixels to obtain a second grayscale representative value of the pixels in the background area. As previously mentioned, in various embodiments, the second grayscale representative value can also be calculated without relying on updating the pixel value. For example, as an equivalent calculation method, the ratio of the number of pixels in each of the at least two pixel sets to the total number of pixels in the background area can be used as the weight of each pixel set, and the second grayscale representative value can be determined as the weighted average of the clustering target grayscale values corresponding to each of the at least two pixel sets calculated using the weights in the clustering process.
[0106] By using any of the above methods, the average grayscale value of the background area can be calculated more accurately, thereby more accurately comparing the grayscale contrast. In actual applications, those skilled in the art can choose which second grayscale representative value calculation method to use in which scenarios according to actual conditions, and the embodiments of the present application are not limited here. For example, taking an actual industrial defect detection application as an example, if a defect is located near the edge of a component, the background area within a predetermined range around it may include the edge of the component and the environmental area outside the component. By using the former method of selecting the grayscale value with the largest number of pixels, the interference of the environmental area on the grayscale contrast calculation can be effectively filtered out, effectively improving the judgment accuracy.
[0107] In the embodiments of the present application, an optional implementation is provided for step SB1. In each embodiment, the following may be included:
[0108] Step SB11: Select a predetermined number of pixels in the background area as initial cluster centers.
[0109] This step can be understood as an initialization step, selecting K (i.e., a predetermined number) points as initial cluster centers. These points can be randomly selected from the pixels in the background area.
[0110] Step SB12: For each pixel in the background area, calculate the distance from the pixel to a predetermined number of cluster centers, and obtain a predetermined number of clusters (i.e., pixel sets) by adding each pixel to the set corresponding to the cluster center with the smallest distance.
[0111] That is, each pixel in the background area is assigned a cluster. For each pixel, the distance from it to each cluster center is calculated, and then it is assigned to the nearest cluster center. In this way, K clusters can be obtained.
[0112] Step SB13: Repeat the following clustering steps until the predetermined conditions are met:
[0113] Step SB131: For each obtained cluster, calculate the mean of all pixels in the cluster and use the mean as the updated cluster center.
[0114] Step SB132: For each pixel in the background area, calculate the distance from the pixel to each updated cluster center, and obtain a predetermined number of updated clusters by adding each pixel to the pixel set corresponding to the updated cluster center with the smallest distance.
[0115] Among them, those skilled in the art can set the predetermined conditions according to actual conditions, such as the cluster center no longer changes, or reaching a preset maximum number of iterations, etc., which are not limited in this embodiment of the present application. Ultimately, a clustering result of K cluster centers and which cluster each pixel belongs to can be obtained.
[0116] In the embodiments of the present application, clustering the grayscale values of the background area can reduce the interference of complex textures and improve the accuracy of the calculation of the second grayscale representative value. Using the above clustering algorithm to filter the grayscale values of the background area provides an effective method in target detection scenarios such as industrial defect detection, helping to accurately calculate the second grayscale representative value and thus more accurately detect the target object.
[0117] In the embodiment of the present application, in step S103, determining the grayscale contrast between the first grayscale representative value and the second grayscale representative value may specifically include:
[0118] Step SC1: Calculate the absolute value of the difference between the first grayscale representative value and the second grayscale representative value.
[0119] Step SC2: taking the absolute value of the difference as the grayscale contrast between the first grayscale representative value and the second grayscale representative value.
[0120] In the embodiment of the present application, the grayscale contrast difference between the foreground and background can be calculated by calculating the absolute difference between the first grayscale representative value and the second grayscale representative value. The calculation method is as follows: contrast = |fg_mean - bg_mean|
[0121] Where fg_mean is the first grayscale representative value (e.g., the weighted average grayscale value of the foreground), and bg_mean is the second grayscale representative value (e.g., the average grayscale value of the background). This absolute difference can be used to evaluate the contrast between the foreground and background. The greater the contrast, the more obvious the target object.
[0122] In the embodiments of the present application, an optional implementation is provided for step S104. In each embodiment, the following may be included:
[0123] Step S1041: for each target object detection frame, determine whether the grayscale contrast corresponding to the target object detection frame is greater than a predetermined threshold.
[0124] Step S1042: Filter out target object detection frames whose grayscale contrast is less than a predetermined threshold. The filtering operation can be to delete the corresponding data or to mark the data. In various embodiments, this step can also be omitted and step S1043 can be directly executed.
[0125] Step S1043: taking the position information and target object category information of the target object detection frame with a grayscale contrast greater than a predetermined threshold in the at least one target object detection frame as the detection result of the image to be detected.
[0126] In an embodiment of the present application, relevant detection parameters can be preset, that is, a predetermined threshold, such as 50, and target objects whose absolute value of grayscale contrast is greater than the predetermined threshold can be detected. For example, obvious defects (such as severe dirt, but not limited to this) will be detected, and visually similar target objects, such as inconspicuous defects (such as slight dirt, but not limited to this) will be filtered out, thereby realizing the function of accurate target object detection.
[0127] In various embodiments, if an image to be processed contains at least one target object detection frame with a grayscale contrast greater than a predetermined threshold, it indicates that a detection result has been found in the image to be processed. For example, in a defect detection scenario, if a defect is found in a workpiece in the image to be processed, an "NG" quality inspection result may be output. The defect location and defect category, such as dirt, discoloration, scratches, damage, or dents, may also be output, but are not limited to these.
[0128] For an image to be processed that does not contain a target object detection frame with a grayscale contrast greater than a predetermined threshold, it indicates that no detection result is found for the image to be processed. Taking the defect detection scenario as an example, if there are no defects in the workpiece in the image to be processed, a quality inspection result of OK can be output for it.
[0129] The target detection method provided in the embodiment of the present application can avoid detecting target objects that are not visually obvious, thereby improving the accuracy of detection.
[0130] In the embodiments of the present application, an optional implementation is provided for step S102. In each embodiment, the following may be included:
[0131] Step S1021: Obtain a grayscale image of the image to be detected.
[0132] Step S1022: Concatenate the grayscale image of the image to be detected and the image to be detected to obtain an input image for the target detection model.
[0133] Step S1023: Use the target detection model to extract image features of the input image and obtain an initial pre-selected area. Based on the image features, perform multiple regression and classification processes on the initial pre-selected area to obtain the position information and target object category information of at least one target object detection frame.
[0134] In the embodiment of the present application, the target detection model based on deep learning is designed to achieve the detection and discrimination of defects under the conditions of general imaging visibility and general distinguishability.
[0135] Among them, the number of input channels of the target detection model is four, that is, the color RGB image corresponding to the image to be detected has three channels, and the grayscale image of the differential contrast of the image to be detected has one channel, which are concatenated into a four-channel input image.
[0136] Specifically, as shown in Figure 5, an object detection model is used to extract image features from an input image. In Figure 5, I represents the input image, couv represents the backbone network, which can be, but is not limited to, an HRNetV2P network, a ResNet network, or a Swin Transformer network. BO represents the obtained initial pre-selected region, which can also be processed by a neural network, such as, but not limited to, an RPN (Region Proposal Network). Pool represents a local region feature extractor (e.g., performing ROI (region of interest) pooling operations), H1-H3 represent network heads, B1-B3 represent target regions (location information of target object detection boxes) obtained by performing hierarchical regression processing on the initial pre-selected regions, and C1-C3 represent classification results obtained by performing hierarchical classification processing on the initial pre-selected regions. Each level of regression and classification processing is performed based on the previous level. This cascaded detection method can improve detection accuracy.
[0137] In an embodiment of the present application, the backbone network in the target detection model can adopt the feature extraction network shown in Figure 6 and the feature fusion network shown in Figure 7. The feature extraction network has a strong ability to extract panoramic information and is extremely effective in detecting workpiece defects in industrial scenarios. Specifically, the feature extraction network can gradually increase the features of different scales obtained by downsampling, and can process features of multiple scales in parallel, and share feature information multiple times in features of various scales to obtain feature maps of multiple scales that contain sufficient semantic information and texture information. The feature maps of multiple scales are then fused through the feature fusion network to obtain image features of the input image, making the image features more accurate both spatially and semantically.
[0138] In the embodiments of the present application, the target detection network can be trained end-to-end. In each embodiment, supervised training can be performed, and the training set contains annotation information of the target object position and category. Taking the defect detection scenario as an example, all parts to be detected including various types of defects can be prepared, and the defect positions and defect categories can be annotated for the part images, that is, the bounding boxes of defects such as dirt and discoloration that need to be detected are marked to train the model. During reasoning, it is only necessary to load the corresponding trained network parameters, directly input the image, and obtain the model output result. That is, during reasoning, the input of the target detection network is the concatenation of RGB color images and contrast difference grayscale images, and the output is the bounding boxes of dirt, discoloration, scratches, damage, dents, etc. and their categories (or the confidence probability of the category).
[0139] In the embodiment of the present application, the loss function of the target detection model training is the same as the loss function of the original model, including the regression loss term, which is used to fit the regression of the detection box coordinates (x, y, w, h), and the classification loss term, which is used to supervise and constrain the category of the detection box. This process can be expressed as follows: Loss = L_reg(x, y, w, h) + L_cls(c)
[0140] Where L_reg is the regression loss term used to fit the regression of the detection box, and L_cls is the classification loss term used to supervise and constrain the category of the detection box. x, y, w, and h represent the center coordinates and width and height of the detection box, respectively, and c represents the category of the detection box.
[0141] During the training process, a large number of labeled samples are used to train the model, allowing it to learn how to accurately identify and locate defects. This process requires a lot of computing resources and time, but once the model is trained, it can quickly and accurately detect defects in actual production.
[0142] In addition, the model training process of the embodiment of the present application is only for training the target detection model, and the steps in the rest of the process are fixed and do not contain learnable parameters. Therefore, when the technical solution provided by the embodiment of the present application is applied on a new machine, it does not require a lot of complicated parameter adjustment work and has good compatibility. This is because only a threshold needs to be set to determine whether defects need to be detected based on the grayscale contrast of the detection frame. This method is simpler and easier to implement than existing deep learning methods. It does not require additional manual annotation and can be applied to actual scenarios that require rapid deployment and iteration.
[0143] In general, the target detection method provided by the embodiment of the present application can effectively detect the target object by jointly using a deep neural network model and a business logic post-processing module, as well as calculating the grayscale contrast of the foreground and background. In particular, in target detection, noise can be filtered by foreground Gaussian weighting and background clustering, which takes into account both the difference between the real target object morphology and the rectangular detection frame, and the actual needs of the business side, more accurately calculating the foreground-background contrast, and assisting in decision-making. Therefore, it has high application value in actual production.
[0144] Based on at least one of the above embodiments, in the embodiments of the present application, a defect detection scenario is taken as an example, and an example process of a defect detection method is provided through FIG8 . In each embodiment, the following steps may be included:
[0145] Step S8.1: Input an image to be inspected. This image is typically a high-definition image of a workpiece and may contain various possible defects. This image is fed into a deep learning object detection network (specifically, a defect detection network). This neural network is pre-trained to identify and locate defects in the image and output several defect detection boxes, each representing a possible defect area.
[0146] Step S8.2: Process each defect detection frame separately. First, calculate the foreground grayscale value (i.e., the first grayscale representative value) within the defect detection frame. The foreground is usually the defect itself, and its color and brightness can be represented by calculating the average grayscale value or by weighting the foreground Gaussian distribution.
[0147] Step S8.3: Calculate the background grayscale value (i.e., the second grayscale representative value) around the defect detection frame. The background is usually the normal part of the workpiece, and its grayscale value average can also be calculated or the background can be clustered and filtered to remove noise grayscale.
[0148] Step S8.4: Calculate the absolute difference between the grayscale values of the foreground and background. This absolute difference can be used to evaluate the contrast between the defect and the background. The greater the contrast, the more obvious the defect.
[0149] Step S8.5: Perform this process on all defect detection frames, and finally obtain the grayscale contrast of each defect detection frame.
[0150] Step S8.6: The grayscale contrast of all defect detection frames is fed into the post-processing module. In this module, a threshold is set based on business logic. If the grayscale contrast of a detection frame exceeds this threshold, it is considered an NG defect and must be detected. The defect's location and category information are then output from the defect detection model. If the grayscale contrast is below this threshold, it is considered an ignorable defect and can be filtered out.
[0151] The defect detection method provided in the embodiment of the present application pre-processes the image to be processed and constructs a corresponding deep neural network (defect detection network) to detect various types of surface defects such as dirt, discoloration, scratches, breakage, and dents on the workpiece. The noise is filtered by weighting the foreground Gaussian distribution and clustering the surrounding background, and the absolute value of the grayscale difference is calculated to accurately calculate the foreground-background contrast of the detection frame. This contrast can be used to distinguish between visually obvious defects and visually inconspicuous defects. It can not only improve the detection rate of visually obvious defects (for example, by lowering the score threshold to improve the detection rate of the neural network, and then using this solution to filter out visually inconspicuous defects), but also assist in the decision-making of business logic.
[0152] In various embodiments, the target detection scheme provided by the embodiments of the present application can also be completed by the collaboration of multiple computer devices or devices with computing capabilities. For example, different computer devices or devices each complete part of the steps of each method provided by the implementation of this application. As an example, as shown in Figure 9, it can be completed by the collaboration of an acquisition device 901, a computing device 902, and a display device 903 connected via a network. Among them, the acquisition device 901 is used to acquire images and obtain images to be detected. The computing device 902 is used to process the images acquired by the acquisition device 901, such as performing defect detection and judgment. The display device 903 is used to display the defect detection results obtained by the computing device 902, which may include the defect location and defect type. Among them, the acquisition device 901 may include but is not limited to image acquisition devices such as cameras and webcams. In some embodiments, when the computing device 902 has a display function, the display device 903 can be a display in the computing device 902.
[0153] The grayscale contrast-based target detection method provided in the embodiments of the present application, such as the industrial defect detection method, can be applied in the fields of defect detection and product quality control, such as steel plate defect detection, floor defect detection, PCB (Printed Circuit Board) defect detection, film defect detection, lamp bead defect detection, metal bar end face detection, fabric wrinkle grade assessment, and automated quality inspection instruments for surface defects of industrial parts.
[0154] The target detection method provided in the embodiments of the present application has the following technical effects:
[0155] 1) The network structure is clear, and each module has good generalization capabilities. Experiments on popular industrial quality inspection datasets have demonstrated that this solution has superior and stable performance, with a high recall rate for defects and a low pass rate.
[0156] 2) The algorithm logic is clear, stable and controllable, and the calculation results of each branch can be visualized, making it easy to quickly locate problems when the algorithm works abnormally.
[0157] 3) Provides better interpretability: By calculating grayscale differences, an intuitive feature can be obtained. This feature can assist in the business decision-making process. The contrast difference is intuitive and easy to adjust.
[0158] 4) High grayscale contrast calculation accuracy: This solution can achieve highly accurate grayscale contrast calculations even within rectangular detection frames. This eliminates the need for complex segmentation network operations and annotations (for a polygonal defect, the segmentation network needs to annotate multiple points (>10) based on the defect, which greatly reduces the speed of annotation and model training iterations. Furthermore, the segmentation model typically consumes more time and memory during inference. These factors make segmentation models difficult to apply in actual industrial inspections), greatly improving the practicality of industrial inspection deployments.
[0159] In various embodiments, the present application provides a target detection device. As shown in FIG10 , the target detection device 100 may include: an image acquisition module 1001, a target object recognition module 1002, a contrast determination module 1003, and a detection result determination module 1004, wherein:
[0160] The image acquisition module 1001 is used to acquire the image to be detected;
[0161] The target object recognition module 1002 is configured to recognize a target object in an image to be detected by using a target detection model to obtain at least one target object detection frame;
[0162] The contrast determination module 1003 is configured to determine, for each target object detection frame, a first grayscale representative value of pixels within the target object detection frame, determine a second grayscale representative value of pixels within a background area of the target object detection frame, and determine a grayscale contrast between the first grayscale representative value and the second grayscale representative value, wherein the background area is an area within a predetermined range surrounding the target object detection frame.
[0163] The detection result determination module 1004 determines the target object detection frame whose grayscale contrast meets the preset conditions in at least one target object detection frame as the detection result of the image to be detected.
[0164] In various embodiments, when used to determine the first grayscale representative value of the pixels within the target object detection frame, the contrast determination module 1003 is specifically configured to:
[0165] determining a weight corresponding to each pixel in the target object detection frame according to a position of each pixel in the target object detection frame, and performing weighted averaging on the grayscale values of each pixel in the target object detection frame based on the weight corresponding to each pixel in the target object detection frame to obtain a first grayscale value and a first grayscale representative value of the pixel in the target object detection frame;
[0166] Or for:
[0167] Determining, based on the position information of the target object detection frame, a weight corresponding to each pixel in the target object detection frame, where the weight corresponding to each pixel represents an influence of the pixel on the first grayscale representative value;
[0168] Get the grayscale value of each pixel in the target object detection box;
[0169] Based on the weight corresponding to each pixel, the grayscale value of each pixel is weighted averaged to obtain the first grayscale representative value of the pixel in the target object detection frame.
[0170] In various embodiments, the weight corresponding to the center pixel in the target object detection frame is greater than the weight corresponding to the edge pixel in the target object detection frame.
[0171] In various embodiments, when the contrast determination module 1003 is used to determine the weight corresponding to each pixel in the target object detection frame based on the position information of the target object detection frame, it is specifically used to:
[0172] The center position of the target object detection frame is used as the origin of the two-dimensional Gaussian distribution function, and the distance between the position of each pixel in the target object detection frame and the center position is used as the variable of the two-dimensional Gaussian distribution function to calculate the Gaussian distribution two-dimensional matrix;
[0173] The values of each element in the Gaussian distribution two-dimensional matrix are determined as the weights corresponding to each pixel in the target object detection frame.
[0174] In various embodiments, when used to determine the second grayscale representative value of pixels in the background area, the contrast determination module 1003 is specifically configured to:
[0175] At least two pixel sets are obtained by clustering each pixel in the background area according to color;
[0176] A second grayscale representative value of the pixels in the background area is determined based on the at least two pixel sets.
[0177] In various embodiments, the contrast determination module 1003 is configured to determine the second grayscale representative value of the pixels in the background area based on the updated grayscale value of each pixel in the background area in at least one of the following ways:
[0178] Determining a pixel set with the largest number of pixels in at least two pixel sets, and using a clustering target grayscale value corresponding to the pixel set with the largest number of pixels in the at least two pixel sets in the clustering process as a second grayscale representative value;
[0179] An average value of the updated grayscale values of the pixels in the background area is calculated as the second grayscale representative value of the pixels in the background area.
[0180] In various embodiments, when the contrast determination module 1003 is used to cluster the pixels in the background area based on the color information of the pixels in the background area, it is specifically used to:
[0181] Selecting a predetermined number of pixels in the background area as initial cluster centers;
[0182] For each pixel in the background area, the distance from the pixel to a predetermined number of cluster centers is calculated, and a predetermined number of clusters are obtained by adding each pixel to the set corresponding to the cluster center with the smallest distance thereto.
[0183] Repeat the following clustering steps until the predetermined conditions are met:
[0184] For each obtained cluster, calculate the mean of all pixels in the pixel set and use the mean as the updated cluster center;
[0185] For each pixel in the background area, the distance from the pixel to each updated cluster center is calculated, and a predetermined number of updated pixel sets are obtained by adding each pixel to the pixel set corresponding to the updated cluster center with the smallest distance thereto.
[0186] In various embodiments, when determining the grayscale contrast between the first grayscale representative value and the second grayscale representative value, the contrast determination module 1003 is specifically configured to:
[0187] calculating an absolute value of a difference between the first grayscale representative value and the second grayscale representative value;
[0188] The absolute value of the difference is taken as the grayscale contrast between the first grayscale representative value and the second grayscale representative value.
[0189] In various embodiments, when the detection result determination module 1004 is used to determine the detection result of the image to be detected based on the grayscale contrast corresponding to at least one target object detection frame, it is specifically used to:
[0190] For each target object detection frame in the at least one target object detection frame, determining whether a grayscale contrast corresponding to the target object detection frame is greater than a predetermined threshold;
[0191] The position information and target object category information of the target object detection frame with a grayscale contrast greater than a predetermined threshold in the at least one target object detection frame are used as the detection result of the image to be detected.
[0192] In various embodiments, when the target object recognition module 1002 is configured to use the target detection model to recognize the target object in the image to be detected and obtain at least one target object detection frame, it is specifically configured to:
[0193] Obtain the grayscale image of the image to be detected;
[0194] Concatenate the grayscale image of the image to be detected with the image to be detected to obtain the input image of the target detection model;
[0195] The target detection model is used to extract image features of the input image and obtain an initial pre-selected area. Based on the image features, multiple regression and classification processes are performed on the initial pre-selected area to obtain the position information and target object category information of at least one target object detection box.
[0196] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device and the beneficial effects produced, please refer to the description of the corresponding method shown in the previous text, and will not be repeated here.
[0197] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, and the processor executes the above computer program to implement the steps of the above method embodiments.
[0198] In an optional embodiment, an electronic device is provided, as shown in FIG11 , and the electronic device 1100 shown in FIG11 includes: a processor 1101 and a memory 1103. The processor 1101 and the memory 1103 are connected, such as through a bus 1102. In various embodiments, the electronic device 1100 may further include a transceiver 1104, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 1104 is not limited to one, and the structure of the electronic device 1100 does not constitute a limitation on the embodiments of the present application.
[0199] Processor 1101 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1101 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0200] Bus 1102 may include a path for transmitting information between the aforementioned components. Bus 1102 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Bus 1102 may be divided into an address bus, a data bus, a control bus, and so on. For ease of illustration, FIG11 shows only one thick line, but this does not indicate that there is only one bus or only one type of bus.
[0201] The memory 1103 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation here.
[0202] The memory 1103 is used to store the computer program for executing the embodiments of the present application, and the execution is controlled by the processor 1101. The processor 1101 is used to execute the computer program stored in the memory 1103 to implement the steps shown in the above method embodiments.
[0203] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0204] An embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0205] The terms "first," "second," "third," "fourth," "1," "2," and the like (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that shown or described in the drawings.
[0206] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0207] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0208] The above are only optional implementation methods for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. A target detection method, characterized in that Executed by an electronic device, the method includes: Identifying target objects in the image to be detected by using a target detection model to obtain at least one target object detection box; For each target object detection box, determining a first grayscale representative value of the pixels within the target object detection box, determining a second grayscale representative value of the pixels within the background area of the target object detection box, and determining the grayscale contrast between the first grayscale representative value and the second grayscale representative value, where the background area is an area within a predetermined range around the target object detection box; Determining the target object detection boxes with grayscale contrast meeting a preset condition among the at least one target object detection box as the detection result of the image to be detected.
2. The object detection method according to claim 1, wherein Determining the first grayscale representative value of the pixels within the target object detection box includes: Determining the weight influence corresponding to each pixel according to the position of each pixel within the target object detection box; Based on the weights corresponding to each pixel within the target object detection box, performing weighted averaging on the grayscale values of each pixel within the target object detection box to obtain the first grayscale representative value of the pixels within the target object detection box.
3. The object detection method according to claim 2, wherein The weight corresponding to the central pixel within the target object detection box is greater than the weight corresponding to the edge pixel within the target object detection box.
4. The object detection method according to claim 3, characterized in that, Determining the weight corresponding to each pixel includes: Taking the central position of the target object detection box as the origin of the two-dimensional Gaussian distribution function, taking the distance between the position of each pixel within the target object detection box and the central position as the variable of the two-dimensional Gaussian distribution function, and calculating to obtain a two-dimensional Gaussian distribution matrix; Determining each element value in the two-dimensional Gaussian distribution matrix as the weight corresponding to each pixel within the target object detection box.
5. The object detection method according to claim 1, wherein Determining the second grayscale representative value of the pixels within the background area includes: Clustering each pixel within the background area according to color to obtain at least two pixel sets; Determining the second grayscale representative value based on the at least two pixel sets.
6. The object detection method according to claim 5, wherein Determining the second grayscale representative value based on the at least two pixel sets includes at least one of the following methods: Determining the pixel set with the largest number of pixels among the at least two pixel sets, and taking the clustering target grayscale value corresponding to the pixel set with the largest number of pixels during the clustering process as the second grayscale representative value; Taking the ratio of the number of pixels in each pixel set among the at least two pixel sets to the total number of pixels within the background area as the weight of each pixel set, and determining the second grayscale representative value as the weighted average of the clustering target grayscale values corresponding to the at least two pixel sets calculated using the weights.
7. The object detection method according to claim 5, wherein Clustering each pixel within the background area according to color to obtain at least two pixel sets includes: Selecting a predetermined number of pixels within the background area as initial clustering centers; For each pixel within the background area, calculating the distance from the pixel to the predetermined number of clustering centers, and obtaining a predetermined number of pixel sets by adding each pixel to the set corresponding to the clustering center with the smallest distance. Repeat the following steps until a predetermined condition is met: For each resulting set of pixels, calculate the mean of all pixels in the set of pixels, and use the mean as the updated clustering center; For each pixel in the background region, calculate the distance from the pixel to each updated clustering center, and obtain a predetermined number of updated sets of pixels by adding each pixel to the set of pixels corresponding to the updated clustering center with the smallest distance thereto.
8. The object detection method according to claim 1, wherein Determine the gray-scale contrast between the first gray-scale representative value and the second gray-scale representative value, including: Calculate the absolute value of the difference between the first gray-scale representative value and the second gray-scale representative value; Use the absolute value of the difference as the gray-scale contrast between the first gray-scale representative value and the second gray-scale representative value.
9. The object detection method according to any one of claims 1-8, characterized in that Determine, as the detection result of the image to be detected, the object detection frame with the gray-scale contrast meeting a preset condition in the at least one object detection frame, including: For each object detection frame in the at least one object detection frame, determine whether the gray-scale contrast corresponding to the object detection frame is greater than a predetermined threshold; Use the position information and object category information of the object detection frame with the gray-scale contrast greater than the predetermined threshold in the at least one object detection frame as the detection result of the image to be detected.
10. The object detection method according to any one of claims 1-8, characterized in that, Use an object detection model to identify an object to be detected in the image to be detected, and obtain at least one object detection frame, including: Obtain a grayscale image of the image to be detected; Concatenate the grayscale image of the image to be detected with the image to be detected to obtain an input image of the object detection model; Use the object detection model to extract image features of the input image, and obtain an initial preselected region. Based on the image features, perform multiple regression processes and classification processes on the initial preselected region to obtain the position information and object category information of at least one object detection frame.
11. A target detection device, characterized in that, Including: An object recognition module, configured to identify an object to be detected in the image to be detected by using an object detection model, and obtain at least one object detection frame; A contrast determination module, configured to, for each object detection frame, determine a first gray-scale representative value of pixels in the object detection frame, determine a second gray-scale representative value of pixels in the background region of the object detection frame, and determine the gray-scale contrast between the first gray-scale representative value and the second gray-scale representative value, where the background region is a region within a predetermined range around the object detection frame; A detection result determination module, which determines, as the detection result of the image to be detected, the object detection frame with the gray-scale contrast meeting a preset condition in the at least one object detection frame.
12. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-10.
Citation Information
Patent Citations
An infrared target detection and tracking method
CN109902578A
Sea surface ship target detection method
CN109961065A
Image segmentation weak labeling method using geometric shape layering
CN113436221A
Projection picture color correction method and device, projection equipment and storage medium
CN116320334A
Boundary box determination method and device, equipment, storage medium and program product
CN116645500A
Cited By
Pavement precision three-dimensional data real-time processing method
CN121563978A
Image edge real-time processing method and system for surface exposure photocuring 3D printing
CN121767680A