Target image detection method, device, equipment, vehicle and storage medium
Through the binarization processing method of adaptive threshold and region of interest, the sensitivity of target detection is improved, and the problem of low calibration accuracy of circum-view cameras during after-sales maintenance is solved, and a higher calibration success rate and accuracy is achieved.
Patent Information
- Application Number
- CN202210699176.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-06-20
AI Technical Summary
During after-sales maintenance, the calibration accuracy of the surround-view camera is low, mainly due to the uneven lighting conditions and dirty calibration scenes, resulting in the failure of target detection.
Adaptive thresholds are used to perform binarization of target images, and the target image is processed by obtaining multiple binarization thresholds, combining the pixel distribution and characteristics of the region of interest, multiple binarized target images are obtained, and the detection results are fused to improve the sensitivity of target detection.
The sensitivity of target detection is improved, thereby improving the success rate and accuracy of calibration, and solving the problem of low calibration accuracy in after-sales maintenance.
Smart Images

Figure CN115049740B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and more specifically, to a target image detection method, device, equipment, vehicle, and storage medium. Background Art
[0002] Before autonomous vehicles leave the factory or undergo maintenance, their surround-view camera parameters must be calibrated. Calibration accuracy directly impacts camera performance. During calibration, the vehicle is placed at a dedicated calibration station, where its surround-view cameras capture images of targets placed around the station. The captured images are then used to locate and identify the targets. However, low calibration accuracy is a common technical issue during after-sales maintenance. Summary of the Invention
[0003] This application proposes a target image detection method, device, equipment, vehicle and storage medium for improving the calibration accuracy of a surround-view camera to a certain extent.
[0004] A first aspect of an embodiment of the present application provides a target image detection method, the method comprising: acquiring a target image; based on the target image, acquiring an adaptive threshold corresponding to the target image, the adaptive threshold comprising a plurality of binarization thresholds; based on the plurality of binarization thresholds, binarizing the target image to obtain a plurality of binarized target images; performing target detection on each of the obtained binarized target images to obtain a detection result for each binarized target image; and fusing the detection results of the plurality of binarized target images to obtain a target detection result.
[0005] A second aspect of an embodiment of the present application provides a target image detection device, which includes: a first acquisition module for acquiring a target image; a second acquisition module for acquiring an adaptive threshold corresponding to the target image based on the target image, wherein the adaptive threshold includes multiple binarization thresholds; a first processing module for binarizing the target image based on the multiple binarization thresholds to obtain multiple binarized target images; a first detection module for performing target detection on each of the obtained binarized target images to obtain a detection result for each binarized target image; and a first fusion module for fusing the detection results of the multiple binarized target images to obtain a target detection result.
[0006] A third aspect of an embodiment of the present application provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the target image detection method described in the first aspect above are implemented.
[0007] A fourth aspect of an embodiment of the present application provides a vehicle, comprising: a processor, a memory, and a computer program stored on the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the target image detection method described in the first aspect above are implemented.
[0008] A fifth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the target image detection method described in the first aspect are implemented.
[0009] In an embodiment of the present application, multiple binarization thresholds corresponding to the acquired target image are obtained, the target image is binarized based on the multiple binarization thresholds to obtain multiple binarized target images, the target in each binarized target image is detected, a detection result for each binarized target image is obtained, and the detection results of the multiple binarized target images are fused to obtain a target detection result. This can improve the sensitivity of detecting targets from target images to a certain extent, thereby improving the calibration success rate and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 A top view of a calibration scene for a surround view camera provided according to an embodiment of the present application is shown.
[0012] Figure 2 An automatic driving control system provided according to an embodiment of the present application is shown.
[0013] Figure 3 A flow chart of a target image detection method provided according to an embodiment of the present application is shown.
[0014] Figure 4 A target image provided according to an embodiment of the present application is shown.
[0015] Figure 5 Shows the Figure 4 The target image shown is the image after outlining the region of interest.
[0016] Figure 6 A specific flow chart of step S302 in the target image detection method provided according to an embodiment of the present application is shown.
[0017] Figure 7 Shows the Figure 5 The image shown is a binarized target image after binarization processing.
[0018] Figure 8 A specific flow chart of step S305 in the target image detection method provided according to an embodiment of the present application is shown.
[0019] Figure 9 A specific flow chart of step S702 in the target image detection method provided according to an embodiment of the present application is shown.
[0020] Figure 10 A schematic structural diagram of a target image detection device provided according to an embodiment of the present application is shown.
[0021] Figure 11 A schematic structural diagram of an electronic device provided according to an embodiment of the present application is shown.
[0022] Figure 12 A schematic structural diagram of a vehicle provided according to an embodiment of the present application is shown.
[0023] Figure 13 A storage medium for storing or carrying program codes for implementing the target image detection method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0025] Vehicles with autonomous driving functions are usually equipped with panoramic image monitoring systems (commonly known as surround-view cameras and surround-view monitoring systems). Surround-view cameras are arranged at the front, rear and both sides of the vehicle body to capture images of all fields of view around the vehicle at all times while the vehicle is driving, and to perform driving planning and / or driving assistance based on the collected images.
[0026] Before a vehicle leaves the factory or undergoes maintenance, it may be necessary to calibrate the surround-view camera's parameters. These parameters include internal and external parameters. Internal parameters refer to device parameters such as focal length, optical center, and lens distortion. External parameters refer to the transformation relationship between the pixel coordinate system and the world coordinate system. Calibration of external parameters is a key technical issue. Calibration accuracy directly affects camera performance.
[0027] like Figure 1 As shown, a top view of the calibration scene of the surround view camera provided by the embodiment of the present application is shown. Figure 1As shown, when setting up the calibration scene, multiple target areas (such as Figure 1 Target areas 101-110 are shown, and at least one target with a predetermined pattern is arranged in each target area. For example, circular targets 113 and square targets 114 (also known as checkerboard targets) are arranged in target area 101. During calibration, the vehicle to be calibrated is placed on a dedicated calibration station 111, with the left front wheel of the vehicle positioned at position 112. The vehicle's surround-view camera captures images of the pre-placed targets around it, and the targets in the captured target images are located and identified.
[0028] The relative positional relationship between each target and the calibrated vehicle at workstation 111 is known (referred to as a priori positional relationship). For example, a vehicle coordinate system (or world coordinate system) is established with the calibrated vehicle at workstation 111 as the reference system. The first world coordinate of each target in this vehicle coordinate system is known. The vehicle coordinate system is a coordinate system used to describe the relative positional relationship between objects surrounding the vehicle and the vehicle itself.
[0029] During calibration, based on the target identified and located in the target image, the relative positional relationship between the identified target and the calibrated vehicle at workstation 111 (referred to as the a posteriori positional relationship) is calculated. For example, the pixel coordinates of the target identified from the target image in the pixel coordinate system of the target image are located, and based on a preset transformation relationship between the pixel coordinates and the world coordinates, the second world coordinates of the detected target in the world coordinate system are calculated. The pixel coordinate system is a coordinate system used to describe the absolute position of pixels in an image.
[0030] For the same target, the calculated posterior relative position relationship is compared with the prior relative position relationship, for example, the first world coordinate is compared with the second world coordinate. If the deviation between the two is greater than the maximum allowable deviation, it is confirmed that the parameters of the surround-view camera have deviated, and the parameters of the surround-view camera are updated, for example, the preset transformation relationship is corrected to reduce the deviation between the first world coordinate and the second world coordinate as much as possible, so that the position relationship of the target relative to the vehicle body output by the surround-view camera is more accurate.
[0031] However, the inventors of the present application have discovered that a technical problem of low calibration accuracy often occurs during calibration during after-sales maintenance.
[0032] This technical problem has long plagued many after-sales service personnel. While they calibrate returned vehicles according to the pre-determined calibration process, they often encounter problems with the calibrated vehicles failing to correctly identify surrounding objects, meaning the calibration accuracy falls short of expectations. The inventors of this application have investigated this issue and discovered that while factory-calibrated vehicles rarely experience substandard calibration, this phenomenon primarily occurs with returned vehicles calibrated at various after-sales centers.
[0033] After further research into the calibration process at after-sales centers, the inventors of this application discovered that the direct cause of substandard calibration accuracy during after-sales center calibration was occasional target detection failure during target recognition and detection. This meant that targets in the target image could sometimes not be identified, or targets could be identified in areas where they were not present. As can be seen from the foregoing, accurate target recognition in the target image is crucial for accurately updating the surround-view camera parameters. This failure directly reduced calibration accuracy.
[0034] The inventors of this application once again visited various after-sales centers for on-site investigations and found that the main reasons for the low quality of the target images taken during calibration in various after-sales centers are the following: there is dirt in the calibration area of the calibration scene in the after-sales center, which causes the dirt and target in the calibration area to be captured synchronously when the target image is taken, and ultimately leads to dirt similar to the target being mistakenly identified as the target during target detection; the lighting conditions of the calibration scenes of various after-sales centers do not meet the standards, mainly due to uneven lighting and inappropriate lighting intensity, which causes the target image quality taken during calibration to not meet the standards.
[0035] The inventors of this application further discovered that the fundamental reason is that the vehicle is calibrated in the production workshop before leaving the factory. The calibration scene in the production workshop is built by highly professional engineers. The target area in the calibration scene in the production workshop is relatively clean (only a preset target pattern), and the lighting conditions are more in line with the calibration standards. However, since the professional engineers in the production workshop cannot be stationed in each after-sales center to provide professional guidance, when each after-sales center arranges and maintains the calibration scene on its own in the later stage, the lighting conditions, ground texture, cleanliness of the target area, etc. of the calibration scene may not meet the calibration standards, resulting in low quality of the target images taken during calibration at each after-sales center, which may cause the target detection to fail, and then lead to low calibration accuracy when calibration is performed at each after-sales center.
[0036] In view of this, it is hoped to provide a target image detection method, which can still successfully detect the target from the target image to a certain extent even if the calibration scene of the after-sales center does not meet the standards, resulting in the quality of the captured target image not meeting the standards.
[0037] Based on this, the inventors of this application have proposed a target image detection method in a specific embodiment of this application, comprising: obtaining a target image; obtaining an adaptive threshold corresponding to the target image based on the target image, the adaptive threshold including multiple binarization thresholds; binarizing the target image based on the multiple binarization thresholds to obtain multiple binarized target images; performing target detection on each of the obtained binarized target images to obtain a detection result for each binarized target image; and fusing the detection results of the multiple binarized target images to obtain a target detection result. For the obtained target image, obtaining adaptive thresholds (multiple binarization thresholds) that are adapted to the pixel distribution and / or pixel characteristics of the target image; binarizing the target image based on the multiple binarization thresholds to obtain multiple binarized target images; detecting the target in each binarized target image to obtain a detection result for each binarized target image; and fusing the detection results of the multiple binarized target images to obtain a target detection result. This eliminates the technical problem of low detection sensitivity in detecting targets from a single binary image obtained after binarization of the target image based on a single binary threshold. It can improve the sensitivity of detecting targets from the target image to a certain extent, thereby improving the calibration success rate and accuracy.
[0038] like Figure 2 As shown, the target image detection method provided in the embodiment of the present application can be applied to an automatic driving control system 200, which may include a vehicle 201, a server 202 and an electronic device 203, wherein the vehicle 201 can be communicatively connected to the server 202 through a communication network, and the vehicle 201 can be communicatively connected to the electronic device 203 through a wired or wireless method, that is, the vehicle 201 can send data to the electronic device 203 through a wired or wireless network, and can also receive data sent by the electronic device 203 through a wired or wireless network.
[0039] In some optional embodiments, electronic device 203 may be an in-vehicle computer, a smart phone, a tablet computer, a wearable smart device, etc. In some optional embodiments, server 202 may be a server of a distributed system, a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0040] exist Figure 2 In the application scenario shown, the target image detection method of the embodiment of the present application can be used to calibrate the parameters of the surround-view camera of the vehicle 201.
[0041] In some optional implementations, the target image detection method provided in the embodiments of the present application is independently executed by the vehicle 201, the electronic device 203 or the server 202.
[0042] In some optional implementations, the target image detection method provided in the embodiments of the present application is collaboratively executed by vehicle 201 and server 202, or collaboratively executed by vehicle 201 and electronic device 203, or collaboratively executed by vehicle 201, electronic device 203, and server 202. During collaborative execution, some steps are executed by vehicle 201, and some steps are executed by server 202 and / or electronic device 203.
[0043] Exemplarily, vehicle 201 acquires a target image and sends the acquired target image to server 202 or electronic device 203. After receiving the target image, server 202 or electronic device 203 acquires an adaptive threshold corresponding to the target image based on the target image, wherein the adaptive threshold includes multiple binarization thresholds; binarizes the target image based on the multiple binarization thresholds to obtain multiple binarized target images; performs target detection on each of the obtained binarized target images to obtain a detection result for each binarized target image; and fuses the detection results of the multiple binarized target images to obtain a target detection result.
[0044] It should be noted that, in the manner in which the vehicle 201 and the server 202 collaborate together, or the vehicle 201 and the electronic device 203 collaborate together, or the vehicle 201, the electronic device 203 and the server 202 collaborate together, the steps respectively executed by the vehicle 201, the server 202 and / or the electronic device 203 are not limited to the manner introduced in the above examples. In actual applications, the steps respectively executed by the vehicle 201, the server 202 and / or the electronic device 203 can be dynamically adjusted according to actual conditions.
[0045] like Figure 3 As shown, Figure 3 A flow chart of a target image detection method provided in an embodiment of the present application is shown. The target image detection method is applied to a vehicle, electronic device, and / or server to detect a target from a target image and calibrate the parameters of the vehicle's surround view camera based on the detection results. The target image detection method includes the following steps S301 to S305, which are described in detail as follows:
[0046] S301, acquiring a target image.
[0047] The target image refers to the target image captured by the surround-view camera on the calibrated vehicle when calibrating the vehicle.
[0048] As mentioned above, when calibrating the surround-view camera on a vehicle, the vehicle will be placed on a dedicated calibration station in the workshop. The surround-view camera on the vehicle will shoot the targets arranged around the dedicated calibration station to collect target images.
[0049] For example, Figure 2 The vehicle 201 is shown positioned Figure 1 In the calibration scene shown, the vehicle 201 is at the calibration station 111 and the left front wheel is placed at the position 112. The surround view camera on the vehicle 201 shoots the target area 101 to the target area 110, thereby collecting the target image.
[0050] In some optional embodiments, the acquired target image may be a target image acquired in real time by a surround-view camera on the calibrated vehicle. In some optional embodiments, the acquired target image may be a target image acquired in advance by a surround-view camera on the calibrated vehicle.
[0051] like Figure 4 As shown in FIG, a target image provided by an embodiment of the present application is shown. For example, the surround view camera located in front of the vehicle 201 is Figure 1 The target areas 101 to 103 shown in FIG are photographed and collected to obtain the following Figure 4 The target image shown in the figure. Surround view cameras mostly use fisheye cameras, and the images they capture are distorted. Figure 4 In the target image shown, image area 401 corresponds to Figure 1 The target area 101 in the image area 402 corresponds to Figure 1 The target area 102 in the image area 403 corresponds to Figure 1 The target area 103 in Figure 4 When performing target detection on the target image in the image, the purpose is to identify all targets from image area 401 to image area 403. In the subsequent embodiments of this application, the identification of targets 404 and 405 from image area 402 will be mainly described.
[0052] It should be noted that when calibrating the surround-view cameras on a vehicle, it is necessary to calibrate the surround-view cameras located at the front, rear and both sides of the vehicle body. That is, all surround-view cameras on the vehicle body will collect target images and calibrate based on the collected target images. Figure 4Only the target image captured by the surround-view camera located at the front of the vehicle is exemplified, and the target image detection method provided in the embodiment of the present application is explained based on the target image. However, this does not limit the target detection method provided in the embodiment of the present application to be applicable only to the target image captured by the surround-view camera at the front of the vehicle. It can also be applied to target images captured by the surround-view camera at other positions on the vehicle body (such as at the rear of the vehicle or on both sides of the vehicle body).
[0053] In some optional embodiments, after acquiring a target image, the target image is preprocessed. For example, when the acquired target image is a color image, such as a three-channel RGB image, the color target image is processed into a grayscale image. Of course, other preprocessing methods may also be performed on the target image, such as performing an enhancement operation on the target image, which is not specifically limited in this application.
[0054] S302 : Based on the target image, obtain an adaptive threshold corresponding to the target image, where the adaptive threshold includes a plurality of binarization thresholds.
[0055] Image binarization involves setting the pixel values of pixels in the original image greater than a set threshold to 255 and those less than or equal to the threshold to 0, transforming the original image into a "black and white" image. This creates a clear distinction between foreground and background pixels in the resulting image. The set threshold is the binarization threshold used during image binarization.
[0056] In some calibration operations, the threshold is set to a fixed value. This means that regardless of the lighting conditions, the target image is binarized using a fixed threshold. However, as mentioned above, lighting conditions vary across after-sales centers, and in particular, it's difficult to achieve complete uniformity in the lighting intensity of calibration scenes across different centers. Furthermore, the overall grayscale of target images captured under varying lighting intensities varies. Using a fixed threshold for binarization doesn't account for the pixel distribution and pixel value characteristics within the target image, and the resulting binarized image won't clearly distinguish between foreground and background pixels.
[0057] In some calibration operations, the threshold is determined based on the pixel average of the entire target image. That is, after obtaining the target image, the pixel average of all pixels in the target image is determined based on the pixel values of all pixels in the target image, and then the pixel average is used as the threshold for binarization of the target image. However, please combine Figure 1 The calibration scene shown and Figure 4 As shown in the target image, it is obvious that in the target image, the foreground pixels (i.e., the target) are significantly less than the background pixels. Moreover, it should be noted that in the actual target image, the background pixels are unlikely to be as Figure 4 The target image shown in is not pure white, but rather has a certain grayscale, and the pixel value of each background pixel is also different. Therefore, the pixel average value determined based on the pixel values of all pixels in the target image will approach the pixel average value determined based on the pixel values of all background pixels in the target image. Therefore, when binarizing the target image based on the pixel average value of the entire target image, some background pixels with pixel values lower than the pixel average value of the entire target image may be processed as the same color as the foreground pixels. As a result, the resulting binarized image cannot clearly distinguish between foreground and background pixels.
[0058] In some calibration operations, the threshold is set as an adaptive threshold based on a sliding window. That is, a set window slides on the target image with a set step size. When the window slides to a set position, the pixel average value of all pixels in the window is used as the binarization threshold for all pixels in the window, and all pixels in the window are binarized using this binarization threshold. However, pixels located on the window boundary are not included in the sliding window, resulting in poor connectivity of pixels located on the window boundary in the resulting binary image. Pixels located on the window boundary appear as isolated black or white dots, which can interfere with target recognition.
[0059] In view of this, in the embodiment of the present application, the adaptive threshold corresponding to the target image is a binarization threshold adapted to the region of interest of the target image.
[0060] During image processing, the region of interest (ROI) is delineated within the image using various bounding boxes, such as boxes, circles, ellipses, and irregular polygons. Machine vision software such as Halcon, OpenCV, and Matlab often use various operators and functions to identify the ROI and facilitate further image processing.
[0061] In an embodiment of the present application, the region of interest of the target image is an area where the target may exist, which is outlined from the target image, so that the target can be subsequently detected from these regions of interest, and an adaptive threshold is determined based on the outlined region of interest, and the target is detected from the outlined region of interest.
[0062] For example, Figure 5 As shown, it shows the Figure 4 The target image shown is an image after outlining the region of interest. Figure 5 All areas within the dotted boxes are regions of interest, that is, the outlined regions of interest include regions of interest 406 to 410.
[0063] Determining the adaptive threshold based on the region of interest can not only eliminate the interference of dirt that may exist in the non-region of interest in the target image, but also determine the adaptive threshold by taking into account the pixel distribution and pixel characteristics in the region of interest, so that the binarized image after binarization based on the adaptive threshold can form a clear "black and white" effect between the foreground pixels and the background pixels in the region of interest, which is conducive to detecting the target from the binary image, thereby improving the target detection sensitivity and, to a certain extent, improving the calibration success rate and accuracy. From another perspective, for the entire target image, its pixel distribution and pixel characteristics indicate that the target is concentrated in the region of interest. Determining the adaptive threshold based on the pixel points in the region of interest also takes into account the pixel distribution and pixel characteristics of the entire target image, excludes the non-region of interest, and eliminates the interference of the non-region of interest on the detection of the target in the region of interest, thereby improving the target detection sensitivity and, to a certain extent, improving the calibration success rate and accuracy.
[0064] In some embodiments, the adaptive threshold corresponding to the target image includes a plurality of binarization thresholds adapted to the region of interest of the target image. Exemplarily, a basic binarization threshold is first determined based on the region of interest of the target image, and then the neighborhood of the basic binarization threshold is determined, and then multiple values are taken within the neighborhood to form a binarization threshold group. The target image is binarized multiple times based on the binarization threshold group, and the targets identified from the multiple binarized images are fused, thereby eliminating the technical problem of low detection sensitivity of detecting targets from a single binarized image after binarizing the target image based on a single binarization threshold. The sensitivity of detecting targets from target images can be improved to a certain extent, thereby improving the calibration success rate and accuracy. Specifically, if Figure 6 As shown, a specific flow chart of step S302 in the target detection method provided in the embodiment of the present application is shown. Figure 6 As shown, step S302 further includes the following steps S601 to S603, as follows:
[0065] S601: Determine a basic binarization threshold of the target image.
[0066] In the embodiment of the present application, the basic binarization threshold of the target image is determined based on the target region of interest of the target image. When there is only one region of interest in the target image, the region of interest is the target region of interest. When there are multiple regions of interest in the target image, the target region of interest is the region of interest from which the target is to be detected in any one of the multiple regions of interest. For example, Figure 5If a target is detected in the region of interest 408 in the target image shown, the region of interest 408 is the target region of interest, and a basic binarization threshold is determined based on the pixels in the region of interest 408. Specifically, step S601 includes the following steps a to b:
[0067] Step a: determining a region of interest from the target image.
[0068] Based on the above, the targets in the calibration scene are relatively regular circular targets and square targets. Even if the targets in the target image captured by the surround-view camera are distorted to a certain extent, the square targets and circular targets in the target image are still similar to square targets and circular targets. Moreover, the color and texture of the targets in the target image are relatively distinct and clear relative to the background. That is, the shape features, color features, and texture features of the pixels in the target image are relatively prominent. Therefore, in some optional embodiments, after acquiring the target image, based on the shape features, color features, and / or texture features of the pixels in the target image, the area in the target image where the target may exist is identified, and the area in the target image where the target may exist is outlined with a preset-shaped border, that is, the region of interest is obtained.
[0069] In some optional embodiments, after acquiring the target image, the pixel distribution in the target image is analyzed, and based on the pixel distribution, the area in the target image where the target may exist is determined, and the area in the target image where the target may exist is outlined with a border of a preset shape, that is, the area of interest is obtained.
[0070] For example, Figure 5 All areas within the dotted boxes are regions of interest, that is, the outlined regions of interest include regions of interest 406 to 410.
[0071] Of course, in some other optional implementations, the region of interest of the target image can also be determined by other methods, which are not specifically limited in this application. Figure 1 The target area 101 and Figure 4 By comparing the image area 401 in FIG. 4 , it can be found that Figure 1 Hit the target 114 Figure 4 The target image shown is target 411. Although target 411 is distorted relative to target 114, the degree of distortion of the intersection of the two square targets in target 114 is very small compared to target 411. Therefore, the intersection of the two square targets in target 114 can be used as a feature point. Figure 4 First, the feature point corresponding to target 114 is identified in the target image in the target image, that is, the intersection vertex of the two square targets in target 411, thereby obtaining the target 411 in Figure 4The pixel coordinates of the target image shown are the pixel coordinates of the target 411 in Figure 4 The first region of interest 406 in the target image is shown, and the world coordinates of target 114 are known. Based on the world coordinates of target 114 and the pixel coordinates of target 411, the transformation relationship between the world coordinates and the pixel coordinates is obtained. Based on the transformation relationship and the world coordinates of targets 115 and 116, it can be inferred that Figure 4 The pixel coordinates of target 405 and target 404 in the target image shown are not accurate, but the pixel coordinates of target 405 and target 404 can be accurately predicted with a high probability. Figure 4 The region in the target image shown is the second region of interest 408. Based on the same method, regions of interest 407, 409 and 410 can also be inferred.
[0072] Step b: determining a basic binarization threshold of the target image based on the pixel points within the region of interest.
[0073] In the embodiment of the present application, the maximum inter-class variance method is used to determine the basic binarization threshold of the target image based on the pixels within the region of interest. Specifically, step a includes the following steps a-1 and a-2:
[0074] Step a-1: calculating the inter-class variance corresponding to different grayscale values within the region of interest.
[0075] In some optional implementations, the image within the region of interest is binarized by traversing grayscale values from 0 to 255 as a binarization threshold to obtain 266 binarized images corresponding to the region of interest, and the inter-class variance of each of the 266 binarized images corresponding to the region of interest is calculated using the following formula ①:
[0076] g=w0w1(u0-u1) 2 ①
[0077] Among them, w0 is the proportion of foreground pixels in the binary image corresponding to the region of interest, w1 is the proportion of background pixels in the binary image corresponding to the region of interest, u0 is the average grayscale of the foreground pixels in the binary image corresponding to the region of interest, and u1 is the average grayscale of the background pixels in the binary image corresponding to the region of interest. Of course, when calculating the between-class variance of any of the 266 binary images corresponding to the region of interest, the parameters w0, w1, u0, and u1 in formula ① are determined for the pixels in the same binary image.
[0078] For example, through the above method, Figure 5The region of interest 408 in the target image shown in the figure is obtained, and the inter-class variances g1, g2, g3, ... g corresponding to the 266 binary images of the region of interest are obtained. k …、g 266 Among them, g k The inter-class variance corresponding to the binarized image obtained by binarizing the image in the region of interest when the grayscale value is k-1 as the binarization threshold, 1≤k≤266, and k is an integer.
[0079] Step a-2: Obtain the grayscale value corresponding to the value with the largest inter-class variance, and use the obtained grayscale value as the basic binarization threshold.
[0080] For example, assuming g k It is g1, g2, g3, ...g k …、g 266 The maximum value among them, the gray value k-1 is used as the basic binarization threshold.
[0081] S602: Determine a floating value range of the basic binarization threshold.
[0082] In some optional embodiments, the floating value range is 10 grayscale values floating up and down on the basis of the basic binarization threshold. In some optional embodiments, the floating value range is 10 grayscale values floating up and 5 grayscale values floating down on the basis of the basic binarization threshold. The floating value range can be flexibly determined according to actual needs and is not specifically limited in this application. However, it should be noted that usually the upper and lower limits of the floating value range are both 10 grayscale values to ensure that after the image in the target region of interest is subsequently binarized based on the binarization threshold obtained within the floating value range, the binarized image of the target region of interest can clearly distinguish between foreground pixels and background pixels. If the floating value range is too wide, it may cause the binarization threshold subsequently obtained from the floating value range to be incompatible with the pixel distribution and / or pixel characteristics of the target region of interest.
[0083] Exemplarily, the floating value range of the basic binarization threshold k-1 is the interval [k-11, k+9], which is the neighborhood of the basic binarization threshold k-1.
[0084] S603 , taking multiple values within the floating value range of the basic binarization threshold to obtain multiple binarization thresholds as the adaptive thresholds.
[0085] After obtaining the floating value range of the basic binarization threshold, multiple values can be arbitrarily selected within the floating value range of the basic binarization threshold, that is, multiple binarization thresholds are obtained from the floating value range of the basic binarization threshold to form a threshold group, and the thresholds in the threshold group are used as adaptive thresholds.
[0086] For example, the neighborhood of the basic binarization threshold k-1 takes multiple values to obtain a threshold group {k-8, k-6, k-3, k-1, k+4, k+7}, and the threshold group {k-8, k-6, k-3, k-1, k+4, k+7} is used as Figure 5 Adaptive thresholding of a region of interest 408 in a target image is shown.
[0087] After step S302 is completed, after obtaining a plurality of binarization thresholds adapted to the regions of interest of the acquired target image, step S303 is continued.
[0088] S303 : performing binarization processing on the target image based on the multiple binarization thresholds to obtain multiple binarized target images.
[0089] After obtaining multiple binarization thresholds (i.e., multiple binarization thresholds in a threshold group) that are adapted to the target region of interest of the target image, the target region of interest in the target image is binarized based on the multiple binarization thresholds. At the same time, the pixel values of pixels in other regions of the image except the target region of interest can also be set to 0. That is, after obtaining the target image and the threshold group, the target region of interest of the target image is binarized by traversing all the binarization thresholds in the threshold group to obtain multiple binary target images.
[0090] For example, for Figure 5 In the target image shown, the target region of interest 408 in the target image is binarized using the binarization threshold k-1 in the threshold group {k-8, k-6, k-3, k-1, k+4, k+7}. This involves comparing the pixel values of all pixels in the region of interest 408 with those of k-1, setting the pixel values of pixels with values greater than k-1 to 255, and setting the pixel values of pixels with values less than or equal to k-1 to 0, thereby obtaining a target image binarized based on the binarization threshold k-1. The binarization of the target region of interest 408 in the target image using other binarization thresholds in the threshold group {k-8, k-6, k-3, k-1, k+4, k+7} is performed in the same manner as the binarization of the target region of interest 408 in the target image using the binarization threshold k-1, and will not be further described herein. Until each binary threshold pair in the threshold group {k-8, k-6, k-3, k-1, k+4, k+7} Figure 5 The target region of interest 408 in the target image shown is binarized to obtain the following: Figure 7 The six binary target images are shown.
[0091] S304 , performing target detection on each obtained binary target image to obtain a detection result of each binary target image.
[0092] In some optional embodiments, a connected domain detection is performed on each binary image to detect possible targets within the binary target image. A connected domain is an image region consisting of adjacent pixels with the same pixel value. In the embodiment of the present application, a connected domain in a binary target image is an image region consisting of adjacent pixels with a pixel value of 255 or 0.
[0093] For example, Figure 7 As shown in the first (a) binary target image in the figure, the image area marked by the numbers 701 and 702 is a connected domain composed of two pixel points with a pixel value of 0. These two connected domains are two targets, which are black in the figure. Target 701 corresponds to Figure 5 Target 405 in, target 702 corresponds to Figure 5 Target 406 in; and Figure 7 The connected domain composed of pixels with a pixel value of 255 in the (a)th binary target image is Figure 5 Pixels in the non-target area of the region of interest 408 are shown in white in the figure.
[0094] Similarly, in Figure 7 In the (b)th binarized target image in FIG, targets 703 and 704 are detected. Figure 7 In the (c)th binary target image in FIG, target 705 is detected, while the area indicated by number 706 is connected to other pixels with a pixel value of 0 and fails to form a separate connected domain. That is, the detection based on the connected domain fails to detect the target 705. Figure 7 In the (c)th binary target image, it is detected Figure 5 The target 406 in corresponds to the target. Figure 7 In the (d)th binary target image in the image, target 707 is detected, but target 707 is not detected. Figure 5 The target 406 in corresponds to the target. Figure 7 In the (e)th binary target image, target 709 was detected, but target 709 was not detected. Figure 5 The target 406 in corresponds to the target. Figure 7 In the (f)th binary target image in the image, target 709 was detected, but target 709 was not detected. Figure 5 The target 406 in the image corresponds to the target.
[0095] S305 , fusing the detection results of the multiple binary target images to obtain a target detection result.
[0096] In some optional embodiments, after connected domain detection is performed on multiple binary images of the same target image, cluster analysis is performed on the targets detected from each binary image, and the target detection results that may be the same target in the multiple binary images of the same target image are fused to obtain the target detection result.
[0097] In some optional embodiments, such as Figure 7 As shown, a specific flow chart of step S305 in the target detection method provided in the embodiment of the present application is shown. Figure 7 As shown, step S305 includes the following steps S701 to S703, which are specifically as follows:
[0098] S701 , based on the detection result of each binary target image, obtaining the position coordinates of the target point corresponding to each binary target image.
[0099] The target point corresponding to the binary target image is the target detected from the binary target image during the connected component detection.
[0100] In some optional embodiments, after the target is detected based on the connected domain, the feature points of the detected target are extracted, and the pixel coordinates of the feature points of the target are used as the detection results of the target (pixel coordinates are used to describe the position of the pixel points in the image, and the "position coordinates" mentioned later in this application are all pixel coordinates). Figure 1 The circular target similar to target 115 shown in FIG can use the center of the target as the characteristic point of the target. Figure 7 In the binary target image shown, the position coordinates of the pixel points close to the center of the circle in target 701, target 703, target 705, target 707, target 709 or target 711 can be used as the position coordinates of each target (corresponding to Figure 1 Target 115 and Figure 4 or Figure 5 In other embodiments, for Figure 1 The square target similar to target 114 shown in the figure can use the intersection vertices of the two square targets as the feature points of the target, and in the corresponding binary target image, the position coordinates of the pixel points corresponding to the intersection vertices of the two square targets can be used as the detection results.
[0101] For example, Figure 7The position coordinates of target 701 are (x1, y1), the position coordinates of target 702 are (x2, y2), the position coordinates of target 703 are (x3, y3), the position coordinates of target 704 are (x4, y4), the position coordinates of target 705 are (x5, y5), the position coordinates of target 707 are (x7, y7), the position coordinates of target 709 are (x9, y9), and the position coordinates of target 711 are (x 11 ,y 11 ).
[0102] It should be noted that in the embodiment of the present application, when detecting the position coordinates of the pixel points in the target image, it should be based on a unified pixel coordinate system, that is, whether the pixel coordinates or position coordinates described in the previous or subsequent text are coordinates located in the same pixel coordinate system.
[0103] S702 : Creating a target point array based on the position coordinates of the target points corresponding to each binary target image.
[0104] In some optional embodiments, after connected domain detection is performed on multiple binary images of the same target image, cluster analysis is performed on the position coordinates of the target detected from each binary image, and the position coordinates of the same target in the multiple binary images of the same target image that may be the same target are added to the same target point array, thereby obtaining multiple target point arrays.
[0105] For example, Figure 7 Target 701, Target 703, Target 705, Target 707, Target 709 and Target 711 all correspond to Figure 4 or Figure 5 Target 405( Figure 1 Target 115 in ); Target 702 and Target 704 both correspond to Figure 4 or Figure 5 Target 404 ( Figure 1 116 in the target). A first target array X={(x1, y1), (x3, y3), (x5, y5), (x7, y7), (x9, y9), ( ...1, y1), (x1, y1), (x1, y1), (x1, y1), 11 ,y 11 )}, based on target 702 and target 704, establish a second target array Y = {(x2, y2), (x4, y4)}.
[0106] In some optional embodiments, such as Figure 8 As shown, a specific flow chart of step S702 in the target detection method provided in the embodiment of the present application is shown. Figure 8As shown, step S702 includes the following steps S801 to S804, which are specifically as follows:
[0107] S801, establishing a first target point array based on first position coordinates, where the first position coordinates are position coordinates of target points corresponding to a first binary image, and the first binary image is any one of the multiple binary target images.
[0108] In some optional embodiments, after connected domain detection is performed on multiple binary images of the same target image, one binary image is randomly selected from the multiple binary images as a first binary image, and then a first target array is established using the position coordinates of each target point detected from the connected domain in the first binary image.
[0109] For example, Figure 7 The (a)th binarized target image is used as the first binarized image, and a first target array A1 = {(x1, y1)} is established for target 701, and a first target array A2 = {(x2, y2)} is established for target 702. Figure 7 The (c)th binarized target image is used as the first binarized image, and a first target array A1 ={(x5, y5)} is established for the target 705.
[0110] In order to more conveniently illustrate the technical solutions provided in the embodiments of the present application and facilitate understanding by those skilled in the art, the following Figure 7 The (c)th binarized target image is taken as the first binarized image for illustration, but this is not a limitation of the present application. As shown in step S801, it can be Figure 7 Any one of the six binary images described in is used as the first binary image.
[0111] S802, determining a positional relationship between second position coordinates and the first position coordinates, where the second position coordinates are position coordinates of a target point corresponding to a second binarized image, and the second binarized image is any one of the multiple binarized target images except the first binarized image.
[0112] S803: When the position relationship is that the second position coordinate falls within a preset range of the first position coordinate, the second position coordinate is added to the first target point array.
[0113] S804: When the position relationship is that the second position coordinates do not fall within the preset range of the first position coordinates, establish a second target point array based on the second position coordinates.
[0114] An arbitrary binary image is selected from the plurality of binary images other than the first binary image as the first second binary image, and then a positional relationship between the position coordinates of each target point detected from the connected domain in the first binary image and the position coordinates of each target point detected from the connected domain in the first second binary image is determined, that is, whether the position of each target point detected from the connected domain in the first binary image is adjacent to the position of each target point detected from the connected domain in the first second binary image is determined, that is, whether the position of each target point detected from the connected domain in the first second binary image (i.e., the second position coordinate) falls within a preset range of the position of each target point detected from the connected domain in the first binary image (i.e., the first position coordinate in the first target point array).
[0115] If the position of a target point detected from the connected domain in the first second binary image is close to the position of a target point detected from the connected domain in the first binary image (i.e., within a preset range), the position coordinates of the target points in the first second binary image that are adjacent to each other in the first binary image are placed in the first target point array corresponding to the adjacent target points in the first binary image.
[0116] If the position of a target point detected from the connected domain in the first second binary image is not adjacent to the positions of all target points detected from the connected domain in the first binary image (i.e., not within a preset range), a second target point array is established using the position coordinates of the non-adjacent target point in the first second binary image.
[0117] If the positions of the multiple target points detected from the connected domain in the first second binary image are not adjacent to the positions of all target points detected from the connected domain in the first binary image (i.e., not within a preset range), multiple second target point arrays are established using the position coordinates of the multiple non-adjacent target points in the first second binary image.
[0118] Then, another binary image is selected from the plurality of binary images other than the first binary image and the first second binary image as a second second binary image. Then, a positional relationship between the mean of the position coordinates in the first target point array or the aforementioned second target point array and the position coordinates of each target point detected from the connected domain in the second second binary image is determined, that is, whether each target point detected from the connected domain in the second second binary image is adjacent to the position represented by the mean of the position coordinates in the first target point array or the aforementioned second target point array is determined, that is, whether the position (i.e., the second position coordinate) of each target point detected from the connected domain in the second second binary image falls within a preset range of the mean of the position coordinates in the first target point array or the aforementioned second target array.
[0119] If the position of a target point detected from the connected area in the second second binary image is close to the position represented by the average of the position coordinates in the first target point array or the aforementioned second target array (that is, within a preset range), then the position coordinates of the target points in the second second binary image that are close to the position represented by the average of the position coordinates in the first target point array are placed in the first target point array, and the position coordinates of the target points in the second second binary image that are close to the position represented by the average of the position coordinates in any of the aforementioned second target point arrays are placed in the corresponding second target point array.
[0120] If the position of a target point detected from the connected area in the second second binary image is not adjacent to the position represented by the average of the position coordinates in the first target point array or any of the aforementioned second target arrays (i.e., not within a preset range), then another second target point array is established using the position coordinates of the non-adjacent target point in the second second binary image.
[0121] If the positions of the multiple target points detected from the connected area in the second second binary image are not adjacent to the positions represented by the average of the position coordinates in the first target point array or any of the aforementioned second target arrays (i.e., not within a preset range), then multiple second target point arrays are established using the position coordinates of the multiple non-adjacent target points in the second second binary image.
[0122] Then, a binary image is selected from the plurality of binary images other than the first binary image and the aforementioned first and second second binary images as a third second binary image. Similarly to the aforementioned step, a positional relationship between the position of each target point in the third second binary image and the average of all position coordinates in the aforementioned first target array and the average of all position coordinates in any of the aforementioned second target arrays is determined. Based on the positional relationship, it is determined whether to add the position coordinates of each target point in the aforementioned second binary image to the aforementioned first target array or any of the aforementioned second target arrays, or to create a new second target array using the position coordinates of the target points in the third second binary image.
[0123] Repeat the above steps until all the binarized images except the first binarized image in the multiple binarized images are traversed.
[0124] In some optional embodiments, the preset range is within 5 basic pixels. For example, if the pixel of the first position coordinate is (m, n), then the preset range of the first position coordinate is within a circular range with (m, n) as the center and a radius of 5 basic pixels.
[0125] For example, in Figure 7 The (c)th binarized target image is used as the first binarized image. After the first target array A={(x5, y5)} is established for the target 705, it is necessary to use Figure 7 In the example, one of the other five binary images except the target image (c) is selected as the second binary image. Figure 7 The target image (a) is removed as the first second binary image, and then it is necessary to determine whether the targets 701 and 702 fall within the preset range of the positions represented by the mean of the position coordinates in the first target array A = {(x5, y5)}. Because targets 701, 703, 705, 707, 709 and 711 all correspond to Figure 4 or Figure 5 Target 405( Figure 1 115 in the first target array A), so target 701 falls within the preset range of positions represented by the mean of the position coordinates in the first target array A = {(x5, y5)}, and the position coordinates of target 701 are added to the first target array A to obtain A = {(x1, y1), (x5, y5)}; and target 702 does not fall within the preset range of positions represented by the mean of the position coordinates in the first target array A = {(x5, y5)}, so the second target array B = {(x2, y2)} is established with the position coordinates (x2, y2) of target 702.
[0126] Then, select Figure 7 Except for the (e)th binarized target image as the second second binarized image, it is necessary to determine whether the target 709 falls within the preset range of the position represented by the mean of the position coordinates in the first target array A = {(x1, y1), (x5, y5)} or the preset range of the position represented by the mean of the position coordinates in the second target array B = {(x2, y2)}. Among them, the position represented by the mean of the position coordinates in the first target array A = {(x1, y1), (x5, y5)} is the coordinate Finally, it is determined that target 709 falls within the preset range of the position represented by the mean of the position coordinates in the first target array A = {(x1, y1), (x5, y5)}, and the position coordinates (x9, y9) of target 701 are added to the first target array A to obtain A = {(x1, y1), (x5, y5), (x9, y9)}.
[0127] Then, select Figure 7Except for the (b)th binarized target image as the third second binarized image, it is necessary to determine whether targets 703 and 704 fall within the preset range of the position represented by the mean of the position coordinates in the first target array A = {(x1, y1), (x5, y5), (x9, y9)} or the preset range of the position represented by the mean of the position coordinates in the second target array B = {(x2, y2)}. Among them, the position represented by the mean of the position coordinates in the first target array A = {(x1, y1), (x5, y5)} is the coordinate Finally, it is determined that target 703 falls within the preset range of the position represented by the average value of the position coordinates in the first target array A = {(x1, y1), (x5, y5), (x9, y9)}, and the position coordinates (x3, y3) of target 703 are added to the first target array A to obtain A = {(x1, y1), (x3, y3), (x5, y5), (x9, y9)}. Targets 702 and 704 both correspond to Figure 4 or Figure 5 Target 404 ( Figure 1 116 in the second target array B), therefore, it is determined that target 704 falls within the preset range of the position represented by the mean of the position coordinates in the second target array B = {(x2, y2)}, and the position coordinates (x4, y4) of target 704 are added to the second target array B to obtain B = {(x2, y2), (x4, y4)}.
[0128] Same as above, then select Figure 7 The (b)th binarized target image is removed as the fourth second binarized image, and it is determined that the target 707 falls within the preset range of the position represented by the average of the position coordinates in the first target array A = {(x1, y1), (x3, y3), (x5, y5), (x9, y9)}, and the position coordinate of the target 707 (x7, y7) is added to the first target array A to obtain A = {(x1, y1), (x3, y3), (x5, y5), (x7, y7), (x9, y9)}.
[0129] Same as above, then select Figure 7 The (f)th binary target image is removed as the fifth second binary image, and the target 711 is determined to fall within the preset range of the position represented by the average value of the position coordinates in the first target array A = {(x1, y1), (x3, y3), (x5, y5), (x9, y9)}, (x9, y9)}, and the position coordinate of the target 711 is (x 11 ,y 11 ) is added to the first target array A to obtain A={(x1, y1), (x3, y3), (x5, y5), (x7, y7), (x9, y9), (x 11 ,y 11)}.
[0130] S703, calculate the confidence of each position coordinate in each target point array, obtain the position coordinates of the target points corresponding to each target point array based on the confidence of each position coordinate in each target point array and the position coordinates in each target point array, and use the position coordinates of the target points corresponding to all target point arrays as the target detection results.
[0131] The confidence level P is used to describe the possibility that a target detected by the connected domain detection method is a target. In some optional embodiments, when the detected target is a circular target, the connectivity, roundness, convexity, etc. of the detected target can be calculated, and the confidence level of the detected target can be determined based on the connectivity, roundness and / or convexity. Connectivity indicates the degree of connectivity of the pixels in the detected target, which is expressed by the proportion of connected pixels in the detected target area to all pixels. Roundness indicates the degree to which the detected target is close to the theoretical circle, which is expressed by the difference between the maximum radius and the minimum radius of the target. Convexity indicates the degree of convexity of the periphery of the detected target, which is expressed by the ratio of the area of the maximum inscribed convex polygon to the area of the minimum circumscribed convex polygon of the detected target.
[0132] The position coordinates describe the position of the detected target in the target image in which the target is detected, and the confidence of the position coordinates is the confidence of the target described by the position coordinates.
[0133] For example, the first target array A is calculated as follows: {(x1, y1), (x3, y3), (x5, y5), (x7, y7), (x9, y9), (x 11 ,y 11 The confidence level of each position coordinate in )} is P1, P3, P5, P7, P9, P 11 The confidence level of each position coordinate in the second target array B = {(x2, y2), (x4, y4)} is P2 and P4 respectively.
[0134] Based on the foregoing, each target point array essentially corresponds to a target point. In some embodiments, for a target point array, the position coordinates of the target point corresponding to the target point array are obtained based on the confidence of each position coordinate in the target point array and all position coordinates in the target point array.
[0135] For example, based on all position coordinates {(x1, y1), (x3, y3), (x5, y5), (x7, y7), (x9, y9), (x 11 ,y 11 )} and the confidence levels P1, P3, P5, P7, P9, P 11, determine the first target array A corresponding to Figure 4 or Figure 5 Based on all the position coordinates {(x2, y2), (x4, y4)} in the second target array B and the confidence P2, P4 corresponding to each position coordinate, determine the second target array B corresponding to Figure 4 or Figure 5 The position coordinates of target 404 in .
[0136] In some optional embodiments, the step of "obtaining the position coordinates of the target point corresponding to each target point array based on the confidence of each position coordinate in each target point array and the position coordinates in each target point array" in the target detection method provided in the embodiment of the present application includes the following steps c to d, as follows:
[0137] Step c: determining the weight of each position coordinate in the target point array according to the confidence of all position coordinates in the target point array and the confidence of each position coordinate in the target point array;
[0138] In some optional implementations, for a target point array, the confidence of a position coordinate in the target point array is divided by the sum of the confidences of all position coordinates in the target point array to obtain the weight of the position coordinate.
[0139] For example, based on all position coordinates {(x1, y1), (x3, y3), (x5, y5), (x7, y7), (x9, y9), (x 11 ,y 11 )} and the confidence levels P1, P3, P5, P7, P9, P 11 , determine the weight of the coordinate (x1, y1) The weight of the coordinate (x3, y3) is The weight of the coordinate (x5, y5) is The weight of the coordinate (x7, y7) is The weight of the coordinate (x9, y9) is Coordinate (x 11 ,y 11 ) has a weight of Based on all the position coordinates {(x2, y2), (x4, y4)} in the second target array B and the confidence P2 and P4 corresponding to each position coordinate, the weight of the coordinate (x2, y2) is determined to be The weight of the coordinate (x4, y4) is
[0140] Step d: performing weighted averaging on the position coordinates in the target point array according to the weight of each position coordinate in the target point array to obtain the position coordinates of the target point corresponding to the target point array.
[0141] For example, the position coordinates of the target 405 corresponding to the first target array A are (s1x1+s3x3+s5x5+s7x7+s9x9+s 11 x 11 ,s1y1+s3y3+s5y5+s7y7+s9y9+s 11 y 11 ). The position coordinates of the target 404 corresponding to the first target array B are (s2x2+s4x4, s2y2+s4y4)
[0142] In an embodiment of the present application, a plurality of binarization thresholds adapted to the acquired target image are obtained, the target image is binarized based on the plurality of binarization thresholds to obtain a plurality of binarized target images, the target in each binarized target image is detected, a detection result for each binarized target image is obtained, and the detection results of the plurality of binarized target images are fused to obtain a target detection result. Based on the region of interest, an adaptive binarization threshold adapted to the pixel distribution and pixel characteristics within the region of interest is determined, which can, to a certain extent, eliminate interference from other dirty regions in the target image and can adapt to the lighting conditions corresponding to the image within the region of interest of the target image. At the same time, within the field of adaptive binarization thresholding, multiple binarization thresholds are taken to form a threshold group. The target image is binarized multiple times based on the multiple binarization thresholds in the threshold group to obtain multiple binarized images. The target detection results of the multiple binarized images are then fused. This overcomes the technical problem that the image obtained by binarization processing based on a single binarization threshold cannot accurately reflect the difference between foreground pixels and background pixels, resulting in the inability to accurately identify the target from the image. This can improve the sensitivity of detecting targets from the target image to a certain extent, thereby improving the calibration success rate and accuracy.
[0143] like Figure 10 As shown, Figure 10 A target image detection device provided by an embodiment of the present application is shown. The target image detection device 100 may include: a first acquisition module 101, a second acquisition module 102, a first processing module 103, a first detection module 104, and a first fusion module 105. Specifically:
[0144] A first acquisition module 101 is used to acquire a target image;
[0145] A second acquisition module 102 is configured to acquire an adaptive threshold corresponding to the target image based on the target image, where the adaptive threshold includes a plurality of binarization thresholds;
[0146] A first processing module 103 is configured to perform binarization processing on the target image based on the multiple binarization thresholds to obtain multiple binarized target images;
[0147] A first detection module 104 is configured to perform target detection on each obtained binary target image to obtain a detection result for each binary target image;
[0148] The first fusion module 105 is configured to fuse the detection results of the multiple binary target images to obtain a target detection result.
[0149] In some optional implementations, the second acquisition module 102 includes:
[0150] A first determination submodule is used to determine a basic binarization threshold of the target image;
[0151] A second determining submodule is used to determine the floating value range of the basic binarization threshold;
[0152] The first value-taking submodule is configured to take multiple values within the floating value range of the basic binarization threshold to obtain multiple binarization thresholds as the adaptive threshold.
[0153] In some optional implementations, the first determining submodule includes:
[0154] A first determining unit, configured to determine a region of interest from the target image;
[0155] The second determining unit is configured to determine a basic binarization threshold of the target image based on the pixels within the region of interest.
[0156] In some optional implementations, the second determining unit includes:
[0157] A first calculation subunit is used to calculate the inter-class variance corresponding to different grayscale values in the region of interest;
[0158] The first acquisition subunit is used to acquire the grayscale value corresponding to the value with the largest inter-class variance, and use the acquired grayscale value as the basic binarization threshold.
[0159] In some optional implementations, the first fusion module 105 includes:
[0160] A first acquiring unit, configured to acquire the position coordinates of a target point corresponding to each binary target image based on a detection result of each binary target image;
[0161] A first establishing unit is configured to establish a target point array based on the position coordinates of the target points corresponding to each binary target image;
[0162] The second calculation unit is used to calculate the confidence of each position coordinate in each target point array, obtain the position coordinates of the target points corresponding to each target point array based on the confidence of each position coordinate in each target point array and the position coordinates in each target point array, and use the position coordinates of the target points corresponding to all target point arrays as the target detection results.
[0163] In some optional implementations, the first establishing unit includes:
[0164] A first establishing subunit is configured to establish a first target point array based on first position coordinates, where the first position coordinates are position coordinates of target points corresponding to a first binarized image, and the first binarized image is any one of the multiple binarized target images;
[0165] a first determining subunit, configured to determine a positional relationship between second position coordinates and the first position coordinates, wherein the second position coordinates are position coordinates of a target point corresponding to a second binarized image, and the second binarized image is any one of the multiple binarized target images except the first binarized image;
[0166] a first adding subunit, configured to add the second position coordinates to the first target point array when the position relationship is such that the second position coordinates fall within a preset range of the first position coordinates;
[0167] The second establishing subunit is configured to establish a second target point array based on the second position coordinates when the position relationship is that the second position coordinates do not fall within a preset range of the first position coordinates.
[0168] In some optional implementations, the second computing unit includes:
[0169] a second determining subunit, configured to determine a weight of each position coordinate in the target point array according to the confidence levels of all position coordinates in the target point array and the confidence level of each position coordinate in the target point array;
[0170] The second calculation subunit is used to perform weighted averaging on the position coordinates in the target point array according to the weight of each position coordinate in the target point array to obtain the position coordinates of the target point corresponding to the target point array.
[0171] like Figure 11 As shown, Figure 11The following is a block diagram of an electronic device provided in an embodiment of the present application. The vehicle 110 may be the vehicle 110 capable of running programs described in the aforementioned embodiments. The vehicle 110 in the present application may include one or more of the following components: a memory 112 for storing one or more computer programs; and one or more processors 111 for accessing and executing the one or more computer programs from the memory 112, so that the electronic device 110 executes the vehicle control method described in the aforementioned embodiments.
[0172] The processor 111 may include one or more processing cores. The processor 111 utilizes various interfaces and circuits to connect various components within the vehicle 110. It executes instructions, programs, code sets, or instruction sets stored in the memory 112, as well as accesses data stored in the memory 112, to perform various functions and process data for the vehicle 110. Optionally, the processor 111 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 111 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 111 and may be implemented separately via a communications chip.
[0173] The memory 112 may include a random access memory (RAM) or a read-only memory (ROM). The memory 112 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 112 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, a target image playback function, a shooting function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the terminal during use (such as a phone book, audio and video data, map data, driving record data), etc.
[0174] like Figure 12 As shown, Figure 12The following is a block diagram of a vehicle structure provided by an embodiment of the present application. The vehicle 120 may be the vehicle 120 capable of running programs described in the aforementioned embodiments. The vehicle 120 in the present application may include one or more of the following components: a memory 122 for storing one or more computer programs; and one or more processors 121 for accessing and executing the one or more computer programs from the memory 122, thereby enabling the vehicle 120 to execute the vehicle control method described in the aforementioned embodiments.
[0175] The processor 121 may include one or more processing cores. The processor 121 utilizes various interfaces and circuits to connect various components within the vehicle 120. It executes instructions, programs, code sets, or instruction sets stored in the memory 122, as well as accesses data stored in the memory 122, to perform various functions and process data within the vehicle 120. Optionally, the processor 121 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 121 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 121 and may instead be implemented via a separate communications chip.
[0176] The memory 122 may include a random access memory (RAM) or a read-only memory (ROM). The memory 122 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 122 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, a shooting function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the terminal during use (such as a phone book, audio and video data, map data, driving record data, etc.).
[0177] like Figure 13 As shown, Figure 13The computer-readable storage medium 130 stores program codes, which can be called by a processor to execute the method described in the above method embodiment.
[0178] The computer-readable storage medium 130 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has storage space for program codes for executing any of the method steps described above. These program codes can be read from or written to one or more computer program products. The program codes can be compressed, for example, in an appropriate form.
[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A target image detection method, characterized in that: The method comprises: Acquire a target image; determining a basic binarization threshold of the target image based on the target image, wherein the basic binarization threshold is used to indicate an inter-class variance of a region of interest in the target image; Determining a floating value range of the basic binarization threshold; Taking multiple values within the floating value range of the basic binarization threshold to obtain multiple binarization thresholds; Based on the multiple binarization thresholds, binarize the target image to obtain multiple binarized target images; Performing target detection on each obtained binary target image to obtain a detection result of each binary target image, wherein the detection result is used to indicate a target point in each binary target image; Based on the detection result of each binary target image, the position coordinates of the target point corresponding to each binary target image are obtained; Establishing a target point array based on the position coordinates of the target point corresponding to each binary target image, wherein each target point array corresponds to one target point; Calculate the confidence of each position coordinate in each of the target point arrays, determine the weight of each position coordinate in the target point array based on the confidence of all position coordinates in the target point array and the confidence of each position coordinate in the target point array, perform weighted averaging on the position coordinates in the target point array based on the weight of each position coordinate in the target point array to obtain the position coordinates of the target point corresponding to the target point array, and use the position coordinates of the target points corresponding to all target point arrays as the detection result of the target image.
2. The method according to claim 1, characterized in that Determining a basic binarization threshold of the target image includes: determining a region of interest from the target image; A basic binarization threshold of the target image is determined based on the pixels within the region of interest.
3. The method according to claim 1, characterized in that The step of establishing a target point array based on the position coordinates of the target points corresponding to each binary target image comprises: Establishing a first target point array based on first position coordinates, where the first position coordinates are position coordinates of a target point corresponding to a first binary image, and the first binary image is any one of the multiple binary target images; Determining a positional relationship between second position coordinates and the first position coordinates, where the second position coordinates are position coordinates of a target point corresponding to a second binarized image, and the second binarized image is any one of the multiple binarized target images except the first binarized image; When the position relationship is that the second position coordinate falls within the preset range of the first position coordinate, adding the second position coordinate to the first target point array; When the position relationship is that the second position coordinate does not fall within the preset range of the first position coordinate, a second target point array is established based on the second position coordinate.
4. A target image detection device, characterized in that: The device comprises: A first acquisition module is used to acquire a target image; A first determining submodule is configured to determine a basic binarization threshold of the target image, wherein the basic binarization threshold is used to indicate an inter-class variance of a region of interest in the target image; A second determining submodule is used to determine the floating value range of the basic binarization threshold; A first value-taking submodule is configured to take multiple values within a floating value range of the basic binarization threshold to obtain multiple binarization thresholds; a first processing module, configured to perform binarization processing on the target image based on the multiple binarization thresholds to obtain multiple binarized target images; a first detection module, configured to perform target detection on each obtained binary target image to obtain a detection result of each binary target image, wherein the detection result is used to indicate a target point in each binary target image; The first fusion module is used to obtain the position coordinates of the target point corresponding to each binary target image based on the detection result of each binary target image; establish a target point array based on the position coordinates of the target point corresponding to each binary target image, wherein each target point array corresponds to one target point; calculate the confidence of each position coordinate in each target point array, determine the weight of each position coordinate in the target point array according to the confidence of all position coordinates in the target point array and the confidence of each position coordinate in the target point array, perform weighted averaging on the position coordinates in the target point array according to the weight of each position coordinate in the target point array, obtain the position coordinates of the target point corresponding to the target point array, and use the position coordinates of the target points corresponding to all target point arrays as the detection result of the target image.
5. An electronic device, characterized in that: include: a memory for storing one or more computer programs; One or more processors, configured to call and run the one or more computer programs from the memory to perform the method according to any one of claims 1 to 3.
6. A vehicle, characterized in that: include: a memory for storing one or more computer programs; One or more processors, configured to call and run the one or more computer programs from the memory to perform the method according to any one of claims 1 to 3.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Rapid calibration method and system based on binocular vision detection technology
CN108765495A
Dynamic target identification and tracking method for autonomous landing of unmanned aerial vehicle
CN110222612A