Target detection method and device
By processing machine vision image signal of the RAW data of the image sensor, images suitable for neural network models are generated, which solves the problem of poor target detection effect in extremely low illumination scenarios in the prior art, and achieves efficient target detection under low latency, low power consumption and low computing power.
Patent Information
- Application Number
- CN202311603693.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
The existing object detection technology is difficult to achieve effective object detection in low latency, low power consumption and low computing power scenarios, especially in extremely low illumination scenarios. When using RAW images to directly perform object detection, the pixel average of the image is concentrated near 0 values, resulting in failure of the quantization process and it is difficult to obtain reliable detection effects.
By performing machine vision image signal processing on the original image RAW data from the image sensor, including level correction, image compression, color space conversion and resolution conversion, images with machine vision characteristics are generated, adapted to the detection task requirements of the neural network model, and adjusted the processing parameters according to the output results of the neural network model to achieve object detection.
It reduces the processing cost and detection delay of image signal processing, improves the accuracy and reliability of object detection, and meets the needs of low latency, low power consumption, and low computing power object detection.
Smart Images

Figure CN120047680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular, to a target detection method. Background Art
[0002] With the rapid development of deep neural network technology, breakthrough progress has also been made in the field of target detection. The purpose of target detection is to locate the position coordinates of the target in the image and identify the category of the target. Current target detection technologies are widely applied to various scenarios in daily life, such as autonomous driving systems (ADS), surveillance, robotics, and healthcare. More and more target detection technologies are deployed on edge computing devices with limited resources and power consumption, such as mobile phones, surveillance cameras, etc., which requires target detection methods to achieve low latency, low power consumption, and low computing power.
[0003] Currently, most target detections are trained based on a labeled image sample dataset that conforms to a certain color language protocol, for example, the sRGB (standard Red Green Blue) standard. Correspondingly, during inference, it is also necessary to be based on the same color language protocol as the training samples to ensure satisfactory results.
[0004] However, images that conform to the color language protocol are usually obtained by subjecting the original image (RAW image) obtained by the camera sensor to a series of complex image signal processing (ISP), which makes it difficult to achieve the low latency, low power consumption, and low computing power required for target detection. Summary of the Invention
[0005] The present invention provides a target detection method to meet at least one of the requirements of low latency, low power consumption, and low computing power for target detection.
[0006] The first aspect of the embodiments of the present application provides a target detection method, which includes:
[0007] Performing machine vision image signal processing on the RAW data of the original image from the image sensor to form an image with machine vision characteristics, and obtaining first image data.
[0008] Inputting the first image data into a trained neural network model for target detection to perform target detection.
[0009] Wherein,
[0010] The parameters used in the machine vision image signal processing are determined according to the target detection metrics of the output results of the trained neural network model.
[0011] The image with machine vision characteristics matches the detection task of the neural network model.
[0012] Preferably, the machine vision image signal processing enables the RAW data to be processed into a target image in an extremely low illumination scene to be detected by the neural network model, and / or the RAW data distribution of the original image is adjusted to a machine vision characteristic image that matches the quantization requirements of the neural network model.
[0013] Preferably, the machine vision image signal processing includes:
[0014] Level correction processing for correcting the output level of each pixel point in the RAW data of the original image,
[0015] Image compression processing for transforming the corrected output level of each pixel point into a brightness value,
[0016] Color space conversion processing for converting the brightness value of each pixel point into color components,
[0017] Resolution conversion processing for converting the first resolution of the image into a second resolution, where the first resolution is greater than the second resolution.
[0018] Preferably, the level correction processing includes black level correction processing, and this processing includes:
[0019] For each pixel point in the RAW data of the original image,
[0020] Taking the result obtained by multiplying the difference between the output level of this pixel point and the white level value by the proportion of the white level value in the difference between the black and white level values as the black level correction result of this pixel point;
[0021] Preferably, the image compression processing includes gamma compression processing, and this processing includes:
[0022] For each pixel point, calculating the gamma power of the black level correction result of this pixel point to obtain the brightness value of this pixel point;
[0023] Preferably, the color space conversion processing includes filtering processing, and this processing includes: performing RAW2y operation on the image data after gamma compression processing;
[0024] Preferably, the resolution conversion processing includes downsampling processing, and this processing includes: performing downsampling on the image data after RAW2y operation.
[0025] Preferably, the machine vision image signal processing further includes: selecting at least one of bad pixel correction, denoising, highlight suppression, backlight compensation, color adjustment, and lens shadow correction according to the detection task of the neural network model,
[0026] Among them,
[0027] Bad pixel correction is used to detect and correct bad pixels in an image.
[0028] Denosing is used to remove noise in an image to improve the image quality.
[0029] Highlight suppression is used to suppress highlight areas in an image to balance the image brightness.
[0030] Backlight compensation is used to compensate for backlight areas in an image to improve the image contrast.
[0031] Color adjustment is used to adjust the color representation of an image to improve the accuracy of image segmentation.
[0032] Lens shadow correction is used to correct image shadows caused by lens angles and lighting conditions.
[0033] Preferably, the determination of the object detection metrics according to the output results of the trained neural network model includes: adjusting parameter by parameter for each parameter, or adjusting the parameters used in each processing step step by step with the processing steps as the granularity, or adjusting all the parameters used as a whole with the overall machine vision image signal processing as the granularity.
[0034] Preferably, the neural network model is trained in the following manner:
[0035] Set the initial values of the parameters used in the machine vision image signal processing.
[0036] Perform machine vision image signal processing on the RAW sample data using the initial values to obtain second sample image data.
[0037] Input the second sample image data into the neural network model to be trained.
[0038] Adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model to be trained and the expected detection result.
[0039] Return to execute the step of performing machine vision image signal processing on the RAW sample data using the initial values until the output result meets the expectation, and obtain the trained neural network model.
[0040] Preferably, the neural network model is trained in the following manner: During the first round of training, set the initial values of the parameters used in the machine vision image signal processing.
[0041] Perform machine vision image signal processing on the RAW sample data using the initial values to obtain second sample image data.
[0042] Input the second sample image data into the neural network model to be trained.
[0043] Adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model to be trained and the expected detection result.
[0044] Return to execute the step of performing machine vision image signal processing on the RAW sample data using the initial value until the output result reaches the expectation, and obtain the neural network model after the first round of training.
[0045] In the next round of training after the first round, determine the parameters used for machine vision image signal processing according to the target detection index of the output result of the neural network model after the first round of training.
[0046] Perform machine vision image signal processing on the RAW sample data using the determined parameters to obtain the second sample image data.
[0047] Input the second sample image data into the current neural network model.
[0048] Adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model and the expected detection result.
[0049] Return to execute the step of performing machine vision image signal processing on the RAW sample data using the determined parameters until the output result reaches the expectation, and obtain the neural network model after this round of training.
[0050] Determine the parameters used for machine vision image signal processing according to the target detection index of the output result of the neural network model after this round of training, and return to execute the step of performing machine vision image signal processing on the RAW sample data using the determined parameters to perform the next round of training until the training ends.
[0051] The second aspect of the embodiments of the present application provides a target detection device, and the device includes:
[0052] A machine vision image signal processing module, configured to perform machine vision image signal processing on the original image RAW data from an image sensor to form an image with machine vision characteristics, and obtain the first image data.
[0053] A target detection module, configured to input the first image data into a trained neural network model for target detection to perform target detection.
[0054] The object detection method provided by this application performs machine vision image signal processing on the original RAW data from an image sensor to form an image with machine vision characteristics, enabling object detection to be based on RAW data, reducing the processing cost and detection latency of image signal processing. By determining the parameters used in machine vision image signal processing according to the object detection metrics of the output results of the trained neural network model, it is beneficial to reduce the computing power required for object detection and improve the accuracy and reliability of detection. Description of the Drawings
[0055] Figure 1 It is a schematic flowchart of an object detection method according to an embodiment of this application.
[0056] Figure 2 It is a schematic flowchart of constructing image signal processing for machine vision according to an embodiment of this application.
[0057] Figure 3 It is a schematic diagram of obtaining RAW sample data in this embodiment.
[0058] Figure 4 It is a schematic flowchart showing that the training of the neural network model and the adjustment of the parameters used in CV-ISP in this embodiment can be carried out alternately.
[0059] Figure 5 It is a schematic diagram of an object detection device according to an embodiment of this application.
[0060] Figure 6 It is a schematic diagram of an image monitoring device applying the object detection device of this application.
[0061] Figure 7 It is another schematic diagram of an object detection device according to an embodiment of this application. Detailed Description of the Embodiment
[0062] In order to make the purpose, technical means, and advantages of this application clearer and more understandable, the following further elaborates on this application with reference to the accompanying drawings.
[0063] The applicant's research found that there are the following problems in object detection based on the image processed by ISP:
[0064] Firstly, due to the complexity of the complete ISP processing, there may be certain redundancy for the object detection task. For example, ISP processing generally includes but is not limited to the following:
[0065] 1. Dead pixel correction: Detect and correct dead pixels in the image.
[0066] 2. Denoising: Remove noise in the image to improve image quality.
[0067] 3. Highlight suppression: Suppress the highlight areas in the image to balance the image brightness.
[0068] 4. Backlight compensation: Compensate the backlight areas in the image to improve the image contrast.
[0069] 5. Color enhancement: Enhance the color performance of the image to improve the visual effect of the image.
[0070] 6. Lens shadow correction: Correct the image shadow problems caused by factors such as lens angle and lighting conditions.
[0071] These processing contents are the core contents of ISP processing and vary with different ISP algorithms and devices.
[0072] There may be some ISP processing contents that are not required for object detection;
[0073] Second, ISP includes several processing links, and the parameters used in each processing link are debugged based on the human eye senses. These parameters may not be optimal for object detection tasks.
[0074] If directly using the RAW image for object detection to completely remove the ISP processing flow, although the overall complexity is reduced, generally, the object detection models deployed on edge computing devices need to perform quantization processing on the image data input to the model. When detecting extremely low-light objects, the pixel mean of the image is generally small and concentrated near 0. After directly quantizing such image data, the loss of the image itself will make it difficult to obtain reliable results.
[0075] The object detection method provided in the embodiments of this application selectively retains non-redundant ISP processing and removes redundant ISP processing according to the ISP processing links that affect the object detection task effect to perform computer vision image signal processing (CV-ISP), and adjusts the parameters used in the selected CV-ISP processing according to the object detection metrics.
[0076] See Figure 1 as shown Figure 1 is a schematic flowchart of an object detection method according to an embodiment of this application. The method includes:
[0077] Step 11, perform computer vision image signal processing on the raw image RAW data from the image sensor to form an image with computer vision characteristics, and obtain the first image data,
[0078] wherein, the parameters used in the computer vision image signal processing are determined according to the object detection metrics of the output result of the trained neural network model.
[0079] The machine vision characteristic image matches the detection task of the neural network model. That is, according to the image characteristics required for the detection task of the neural network model, the original RAW image data from the image sensor is processed into a machine vision characteristic image that conforms to the required image characteristics.
[0080] As an example, machine vision image signal processing enables the RAW data to be processed into a target image in an extremely low illumination scene for detection by the neural network model, and / or the distribution of the original RAW image data is adjusted to match the quantization requirements of the neural network model to form a machine vision characteristic image.
[0081] As an example, machine vision image signal processing includes:
[0082] Level correction processing for correcting the output level of each pixel point in the original RAW image data. For example, black level correction.
[0083] Image compression processing for transforming the corrected output level of each pixel point into a brightness value. For example, gamma compression.
[0084] Color space conversion processing for converting the brightness value of each pixel point into color components. For example, RAW2y operation.
[0085] Resolution conversion processing for converting the first resolution of the image into a second resolution, where the first resolution is greater than the second resolution. For example, downsampling.
[0086] When facing an extremely low illumination scene where the RAW data distribution is concentrated near 0, the above-mentioned level correction processing and image compression processing can adjust the data distribution of the RAW data to make the data distribution more adaptable to the quantization requirements. In this way, the effect of the target detection system is improved. Through color space conversion processing, not only can the noise be reduced to a certain extent and the effect of target detection in a high-noise scene be improved, but also the RAW data can be converted into a grayscale image and fed into the network as a single-channel input, which saves more calculations compared to the traditional three-channel input of RGB.
[0087] As another example, machine vision image signal processing further includes: selecting at least one of bad pixel correction, denoising, strong light suppression, backlight compensation, color adjustment, and lens shadow correction according to the detection task of the neural network model.
[0088] Among them,
[0089] Bad pixel correction is used to detect and correct bad pixels in the image.
[0090] Denoising is used to remove noise in the image to improve the image quality.
[0091] Highlight suppression is used to suppress the highlight areas in an image to balance the image brightness.
[0092] Backlight compensation is used to compensate the backlight areas in an image to improve the image contrast.
[0093] Color adjustment is used to adjust the color representation of an image to improve the accuracy of image segmentation.
[0094] Lens shadow correction is used to correct the image shadows caused by the lens angle and lighting conditions.
[0095] Step 12: Input the first image data into the trained neural network model for object detection to perform object detection.
[0096] As an example, the neural network model can be a model of any structure, including but not limited to convolutional neural network (CNN), recurrent neural network (RNN), temporal convolutional network (TCN), etc.
[0097] The object detection method of the embodiments of the present application can achieve object detection directly based on the original RAW image data. Compared with the object detection based on the image signal processing of visual perception, the embodiments of the present application reduce a large number of redundant image signal processing modules, reduce the storage and computing costs, and thus reduce the overall power consumption. Moreover, since the machine vision image signal processing selectively performs corresponding image processing according to the image characteristics required by the detection task of the neural network model, and different machine vision image signal processing is selected with different detection tasks, in this way, both the full storage and full computing of all machine vision image signal processing are avoided, and the machine vision image signal processing is only used during object detection, which further helps to reduce the storage and computing costs. For the parameters used in the CV-ISP processing, the object detection metrics can be adjusted as the goal, avoiding the influence of the parameter selection for satisfying the human eye visual effect on the object detection effect and improving the reliability of the object detection system.
[0098] For ease of understanding the present application, the following takes the object detection in an extremely low illumination scenario as an example for illustration. It should be understood that the present application is not limited to this scenario, and the object detection in any other scenario can be adjusted and applied as needed.
[0099] See Figure 2 as shown Figure 2 is a schematic flowchart of constructing the machine vision image signal processing according to the embodiments of the present application. It includes:
[0100] Step 21: Perform machine vision image signal processing on the RAW sample data from the image sensor to form a machine vision sample image.
[0101] As an example, machine vision image signal processing, i.e., CV-ISP, sequentially includes: level correction processing, image compression processing, color space conversion, and resolution conversion processing.
[0102] Since the accuracy of the analog-to-digital conversion chip in the actual image sensor cannot accurately convert too small voltage values, and there will be dark current in the circuit of the image sensor itself. These dark currents cause certain output levels for pixel points that are not exposed to light, which is equivalent to all pixel points outputting a minimum level. This results in the effective representation range of the RAW sample data not being fully utilized. For example, the 8-bit representation range is [0, 255], but the actual RAW sample data only uses [25, 255]. Therefore, level correction is required. Given that the minimum level is perceived as black by the human eye, it is called black level correction (BLC, BlackLevel Correct).
[0103] As an example, for each pixel point in the RAW sample data, the result obtained by multiplying the difference between the output level of the pixel point and the black level value by the proportion of the white level value in the difference between the white and black level values is used as the black level correction result of the pixel point; expressed in a mathematical formula as:
[0104] y = (x – black_level) × white_level / (white_level – black_level),
[0105] where x is the output level value of each pixel point before correction, y is the output level value of the pixel point after correction, black_level is the black level value, that is, the minimum level value output by the image sensor, and white_level is the white level value, that is, the maximum level value output by the image sensor.
[0106] Given that the output of the sensor often has a large number of bits, to further reduce the overall bit width, image data bit width compression can be performed. For example, gamma compression. Taking gamma compression as an example, calculate the gamma power of the output level of each pixel point after correction to obtain the brightness value of the pixel point. Expressed in a mathematical formula as:
[0107] B = y gamma
[0108] where y is the output level of each pixel point after correction, B is the brightness value of the pixel point, and gamma can be determined according to the target detection index.
[0109] Through gamma compression processing, both the data bit width is compressed and the output level of the pixel point is converted into a brightness value.
[0110] As an example, the color space conversion can be achieved through a RAW2y operation. The filtering process can be performed by sliding convolution of the raw sample data with a fixed filter.
[0111] The RAW2y operation converts the RAW sample data into the Y component in the YCbCr color space. In digital image processing, YCbCr is a commonly used color space, where Y represents the luminance component, and Cb and Cr represent the chrominance components. By converting to the YCbCr color space, it is more convenient to adjust attributes such as the brightness, contrast, and color balance of the image.
[0112] Specifically, it can be performed by a fixed filter, such as sliding convolution on the raw sample data in the Bayer pattern. For example, for the raw sample data in the Bayer pattern, since the Bayer pattern is arranged as or The result of the 3×3 filter is to perform weighted averaging of the pixel values of the R, G, and B channels within a 3×3 window to reduce noise. In addition, the filtering process can also convert the RAW sample data into a grayscale image to be input into the target detection neural network as single-channel data, which saves computing power compared to the three-channel input of RGB.
[0113] To further reduce the computational and memory overhead of target detection, the original resolution (the first resolution), such as images with resolutions of 1280*720, 1920*1080, etc., is downsampled to convert the first resolution of the image into a second resolution, where the first resolution is greater than the second resolution, so as to perform target detection based on the low-resolution image data.
[0114] See Figure 3 as shown in Figure 3 This is a schematic diagram for obtaining RAW sample data in this embodiment. The image sample data that conforms to the color language protocol, such as sREG, which has been labeled, is successively subjected to inverse global tone mapping, inverse gamma correction, inverse color correction, inverse white balance, and inverse Bayer arrangement to obtain the RAW sample data.
[0115] Among them,
[0116] The inverse global tone mapping is used to map the luminous intensity of the display device into the signal intensity output by the camera sensor.
[0117] The inverse gamma correction is used to convert the optical signal into an electrical signal.
[0118] The inverse color correction is used to simulate the inverse process of color correction.
[0119] The inverse white balance processing is used to simulate the inverse process of white balance to obtain the original image without white balance processing.
[0120] The inverse Bayer arrangement processing is used to convert the image conforming to the color language protocol into the original image in Bayer arrangement.
[0121] In this step, any parameter value set can be used as the initial value for the parameters used in the machine vision image signal processing.
[0122] Step 22: Input the first sample image data after CV-ISP processing into the trained neural network model for object detection to perform sample object detection.
[0123] Step 23: Adjust the parameters used in the CV-ISP processing according to the sample object detection metrics output by the trained neural network model.
[0124] As an example, use the automatic machine learning method (Auto-ML) to adjust the parameters used in the CV-ISP processing to improve the object detection effect oriented by the object detection metrics.
[0125] For example, the object detection metrics can include mean Average Precision (mAP), the Auto-ML parameter search method is random search, assuming the number of searches is 10 times. Taking the adjustment of the parameter Gamma value as an example, the search space is [0.1, 3.0], and the search interval is 0.2. The adjustment process is as follows:
[0126] Step 1: First, randomly search for K parameters in the search space at the search interval, for example, 4 parameters.
[0127] Step 2: Use each parameter obtained by random search for Gamma compression to obtain 4 different detection metrics output by the object detection model.
[0128] Step 3: Repeat Step 1 and Step 2 until the preset number of searches is reached, and record the Gamma parameter used when the detection metric is the highest as the final Gamma parameter.
[0129] As an example, the adjustment of the parameters used in the CV-ISP processing can be to adjust each parameter one by one, or to adjust the parameters used in each processing link granularly by processing link, or to adjust all the parameters used in the overall CV-ISP granularly as a whole.
[0130] For example, the way to adjust each parameter used in the CV-ISP processing one by one can be as follows:
[0131] Step 231: For any parameter used, use the automated machine learning method to search within the search parameter space of the parameter at the search parameter interval of the parameter.
[0132] Step 232: Use the searched parameter to perform machine vision image signal processing on the RAW sample data to obtain first sample image data.
[0133] Step 233: Input the first sample image data into the trained neural network model and record the target detection metrics of the output result of the neural network model.
[0134] Step 234: Return to Step 231 until the search for the parameter is completed, and use the parameter used for machine vision image signal processing at the best target detection metrics as the parameter used for machine vision image signal processing during target detection.
[0135] Step 235: Repeatedly execute Steps 231 - 234 until each parameter used has been adjusted.
[0136] For another example, the parameters used in the CV - ISP processing process can be adjusted as a whole in the following way:
[0137] Step 231': For each parameter used, use the automated machine learning method to search within the search parameter space of the parameter at the search parameter interval of the parameter.
[0138] Step 232': Use the searched parameters to perform machine vision image signal processing on the RAW sample data to obtain first sample image data.
[0139] Step 233': Input the first sample image data into the trained neural network model and record the target detection metrics of the output result of the neural network model.
[0140] Step 234': Return to execute Step 231' until the search is completed, and use the parameters used for machine vision image signal processing at the best target detection metrics as the parameters used for machine vision image signal processing during target detection.
[0141] For yet another example, the parameters used in the CV - ISP processing process can be adjusted for each processing link in the following way:
[0142] Step 231'': For each parameter used in any processing link, use the automated machine learning method to search within the search parameter space of the parameter at the search parameter interval of the parameter.
[0143] Step 232〃, in this processing step, perform machine vision image signal processing on the RAW sample data using the searched parameters to obtain the first sample image data.
[0144] Step 233〃, input the first sample image data into the trained neural network model, and record the target detection metrics of the output result of this neural network model.
[0145] Step 234〃, return to execute Step 231〃 until the search for each parameter used in this processing step is completed. Take the parameter used in this processing step at the time of the best target detection metrics as the parameter used in this processing step in the machine vision image signal processing during target detection.
[0146] Step 235〃, repeatedly execute Steps 231〃 to 234〃 until the parameters used in all processing steps are adjusted.
[0147] As an example, the search end condition can be that the number of searches reaches a set number threshold, or that the target detection metrics reach the expected value, etc.
[0148] In view of the fact that the neural network model can have a certain generalization ability after training and can adapt to the parameters used in different CV-ISPs, the above-mentioned trained neural network model can be obtained through training in the following manner:
[0149] A1. Set the initial values of the parameters used in the machine vision image signal processing.
[0150] A2. Perform machine vision image signal processing on the RAW sample data using the initial values to obtain the second sample image data.
[0151] A3. Input the second sample image data into the neural network model to be trained.
[0152] A4. According to the loss function value between the output result of the neural network model to be trained and the expected detection result, adjust the model parameters of the neural network model.
[0153] Repeatedly execute Steps A2 to A4 until the output result reaches the expectation to obtain the trained neural network model.
[0154] To improve the training efficiency and the accuracy of target detection, the training of the above-mentioned neural network model and the adjustment of the parameters used in CV-ISP can be carried out alternately. See Figure 4 as shown. Figure 4 This is a schematic flowchart showing that the training of the neural network model in this embodiment and the adjustment of the parameters used in CV-ISP can be carried out alternately. It includes:
[0155] Step 41, during the first-round training process, set the initial values of the parameters used in machine vision image signal processing.
[0156] Step 42, perform machine vision image signal processing on the RAW sample data using the initial values to obtain the second sample image data.
[0157] Step 43, input the second sample image data into the neural network model to be trained.
[0158] Step 44, adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model to be trained and the expected detection result.
[0159] Step 45, return to execute Step 42 until the output result meets the expectation, and obtain the neural network model after the first-round training.
[0160] Step 46, during the next-round training process after the first round, determine the parameters used in machine vision image signal processing according to the target detection index of the output result of the neural network model after the first-round training.
[0161] Step 47, perform machine vision image signal processing on the RAW sample data using the determined parameters to obtain the second sample image data.
[0162] Step 48, input the second sample image data into the current neural network model.
[0163] Step 49, adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model and the expected detection result.
[0164] Step 50, return to execute Step 47 until the output result meets the expectation, and obtain the neural network model after this round of training.
[0165] Step 51, determine the parameters used in machine vision image signal processing according to the target detection index of the output result of the neural network model after this round of training.
[0166] In this step, the parameters used in machine vision image signal processing can be determined according to Step 23.
[0167] Step 52, return to execute Step 47 to perform the next-round training until the training ends.
[0168] In this embodiment, the target detection index of the output result of the neural network model is used to adjust the parameters used in the machine vision image signal processing, so that the image data obtained by the machine vision image signal processing is more matched and adapted to the target detection task performed by the neural network model, avoiding the mismatch between the image data processed by the traditional ISP and the target detection task performed by the neural network model due to being suitable for human eye perception, which is beneficial to improving the accuracy of target detection.
[0169] See Figure 5 as shown Figure 5 This is a schematic diagram of the target detection device according to an embodiment of the present application. The device includes:
[0170] A machine vision image signal processing module, configured to perform machine vision image signal processing on the original RAW image data from an image sensor to form an image with machine vision characteristics, and obtain first image data.
[0171] A target detection module, configured to input the first image data into a trained neural network model for target detection to perform target detection.
[0172] As an example, the machine vision image signal processing module includes:
[0173] A level correction processing sub-module, configured to correct the output level of each pixel point in the original RAW image data.
[0174] An image compression processing sub-module, configured to transform the corrected output level of each pixel point into a brightness value.
[0175] A color space conversion processing sub-module, configured to convert the brightness value of each pixel point into color components.
[0176] A resolution conversion processing sub-module, configured to convert the first resolution of the image into a second resolution, where the first resolution is greater than the second resolution.
[0177] The machine vision image signal processing module further includes:
[0178] A selection sub-module, configured to select at least one of bad pixel correction, denoising, strong light suppression, backlight compensation, color adjustment, and lens shadow correction for processing according to the detection task of the neural network model.
[0179] A bad pixel correction processing sub-module, configured to detect and correct bad pixels in the image.
[0180] A denoising sub-module, configured to remove noise in the image.
[0181] A strong light suppression sub-module, configured to suppress strong light regions in the image to balance the brightness of the image.
[0182] A backlight compensation sub-module, which is used to compensate for the backlight area in the image to improve the contrast of the image.
[0183] A color adjustment sub-module, which is used to adjust the color performance of the image to improve the accuracy of image segmentation.
[0184] A lens shadow correction sub-module, which is used to correct the image shadow caused by the lens angle and lighting conditions.
[0185] See Figure 6 as shown Figure 6 is a schematic diagram of an image monitoring device to which the target detection device of the present application is applied. The image monitoring device includes:
[0186] An image sensor, which is used to collect image data to obtain RAW data.
[0187] An ISP processing module, which is used to perform image signal processing on the RAW data from the image sensor to present a visual image suitable for human eye perception to the user.
[0188] A machine vision image signal processing module, which is used to perform machine vision image signal processing on the original image RAW data from the image sensor to form a machine vision image with machine vision characteristics, so as to form a machine vision image that matches the task of the neural network model for target detection.
[0189] A target detection module, which is used to input the machine vision image data into a trained neural network model for target detection.
[0190] A wake-up module, which is used to trigger system wake-up according to the comparison result between the detection result output by the target detection module and the set wake-up event. When the detection result meets the set wake-up event, the system is woken up. When the detection result does not meet the set wake-up event, the system wake-up is prohibited, so as to keep the system in the standby mode when the non-wake-up event occurs.
[0191] The image monitoring device of this embodiment processes RAW data into machine vision data for target detection, which is beneficial to reducing power consumption. Thus, it can be realized to trigger system wake-up by using the set wake-up event. For example, when a pedestrian or a motor vehicle appears in front of the camera, event wake-up can be realized under extremely low system power consumption.
[0192] See Figure 7 as shown Figure 7 is another schematic diagram of the target detection device according to the embodiment of the present application. The device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of the target detection method described in the embodiment of the present application, and / or the steps of constructing the machine vision image signal processing described in the embodiment of the present application.
[0193] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0194] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0195] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the target detection step and / or the step of constructing machine vision image signal processing.
[0196] For the device / network-side device / storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0197] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0198] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A target detection method, characterized in that, the method includes: Performing machine vision image signal processing on the original image RAW data from an image sensor to form a machine vision characteristic image, and obtaining first image data, Inputting the first image data into a trained neural network model for target detection to perform target detection, wherein, The parameters used in the machine vision image signal processing are determined according to the target detection metrics of the output results of the trained neural network model, The machine vision characteristic image matches the detection task of the neural network model.
2. The method according to claim 1, characterized in that, The machine vision image signal processing enables the RAW data to be processed into a target image in an extremely low illumination scene to be detected by the neural network model, and / or the distribution of the original image RAW data is adjusted to match the quantization requirements of the neural network model to form a machine vision characteristic image.
3. The method according to claim 2, characterized in that, The machine vision image signal processing includes: Level correction processing, for correcting the output level of each pixel point in the original image RAW data, Image compression processing, for transforming the corrected output level of each pixel point into a brightness value, Color space conversion processing, for converting the brightness value of each pixel point into color components, Resolution conversion processing, for converting the first resolution of the image into a second resolution, where the first resolution is greater than the second resolution.
4. The method according to claim 3, characterized in that, The level correction processing includes black level correction processing, and this processing includes: For each pixel point in the original image RAW data, Taking the result obtained by multiplying the difference between the output level of this pixel point and the white level value by the proportion of the white level value in the difference between the black and white level values as the black level correction result of this pixel point; The image compression processing includes gamma compression processing, and this processing includes: For each pixel point, calculating the gamma power of the black level correction result of this pixel point to obtain the brightness value of this pixel point; The color space conversion processing includes filtering processing, and this processing includes: performing a RAW2y operation on the image data after gamma compression processing; The resolution conversion processing includes downsampling processing, and this processing includes: performing downsampling on the image data after the RAW2y operation.
5. The method according to claim 4, characterized in that, The machine vision image signal processing further includes: selecting at least one of bad pixel correction, denoising, strong light suppression, backlight compensation, color adjustment, and lens shadow correction according to the detection task of the neural network model, wherein, Bad pixel correction is used to detect and correct bad pixels in the image, Denoising is used to remove noise in the image to improve image quality, Strong light suppression is used to suppress strong light areas in the image to balance the brightness of the image, Backlight compensation is used to compensate for backlight areas in the image to improve image contrast, Color adjustment is used to adjust the color performance of the image to improve the accuracy of image segmentation, Lens shadow correction is used to correct image shadows caused by lens angles and lighting conditions.
6. The method according to claim 5, characterized in that the determination of the object detection index according to the output result of the trained neural network model includes: adjusting parameter by parameter for each parameter, or adjusting the parameters used in each processing link by processing link granularity, or adjusting all the parameters used as a whole with the whole machine vision image signal processing as the granularity, wherein the adjusting all the parameters used as a whole with the whole machine vision image signal processing as the granularity includes: for each parameter used, using an automated machine learning method to search in the search parameter space of the parameter at the search parameter interval of the parameter, performing machine vision image signal processing on the RAW sample data using the searched parameters to obtain first sample image data, inputting the first sample image data into the trained neural network model, and recording the object detection index of the output result of the neural network model, returning to execute the step of using an automated machine learning method to search in the search parameter space of the parameter at the search parameter interval of the parameter for each parameter used to perform the next search until the search ends, using the parameters used at the best object detection index as the parameters used for machine vision image signal processing; the adjusting parameter by parameter for each parameter includes: for any parameter used, using an automated machine learning method to search in the search parameter space of the parameter at the search parameter interval of the parameter, performing machine vision image signal processing on the RAW sample data using the searched parameter to obtain first sample image data, inputting the first sample image data into the trained neural network model, and recording the object detection index of the output result of the neural network model, returning to execute the step of using an automated machine learning method to search in the search parameter space of the parameter at the search parameter interval of the parameter for any parameter used to search for the next parameter until the search ends, using the parameter used at the best object detection index as the parameter used for the machine vision image signal processing; the adjusting the parameters used in each processing link by processing link granularity includes: for each parameter used in any processing link, using an automated machine learning method to search in the search parameter space of the parameter at the search parameter interval of the parameter, performing machine vision image signal processing on the RAW sample data using the searched parameters in this processing link to obtain first sample image data, inputting the first sample image data into the trained neural network model, and recording the object detection index of the output result of the neural network model, returning to execute the step of using an automated machine learning method to search in the search parameter space of the parameter at the search parameter interval of the parameter for each parameter used in any processing link to search for the parameters used in the next processing link until the search ends, using the parameters used in this processing link at the best object detection index as the parameters used in this processing link for the machine vision image signal processing.
7. The method according to claim 6, wherein, the neural network model is trained in the following manner: Set the initial values of the parameters used for machine vision image signal processing, Perform machine vision image signal processing on the RAW sample data using the initial values to obtain second sample image data, Input the second sample image data into the neural network model to be trained, Adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model to be trained and the expected detection result, Return to execute the step of performing machine vision image signal processing on the RAW sample data using the initial values until the output result meets the expectation, and obtain the trained neural network model.
8. The method according to claim 6, wherein, the neural network model is trained in the following manner: During the first round of training, set the initial values of the parameters used for machine vision image signal processing, Perform machine vision image signal processing on the RAW sample data using the initial values to obtain second sample image data, Input the second sample image data into the neural network model to be trained, Adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model to be trained and the expected detection result, Return to execute the step of performing machine vision image signal processing on the RAW sample data using the initial values until the output result meets the expectation, and obtain the neural network model after the first round of training, During the next round of training after the first round, determine the parameters used for machine vision image signal processing according to the target detection index of the output result of the neural network model after the first round of training, Perform machine vision image signal processing on the RAW sample data using the determined parameters to obtain second sample image data, Input the second sample image data into the current neural network model, Adjust the model parameters of the neural network model according to the loss function value between the output result of the neural network model and the expected detection result, Return to execute the step of performing machine vision image signal processing on the RAW sample data using the determined parameters until the output result meets the expectation, and obtain the neural network model after this round of training, Determine the parameters used for machine vision image signal processing according to the target detection index of the output result of the neural network model after this round of training, and return to execute the step of performing machine vision image signal processing on the RAW sample data using the determined parameters to perform the next round of training until the training ends.
9. A target detection device, wherein, the device includes: A machine vision image signal processing module, configured to perform machine vision image signal processing on the original image RAW data from an image sensor to form an image with machine vision characteristics, and obtain first image data, A target detection module, configured to input the first image data into a trained neural network model for target detection.
10. The device according to claim 9, wherein, the machine vision image signal processing module includes: A level correction processing sub-module, configured to correct the output level of each pixel point in the original image RAW data, The image compression processing sub-module is used to transform the output level after correcting each pixel into a brightness value. The color space conversion processing sub-module is used to convert the brightness value of each pixel into color components. The resolution conversion processing sub-module is used to convert the first resolution of the image into a second resolution, where the first resolution is greater than the second resolution.