RAW image processing method for target detection
By building a semantic fusion ISP network, we can directly generate images that are more suitable for detection from RAW images, solving the problem of insufficient detection accuracy of JPEG images in object detection, and achieving efficient image compression and detection rate improvement.
Patent Information
- Application Number
- CN202510329394.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the prior art, JPEG images have overemphasized the visual optimization of human eyes in the object detection task, resulting in the distinction between the foreground and background that is not clear enough, limiting the detection accuracy.
Build an ISP network with semantic fusion to directly generate images that are more suitable for detection and occupancy of storage from RAW images. Through the combination of feature extraction network and ISP network, the key features of RAW images are retained and enhanced, and efficient loss functions are used for training.
The recognition accuracy of object detection is significantly improved and the storage space requirement is reduced, the number of bits required for image transmission is reduced, and the detection rate is improved.
Smart Images

Figure CN120374933A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of image processing and object detection, and specifically provides a RAW image processing method for object detection. Background Art
[0002] With the continuous improvement of the social automation level, the application of image information is no longer limited to human vision, and the image communication between machines is becoming increasingly frequent, with its importance increasing significantly. As one of the core tasks in the field of computer vision, object detection aims to accurately identify and locate specific target objects in images or videos; currently, most image object detection tasks are implemented using JPEG images, and JPEG images are usually generated after a series of processes by an Image Signal Processor (ISP) on the original RAW images. RAW images are the digital images directly generated by the image sensors of digital cameras and contain a large amount of detailed information. In addition, in order to provide a better visual experience for the human eye, major camera manufacturers optimize the parameter values of each step of the image signal processor according to empirical values. However, this optimization method for the human eye visual experience conflicts with the requirements of the object detection task. In the object detection task, an image is divided into a foreground and a background. The foreground is the region of interest (ROI) where the target is located, and the background is the background information irrelevant to the task; in order to improve the detection accuracy and efficiency, the features of the foreground target need to be enhanced, and the background information should be suppressed. However, the optimization method based on the human eye visual experience enhances the foreground and the background without discrimination, overemphasizing the visual effect and details, resulting in unclear distinction between the foreground and the background. When the optimized JPEG image is applied to object detection, the detection accuracy is greatly limited. Therefore, the present invention provides a RAW image processing method for object detection, which retains the key features of the RAW image as much as possible and enhances them emphatically, so that the processed image is still applicable to the object detection task. Summary of the Invention
[0003] The purpose of the present invention is to provide a RAW image processing method for object detection, which can improve the compression efficiency of the image and significantly improve the recognition accuracy when the processed image is applied to the object detection task.
[0004] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0005] A RAW image processing method for object detection, characterized by comprising the following steps:
[0006] Step 1. Collect the RAW original image generated by the image sensor of the digital camera, and generate a JPEG input image through an Image Signal Processor (ISP); meanwhile, convert the RAW original image according to the single-channel RGGB segmentation to form a RAW input image in 4-channel format;
[0007] Step 2. Use a feature extraction network to extract features from the JPEG input image to obtain multi-scale features;
[0008] Step 3. Construct an ISP network with semantic fusion and complete the training. Input the RAW input image and the multi-scale features into the ISP network, and the ISP network outputs a 3-channel RGB image.
[0009] Further, in Step 2, the feature extraction network uses the pre-trained RESNET101 network.
[0010] Further, in Step 3, the input of the ISP network with semantic fusion is the RAW input image and the multi-scale features Feature1 to Feature3, and the output is a 3-channel RGB image; the ISP network with semantic fusion includes: network units U1 to U6, a downsampling module, an upsampling module, an SFT module, a residual network, and an output layer, where:
[0011] The structure of network unit U1 is: CONV 3×3×3×32 + LeakyReLU + CONV 3×3×32×32 + LeakyReLU, its input U1in is the RAW input image, and the output is the feature U1out;
[0012] After U1out is downsampled by the downsampling module, it is input into the SFT module and fused with the feature Feature1, and the SFT module outputs the feature U1sft;
[0013] The structure of network unit U2 is: CONV 3×3×32×64 + LeakyReLU + CONV 3×3×64×64 + LeakyReLU, its input U2in is the feature U1sft, and the output is the feature U2out;
[0014] After the feature U2out is downsampled by the downsampling module, it is input into the SFT module and fused with the feature Feature2, and the SFT module outputs the feature U2sft;
[0015] The structure of network unit U3 is: CONV3×3×64×128 + LeakyReLU + CONV 3×3×128×128 + LeakyReLU, its input U3in is the feature U2sft, and the output is the feature U3out;
[0016] Feature U3out is downsampled by the downsampling module and then input into the SFT module, where it is feature-fused with Feature 3, and the SFT module outputs Feature U3sft;
[0017] Feature U3sft passes through the residual network and then is input into the SFT module, where it is feature-fused with Feature 3, and the SFT module outputs Feature U4sft;
[0018] Feature U4sft is upsampled by the upsampling module and then concatenated with Feature U3out in channels (concat), and after concatenation, it is input into network unit U4;
[0019] The structure of network unit U4 is: CONV3×3×(256 + 128)×256+LeakyReLU+CONV 3×3×256×128+LeakyReLU, and its output is Feature U4out;
[0020] Feature U4out is input into the SFT module, where it is feature-fused with Feature 2, and the SFT module outputs Feature U5sft;
[0021] Feature U5sft is upsampled by the upsampling module and then concatenated with Feature U2out in channels (concat), and after concatenation, it is input into network unit U5;
[0022] The structure of network unit U5 is: CONV 3×3×(128 + 64)×128+LeakyReLU+CONV 3×3×128×64+Lea kyReLU, and its output is Feature U5out;
[0023] Feature U5out is input into the SFT module, where it is feature-fused with Feature 1, and the SFT module outputs Feature U6sft;
[0024] Feature U6sft is upsampled by the upsampling module and then concatenated with Feature U1out in channels (concat), and after concatenation, it is input into network unit U6;
[0025] The structure of network unit U6 is: CONV 3×3×(64 + 32)×64+LeakyReLU+CONV 3×3×64×32+Leaky ReLU, and its output is Feature U6out;
[0026] The structure of the output layer is: CONV 1×1×32×3+Sigmoid, its input is U6out, and the output is Out, which is the 3-channel RGB image output by the feature fusion network.
[0027] Further, in the ISP network, the downsampling module is composed of dilated convolutional layers, and its structure is: Dil-CONV 3×3×(in_channel)×(in_channel), with a stride of 2, a padding of 2, and a dilation of 2.
[0028] Further, in the ISP network, the structure of the upsampling module is: CONV 3×3×(in_channel)×(in_channel×4)+PixelShuffle(2); a convolution is used to change the number of output channels to 4 times that of the input channels, and after convolution, PixelShuffle(2) is used to complete upsampling.
[0029] Further, in step 3, the training process of the ISP network is: setting a loss function and performing end-to-end offline training on the ISP network; the specific loss function is:
[0030] L = Perceptual Loss +Task Loss +Pixel Loss +Bpp
[0031] where L is the total loss function, Perceptual_Loss is the perceptual loss, Task Loss is the task loss, Pixel_Loss is the pixel loss, and Bpp is the compression degree characterization value.
[0032] Based on the above technical solutions, the beneficial effects of the present invention are as follows:
[0033] The present invention proposes a RAW image processing method for object detection. By constructing an ISP network with semantic fusion, it directly generates an image that is more suitable for detection and occupies less storage from the RAW image; in the present invention, the RAW image retains a large amount of original information, and semantic information is obtained through the JPEG image and its provided labels and integrated into the RAW image processed by the ISP. While taking into account the detection accuracy, the number of bits required for image transmission is effectively reduced; at the same time, an efficient loss function close to the training target is used to constrain the entire training process, which can significantly reduce the space occupied by the image and improve the detection rate; in summary, after introducing the RAW image processing method for object detection, the present invention can improve the recognition accuracy and reduce the storage space requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic flowchart of the RAW image processing method for object detection in the embodiment of the present invention.
[0035] Figure 2Schematic diagram of the ISP network with semantic fusion in the embodiments of the present invention.
[0036] Figure 3 Comparison chart of bpp-mAP between the embodiments of the present invention and the comparative examples. Detailed implementation manners
[0037] To make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0038] This embodiment provides a RAW image processing method for object detection, and its process is as Figure 1 shown, and specifically includes the following steps:
[0039] Step 1. Read the RAW image and complete data preprocessing;
[0040] Collect the RAW original image generated by the image sensor of the digital camera, and generate a JPEG input image through the image signal processor (ISP); meanwhile, convert the RAW original image according to the single-channel RGGB segmentation to form a RAW input image in 4-channel format;
[0041] Step 2. Use the feature extraction network to extract features from the JPEG input image to obtain multi-scale features Feature1 to Feature3, and the feature extraction network uses the pre-trained RESNET101 network;
[0042] Step 3. Build an ISP network with semantic fusion and complete training, input the RAW input image and multi-scale features into the ISP network, and output a 3-channel RGB image by the ISP network;
[0043] The ISP network is as Figure 2 shown, its input is the RAW input image and multi-scale features Feature1 to Feature3, and the output is a 3-channel RGB image; specifically, the ISP network with semantic fusion adopts the Unet framework, including: network units U1 to U6, downsampling module, upsampling module, SFT module, residual network and output layer, where:
[0044] The structure of network unit U1 is: CONV 3×3×3×32 + LeakyReLU + CONV 3×3×32×32 + LeakyReLU, its input U1in is the RAW input image, and the output is feature U1out;
[0045] After U1out is downsampled by the downsampling module and then input into the SFT module, it is feature-fused with feature Feature1, and the SFT module outputs feature U1sft;
[0046] The downsampling module is composed of dilated convolutional layers, with the structure: Dil-CONV 3×3×(in_channel)×(in_channel), with a stride of 2, padding of 2, and dilation of 2;
[0047] The structure of network unit U2 is: CONV 3×3×32×64+LeakyReLU+CONV 3×3×64×64+LeakyReLU, with its input U2in being feature U1sft and its output being feature U2out;
[0048] Feature U2out is downsampled by the downsampling module and then input into the SFT module, where it is feature-fused with feature Feature2, and the SFT module outputs feature U2sft;
[0049] The structure of network unit U3 is: CONV3×3×64×128+LeakyReLU+CONV 3×3×128×128+LeakyReLU, with its input U3in being feature U2sft and its output being feature U3out;
[0050] Feature U3out is downsampled by the downsampling module and then input into the SFT module, where it is feature-fused with feature Feature3, and the SFT module outputs feature U3sft;
[0051] Feature U3sft is input into the SFT module after passing through the residual network, where it is feature-fused with feature Feature3, and the SFT module outputs feature U4sft;
[0052] Feature U4sft is upsampled by the upsampling module and then concatenated (concat) with feature U3out in channels, and the concatenated result is input into network unit U4;
[0053] The structure of the upsampling module is: CONV 3×3×(in_channel)×(in_channel×4)+PixelShuffle(2); uses a convolutional layer to change the number of output channels to 4 times that of the input channels, and after convolution, PixelShuffle(2) is used to complete upsampling;
[0054] The structure of network unit U4 is: CONV3×3×(256+128)×256+LeakyReLU+CONV 3×3×256×128+LeakyReLU, and its output is feature U4out;
[0055] Feature U4out is input into the SFT module, where it is feature-fused with feature Feature2, and the SFT module outputs feature U5sft;
[0056] After the feature U5sft is upsampled by the upsampling module, it is concatenated with the feature U2out in the channel dimension (concat), and the concatenated result is input into the network unit U5;
[0057] The structure of the network unit U5 is: CONV 3×3×(128+64)×128+LeakyReLU+CONV 3×3×128×64+LeakyReLU, and its output is the feature U5out;
[0058] The feature U5out is input into the SFT module, where it is feature-fused with the feature Feature1, and the SFT module outputs the feature U6sft;
[0059] After the feature U6sft is upsampled by the upsampling module, it is concatenated with the feature U1out in the channel dimension (concat), and the concatenated result is input into the network unit U6;
[0060] The structure of the network unit U6 is: CONV 3×3×(64+32)×64+LeakyReLU+CONV 3×3×64×32+LeakyReLU, and its output is the feature U6out;
[0061] The structure of the output layer is: CONV 1×1×32×3+Sigmoid, its input is U6out, and its output is Out, which is the 3-channel RGB image output by the feature fusion network;
[0062] It should be noted that in the present invention, "CONV A×B×C×D" represents a convolutional layer with a convolutional kernel size of A×B, an input channel number of C, and an output channel number of D, "Dil-CONV A×B×C×D" represents a dilated convolutional layer with a convolutional kernel size of A×B, an input channel number of C, and an output channel number of D, LeakyReLU represents the LeakyReLU (Leaky Rectified Linear Unit) activation function, Sigmoid represents the Sigmoid activation function, and PixelShuffle(2) represents 2-fold upsampling;
[0063] The training process of the ISP network is: setting the loss function and performing end-to-end offline training on the ISP network;
[0064] In this embodiment, the PASCAL RAW dataset is used to provide JPEG images of 600×400 and RAW images of 6034×4012. The RAW images of 6034×4012 are cropped to remove the redundant data on the right side and the lower side, and downsampled to 600×400 and then stored in the.npy format; multi-scale features extracted from the JPEG images and the RAW images are used as inputs, and the COCO annotations provided officially are used as labels to form the training dataset;
[0065] The specific loss function is as follows:
[0066] L = Perceptual Loss +Task Loss +Pixel Loss +Bpp
[0067] Among them, L is the total loss function, Perceptual_Loss is the perceptual loss, Task Loss is the task loss, Pixel_Loss is the pixel loss (L2 Loss), and Bpp is the compression degree representation value;
[0068]
[0069] Among them, x n and y n are the values of the input image x and the target image y at pixel n, and N is the number of pixel points;
[0070]
[0071] Among them, φ i (x) and φ i (y) are the feature representations of the input image x and the target image y at the i-th layer, λ i is the weight of each layer of features, and ||·||2 is the L2 norm;
[0072] Task Loss = L cls +L reg +L preg +L obj
[0073] Among them, L cls is the classification loss, which is used to evaluate the performance of the model in the classification task and is obtained by calculating the difference between the class probabilities predicted by the model and the true labels, and the cross-entropy loss function is used;
[0074] L regis the regression loss, which is used to evaluate the gap between the continuous values predicted by the model (such as the coordinates of the bounding box) and the true values. The regression loss function used in this embodiment is the mean squared error (MSE);
[0075] L preg is the pre-regression loss, which is used to adjust the intermediate features to prevent the model from overfitting. The L2 norm is used to regularize the features, and the expression is:
[0076]
[0077] where ||·||2 is the L2 norm, λ is the regularization coefficient, and f is the feature extracted by the network;
[0078] L obj is the objectness loss, which is used to evaluate whether the predicted region of the model contains an object in the object detection task, and is used to calculate the probability that the model predicts an object in each candidate region. The binary cross-entropy loss function is used;
[0079] In addition, it should be noted that: The present invention provides a RAW image processing method for object detection. When the image processed by the present invention is applied to different detection tasks, the task loss can be adaptively modified, which can further improve the detection rate.
[0080] Based on the above technical solutions, this embodiment is tested on the dataset PASCALRAW, which contains 4259 pictures (each picture contains its corresponding.NEF file and.JPG file, with a resolution of 6016×4000); there are 3 categories in this dataset, and 300 pictures of each category are randomly selected, a total of 900 pictures are used as the test set. In this embodiment, the pictures are downsampled to 600×400, and 4 qualities (quality = 2, 3, 4, 5) are selected for comparison. During the test process, the RGB images output by the ISP network are sequentially passed through the compression network and the object detection network to complete the object detection task; To illustrate the beneficial effects of the present invention, Comparative Example 1 is set: Use the software Nikon Workshop provided by Nikon for ISP processing and use Cheng-2020 for end-to-end compression; Comparative Example 2: Use the demosaicing network for processing and use Cheng-2020 for compression; The test results of this embodiment and Comparative Examples 1 and 2 are as Figure 3 shown, where the horizontal axis is bpp (bits per pixel), which represents the number of bits required to transmit each pixel of the image, and the vertical axis is mAP (mean Average Precision), which is used to measure the overall accuracy of the object detection task; The four points of each method represent the images of 4 qualities respectively. The four points of each method correspond to 4 images of different qualities, and the quality increases sequentially from left to right; ByFigure 3 It can be seen that the method proposed by the present invention has a significantly higher detection accuracy than the comparative example under the same or similar bpp; especially when the image quality is high, the present invention can still obtain a higher mAP at a lower bit rate, which proves that the proposed method has better overall performance and practical value in the application of combining image compression and object detection.
[0081] In summary, the RAW image processing method for object detection proposed by the present invention has excellent performance. In the PASCAL RAW dataset, compared with Nikon Workshop, under the same image quality, the mAP can be increased by more than 10%.
[0082] As described above, the above are only specific embodiments of the present invention. Any feature disclosed in this specification, unless specifically stated, can be replaced by other equivalent or similar-purpose alternative features; all the features disclosed, or all the steps in any method or process, except for mutually exclusive features and / or steps, can be combined in any manner.
Claims
1. A RAW image processing method for object detection, characterized in that, It includes the following steps: Step 1. Collect the RAW original image generated by the image sensor of the digital camera, and generate a JPEG input image through an Image Signal Processor (ISP); meanwhile, convert the RAW original image according to single-channel RGGB segmentation to form a RAW input image in 4-channel format; Step 2. Use a feature extraction network to extract features from the JPEG input image to obtain multi-scale features; Step 3. Construct an ISP network with semantic fusion and complete the training. Input the RAW input image and the multi-scale features into the ISP network, and the ISP network outputs a 3-channel RGB image.
2. The RAW image processing method for object detection according to claim 1, wherein In Step 2, the feature extraction network uses the pre-trained RESNET101 network.
3. The RAW image processing method for target detection according to claim 1, characterized in that, In Step 3, the input of the ISP network with semantic fusion is the RAW input image and multi-scale features Feature1 to Feature3, and the feature map with fused semantics; the ISP network with semantic fusion includes: network units U1 to U6, a downsampling module, an upsampling module, an SFT module, a residual network, and an output layer, where: The structure of network unit U1 is: CONV 3×3×3×32 + LeakyReLU + CONV 3×3×32×32 + LeakyReLU, its input U1in is the RAW input image, and the output is feature U1out; After U1out is downsampled by the downsampling module and input into the SFT module, it performs feature fusion with feature Feature1, and the SFT module outputs feature U1sft; The structure of network unit U2 is: CONV 3×3×32×64 + LeakyReLU + CONV 3×3×64×64 + LeakyReLU, its input U2in is feature U1sft, and the output is feature U2out; After feature U2out is downsampled by the downsampling module and input into the SFT module, it performs feature fusion with feature Feature2, and the SFT module outputs feature U2sft; The structure of network unit U3 is: CONV3×3×64×128 + LeakyReLU + CONV 3×3×128×128 + LeakyReLU, its input U3in is feature U2sft, and the output is feature U3out; After feature U3out is downsampled by the downsampling module and input into the SFT module, it performs feature fusion with feature Feature3, and the SFT module outputs feature U3sft; After feature U3sft passes through the residual network and is input into the SFT module, it performs feature fusion with feature Feature3, and the SFT module outputs feature U4sft; After feature U4sft is upsampled by the upsampling module, it is concatenated (concat) with feature U3out in channels, and the concatenated result is input into network unit U4; The structure of network unit U4 is: CONV3×3×(256 + 128)×256 + LeakyReLU + CONV 3×3×256×128 + LeakyReLU, and its output is feature U4out; Feature U4out is input into the SFT module, and feature fusion is performed with Feature2, and the feature U5sft is output by the SFT module; After the feature U5sft is upsampled by the upsampling module, it is concatenated (concat) with the feature U2out in channels, and the concatenated result is input into the network unit U5; The structure of the network unit U5 is: CONV 3×3×(128 + 64)×128 + LeakyReLU + CONV 3×3×128×64 + LeakyReLU, and its output is the feature U5out; Feature U5out is input into the SFT module, and feature fusion is performed with Feature1, and the feature U6sft is output by the SFT module; After the feature U6sft is upsampled by the upsampling module, it is concatenated (concat) with the feature U1out in channels, and the concatenated result is input into the network unit U6; The structure of the network unit U6 is: CONV 3×3×(64 + 32)×64 + LeakyReLU + CONV 3×3×64×32 + LeakyReLU, and its output is the feature U6out; The structure of the output layer is: CONV 1×1×32×3 + Sigmoid, its input is U6out, and the output is Out, which is the 3-channel RGB image output by the feature fusion network.
4. The RAW image processing method for object detection according to claim 3, wherein In the ISP network, the downsampling module is composed of a dilated convolutional layer, and the structure is: Dil-CONV 3×3×(in_channel)×(in_channel), its stride is 2, the padding is 2, and the dilation is 2.
5. The RAW image processing method for object detection according to claim 3, wherein In the ISP network, the structure of the upsampling module is: CONV 3×3×(in_channel)×(in_channel×4) + PixelShuffle(2), where in_channel represents the number of input channels.
6. The RAW image processing method for object detection according to claim 1, characterized in that, In step 3, the training process of the ISP network is: set the loss function and perform end-to-end offline training on the ISP network; the loss function is specifically: L = Perceptual Loss + Task Loss + Pixel Loss + Bpp Among them, L is the total loss function, Perceptual_Loss is the perceptual loss, Task Loss is the task loss, Pixel_Loss is the pixel loss, and Bpp is the compression degree characterization value.
Citation Information
Patent Citations
Real-time low-illumination image enhancement method based on convolutional neural network
CN116579940A
Low-illumination small target detection method based on SCKConv multi-scale feature fusion enhancement
CN117409244A
Medical image segmentation method based on u-net
US20220309674A1