An Adaptive Image Signal Processing Method, Device and Storage Medium
By modeling the configuration process of the image signal processing system as a Markov decision-making process, and optimizing image processing modules and parameters using a convolutional neural network-based policy network and reinforcement learning algorithm, the problem of difficulty in optimizing image processing pipelines and parameters in the existing technology is solved, and dynamic pipelines and parameter prediction and switching in the inference stage are realized, and the target detection performance is improved.
Patent Information
- Application Number
- CN202410068939.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-01-17
AI Technical Summary
The prior art is difficult to optimize image processing pipelines and parameters at the same time, especially in downstream recognition tasks such as object detection, and it is impossible to dynamically predict the optimal image processing pipelines and parameters, and it is impossible to dynamically switch the image processing pipelines during the inference stage.
The configuration process of the image signal processing system is modeled as a Markov decision-making process, and the selection of image processing modules and its parameter prediction is used to use a policy network based on a convolutional neural network, and the entire image signal processing system is optimized through reinforcement learning algorithms.
It realizes the optimal image processing pipeline and parameters dynamically predicts different inputs in the inference stage, can maximize the object detection performance in different scenarios, and has the ability to switch dynamically from high-precision to low-latency image processing pipelines.
Smart Images

Figure CN118096491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology, and more particularly to an adaptive image signal processing method, device, and storage medium. Background Art
[0002] An image signal processing system consists of multiple different processing modules and parameters. During the design and optimization of an image signal processing system, it is necessary to design the image processing pipeline and optimize the corresponding parameters simultaneously. Therefore, this is a very complex task, usually completed manually by image processing experts. The initial image signal processing system was designed for image quality. For example, in photography-related applications, the goal of image signal processing is to enhance the perceived quality of the image. However, downstream recognition tasks such as object detection have different requirements for the image signal processing system, and most image processing systems tend to inherit the original camera design, with static manually designed pipelines and manually adjusted parameters. This makes it not the best choice for downstream recognition tasks and is also complex to adjust.
[0003] After retrieval, Tseng et al. proposed a parameter optimization algorithm for an image signal processing system based on gradient optimization. They first used a convolutional neural network to proxy the image signal processing system and then used the gradient optimization algorithm to optimize the parameters of the image processing system. However, this method optimizes the parameters based on a manually designed image processing pipeline and does not optimize the image processing pipeline, and the optimized parameters are also fixed during inference and may not be suitable for different scenarios. Qin et al. proposed a method for predicting image processing parameters based on an attention mechanism, which can dynamically predict the corresponding optimal parameters according to different input images during the inference stage, but still does not optimize the image processing pipeline. Although these methods optimize the image processing parameters, they do not optimize the image processing pipeline.
[0004] Yu et al. considered that downstream recognition tasks such as object detection have different requirements for the pipeline of the image processing system. They first used a neural network search algorithm to search for an optimal image processing pipeline and then used the gradient optimization algorithm to optimize its parameters. However, during the inference stage, both the image processing pipeline and the parameters are fixed. Shi et al. modeled the module selection and parameter optimization process of the image processing pipeline as a genetic evolution process and used the genetic evolution algorithm to optimize a set of optimal image processing pipelines and parameters. During the inference stage, both the image processing pipeline and the parameters are fixed. These methods optimize the image processing pipeline and parameters, but during the inference stage, both the image processing pipeline and the parameters are fixed, which cannot meet the requirements of different scenarios simultaneously.
[0005] Hu et al. designed an image processing method for retouching tasks based on reinforcement learning algorithms, which can dynamically modify images to their preferred styles. However, the requirements of image quality tasks and downstream recognition tasks for image signal processing systems are different.
[0006] However, the above-mentioned existing technologies mainly have the following disadvantages:
[0007] 1. During the optimization process of the current method, it is impossible to optimize both the image processing pipeline and parameters for downstream recognition tasks such as object detection at the same time, and this is a complex task.
[0008] 2. During the optimization process of the current method, it does not consider that different image processing modules have different computational costs, which is also necessary for high efficiency in scenarios such as autonomous driving.
[0009] 3. During inference, the current method cannot dynamically predict the optimal image processing pipeline and parameters.
[0010] 4. During inference, the current method layout does not have the ability to dynamically switch from a high-precision image processing pipeline to a low-latency image processing pipeline. Summary of the Invention
[0011] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide an adaptive image signal processing method, device, and storage medium.
[0012] The purpose of the present invention can be achieved by the following technical solutions:
[0013] According to the first aspect of the present invention, an adaptive image signal processing method is provided, which includes:
[0014] Step S1, modeling the configuration process of the image signal processing system as a Markov decision process, where the image signal processing system includes selected image processing modules and their parameters;
[0015] Step S2, selecting image processing modules and predicting their parameters based on the policy network of the convolutional neural network, and estimating the current state value based on the value network of the convolutional neural network;
[0016] Step S3, integrating the pre-trained object detection algorithm into the system as a loss function to calculate the reward of the current policy;
[0017] Step S4, using the reinforcement learning algorithm to optimize and train the entire image signal processing system.
[0018] As a preferred technical solution, the Markov decision process in step S1 is specifically:
[0019] Take the initial image as state s0, and after one processing by the image processing module, the next state obtained is s1, where the total length of the Markov decision is T.
[0020] As a preferred technical solution, the T is adjusted according to actual requirements.
[0021] As a preferred technical solution, each of the image processing modules comes from a predefined image processing module pool.
[0022] As a preferred technical solution, the selection of the image processing module and the prediction of its parameters by the policy network based on the convolutional neural network in step S2 are specifically as follows:
[0023] The policy network based on the convolutional neural network has two heads, namely:
[0024] One network head is responsible for outputting the classification probability, which is used to represent the probability of each module being selected. The module with the highest probability is represented by M t denoted;
[0025] The other network head is responsible for predicting the parameters of the image processing module, and the predicted parameters are represented by Θ t denoted.
[0026] As a preferred technical solution, the value network based on the convolutional neural network in step S2 is used to estimate the current state value specifically as follows:
[0027] The value network based on the convolutional neural network inputs the current state s t and then outputs a value to judge the quality of the current state.
[0028] As a preferred technical solution, in step S3, the pre-trained object detection algorithm is integrated into the system as the loss function, and the reward of the current policy is calculated specifically as follows:
[0029] By inputting the current state and the next state into the object detection network respectively, the change of the loss function is obtained, and this change is used as the reward of the current policy;
[0030] At the same time, the computational consumption of the image processing module is regarded as a constraint term and subtracted from the reward of the current policy.
[0031] As a preferred technical solution, in step S4, the reinforcement learning algorithm is used to optimize and train the entire image signal processing system specifically as follows:
[0032] For the calculation methods of the reward in step S4 and the value in step S3, the reinforcement learning algorithm is used to optimize and train the policy network and the value network.
[0033] According to a second aspect of the present invention, there is provided an electronic device including a memory and a processor, wherein a computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.
[0034] According to a third aspect of the present invention, there is provided a computer-readable storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the method described above is implemented.
[0035] Compared with the prior art, the present invention has the following advantages:
[0036] 1) For each input image, the system of the present invention can automatically generate an optimal image processing pipeline and related image parameters to maximize the object detection performance.
[0037] 2) In the process of designing and optimizing the image processing pipeline and parameters, the present invention can take into account the computational consumption of different image processing modules, and can optimize the optimal image processing pipeline and parameters for specific tasks and computational resource requirements.
[0038] 3) The present invention can dynamically complete the switching from a high-precision image processing pipeline to a low-latency image processing pipeline during the inference stage to meet different computational resource constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of a typical image signal processing pipeline;
[0040] Figure 2 It is a schematic diagram of the adaptive image signal processing method of the present invention;
[0041] Figure 3 It is a schematic diagram of the specific implementation process of the present invention;
[0042] Figure 4 It is a curve graph of the specific effect of the present invention;
[0043] Figure 5 It is a comparison schematic diagram between the present invention and other existing methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] The present invention has a task-driven and scene-adaptive image signal processing method with deep reinforcement learning. Taking object detection as an example of computer vision tasks, for each input image, the system can automatically generate the best image processing pipeline and related image parameters to maximize the object detection performance.
[0046] As Figure 1 shown, a typical image signal processing pipeline includes two parts: the raw domain and the RGB domain. The raw domain converts the raw sensor signal into a linear RGB image, which includes bad pixel correction, black level correction, lens shading correction, and demosaicing. The RGB domain processing further applies custom rendering and post-processing to generate the final image. This processing includes tone mapping, color correction, noise reduction, sharpening / blurring, white balance, gamma correction, exposure adjustment, contrast adjustment, saturation enhancement, and saturation reduction. In the present invention, it is assumed that the captured raw sensor data has been converted into a linear RGB image using simple static raw domain processing, and the focus is on the image processing design and optimization in the RGB domain.
[0047] As Figure 2 shown, the adaptive image processing method of the present invention specifically includes: First, the configuration process of the image signal processing system is modeled as a Markov decision process. The initial image is used as the state s0, and after being processed by an image processing module once, the next state s1 is obtained. The total length of the Markov decision is T, and T can be adjusted according to actual needs. The image processing module includes the selected image processing module and its parameters, and each image processing module comes from a predefined image processing module pool. Second, a policy network based on a convolutional neural network is designed to learn the prediction of the selection of the image processing module and its parameters, and a value network based on a convolutional neural network is designed to estimate the current state value. Immediately afterwards, a pre-trained object detection algorithm is integrated into the optimization system as a loss function to calculate the reward of the current policy. At the same time, the computational consumption of the image processing module is also regarded as a constraint term and subtracted from the reward of the current policy, so that the system can also consider its computational consumption when selecting the module. Finally, a reinforcement learning algorithm is used to optimize and train the entire system.
[0048] Among them, the selection of the image processing module and the prediction of its parameters by the policy network based on a convolutional neural network are specifically as follows: The policy network based on a convolutional neural network has two heads, namely: one network head is responsible for outputting the classification probability, which is used to represent the probability of each module being selected. The module with the highest probability is represented by M t ; the other network head is responsible for predicting the parameters of the image processing module, and the predicted parameters are represented by Θ t .
[0049] The value network based on the convolutional neural network is used to estimate the current state value, specifically: the value network based on the convolutional neural network takes the current state s t as input, and then outputs a value to judge the quality of the current state.
[0050] Integrate the pre-trained object detection algorithm into the system as the loss function to calculate the reward of the current policy, specifically: by inputting the current state and the next state into the object detection network respectively, obtain the change of the loss function, and use this change as the reward of the current policy. At the same time, regard the computational consumption of the image processing module as a constraint term and subtract it from the reward of the current policy, so that the system can also consider its computational consumption when selecting the module.
[0051] Use the reinforcement learning algorithm to optimize and train the entire image signal processing system, specifically: for the above calculation methods of reward and value, adopt the reinforcement learning algorithm to optimize and train the policy network and the value network.
[0052] During inference, as Figure 3 and Figure 4 shown, the adaptive image processing method of the present invention can dynamically predict the image processing pipeline and its corresponding parameters according to different scene inputs to maximize the object detection performance. And it can dynamically complete the switching from the high-precision image processing pipeline to the low-latency image processing pipeline (for example, only running 3 stages) to meet different computational resource constraints.
[0053] Through the above process, the present invention has the following technical effects:
[0054] 1. During the inference stage, it can dynamically predict the image processing pipeline and corresponding parameters according to different inputs.
[0055] 2. During the process of optimizing the design of the image processing pipeline and parameters, it can take into account the computational consumption of different image processing modules and design the most suitable image processing pipeline and parameters. As shown in Table 1, after adding the computational consumption constraint, the probability of selecting the module with slow running speed decreases, and the probability of selecting the module with fast running speed increases, and the overall running speed is increased by 21%, and the detection performance decreases slightly. 3. It has the ability to dynamically switch from the high-precision image processing pipeline to the low-latency image processing pipeline, can dynamically switch the image processing pipeline according to the computational resource constraints during inference, and does not require retraining.
[0056] Table 1
[0057]
[0058] After experimental verification, it can dynamically predict the image processing pipeline and parameters according to different input images, and each prediction only takes 2.2ms, which is very efficient. As Figure 5As shown, compared with several other methods, the present invention can achieve the optimal object detection effect.
[0059] The above is the introduction of the method embodiment. The following further illustrates the solution of the present invention through embodiments of electronic devices and storage media.
[0060] An embodiment of the present invention further provides an electronic device including a central processing unit (CPU), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0061] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0062] The processing unit executes the various methods and processes described above, such as the method of the present invention. For example, in some embodiments, the method of the present invention can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the method of the present invention described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute the method of the present invention in any other suitable manner (e.g., by means of firmware).
[0063] The functions described above herein can be at least partially executed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0064] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0065] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0066] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An adaptive image signal processing method, characterized in that: The method includes: Step S1, modeling the configuration process of the image signal processing system as a Markov decision process, wherein the image signal processing system includes a selected image processing module and its parameters; Step S2, selecting an image processing module and predicting its parameters based on the policy network of the convolutional neural network, and estimating the current state value based on the value network of the convolutional neural network; Step S3, integrating the pre-trained object detection algorithm into the system as a loss function to calculate the reward of the current strategy; Step S4, using a reinforcement learning algorithm to optimize and train the entire image signal processing system; The selection of the image processing module and parameter prediction of the image processing module by the convolutional neural network-based policy network in step S2 are specifically as follows: The policy network based on convolutional neural network has two heads: A network head is responsible for outputting the probability of classification, which is used to represent the probability of each module being selected, where the module with the highest probability is represented by M t express; The other network head is responsible for predicting the parameters of the image processing module. The predicted parameters are expressed as Θ t express; The step S3 integrates the pre-trained target detection algorithm into the system as a loss function, and calculates the reward of the current strategy as follows: By inputting the current state and the next state into the target detection network respectively, the change of the loss function is obtained, and the change is used as the reward of the current strategy; At the same time, the computational cost of the image processing module is taken as a constraint and subtracted from the reward of the current strategy.
2. The adaptive image signal processing method according to claim 1, characterized in that: The Markov decision process in step S1 is specifically as follows: The initial image is taken as state s0. After being processed by the image processing module once, the next state is s1, where the total length of the Markov decision is T.
3. The adaptive image signal processing method according to claim 2, characterized in that: The T is adjusted according to actual needs.
4. The adaptive image signal processing method according to claim 1, characterized in that: Each of the image processing modules is from a predefined image processing module pool.
5. The adaptive image signal processing method according to claim 1, characterized in that: The value network based on the convolutional neural network in step S2 is used to estimate the current state value: The value network based on convolutional neural network converts the current state s t Take input and then output a value to judge whether the current state is good or bad.
6. The adaptive image signal processing method according to claim 1, characterized in that: The step S4 uses a reinforcement learning algorithm to optimize and train the entire image signal processing system, specifically: For the calculation method of the reward of step S4 and the value of step S3, a reinforcement learning algorithm is used to optimize the training strategy network and the value network.
7. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Energy-saving unmanned vehicle path navigation method capable of minimizing useless actions
CN112857373A
Mobile robot navigation method and device, computer equipment and storage medium
CN113609786A