Agent-free ISP hyper-parameter high-speed optimization method and system

Through the agentless ISP hyperparameter optimization method, the parallel decision-making of master-slave agents and nonlinear quantization update registers are solved, and the adaptability and real-time problems of the image processor in dynamic scenarios are realized, and high-speed collaborative optimization and stable imaging of ISP parameters are achieved.

CN120355934APending Publication Date: 2025-07-22上海芯开技术有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510479187.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art image processor hyperparameter optimization method has poor adaptability in dynamic scenarios, distortion of proxy model modeling and lag in response to single agents, resulting in image quality degradation and insufficient system real-time performance.

Method used

The agentless ISP hyperparameter high-speed optimization method is adopted to collect RAW data through the image sensor for black level correction, initialize the ISP register parameters, combine the parallel decision-making of the master and slave agent and nonlinear quantization update registers, and calculate the composite reward based on object detection and image quality to realize the optimization of the multidimensional state vector.

Benefits of technology

High-speed collaborative optimization and stable imaging of ISP parameters in dynamic environments are realized, image quality is improved, object detection error and system response delay are reduced, and data acquisition costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355934A_ABST
    Figure CN120355934A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image signal processing, and discloses an agent-free ISP hyper-parameter high-speed optimization method and system, and the method comprises the following steps: collecting RAW data through a sensor, carrying out the black level correction, and initializing the ISP register parameter mapping; a multi-dimensional state vector is constructed by combining the coded image and the parameters, and the master and slave agents generate discrete adjustment amounts in parallel; after a register is updated through nonlinear quantization, a composite reward is calculated based on target detection and image quality, and data is stored and network parameters are updated through a priority experience pool; the system comprises a sensor module, an ISP hardware processing unit, a master-slave agent decision module, a reward calculation module, a register mapping module and an experience playback module. According to the method, hardware-in-the-loop and deep reinforcement learning are combined, and the problems of proxy model distortion, response lag, high data dependence and parameter mutation are solved through master-slave agent parallel decision, a self-supervision reward mechanism and nonlinear parameter mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image signal processing, and specifically to a method and system for high-speed optimization of ISP hyperparameters without an agent. Background Art

[0002] Traditional methods for optimizing hyperparameters of image processors have long relied on offline parameter tuning and surrogate model simulation, and their inherent defects have gradually emerged in dynamic scenarios. Taking night imaging in autonomous driving as an example, the offline method needs to pre-define hundreds of groups of parameters to cope with different light intensities. However, in actual road conditions, the scene of the superposition of strong light at the tunnel exit and rain and fog scattering exceeds the preset range, resulting in the failure of image dynamic range compression and a sharp drop of more than 30% in the license plate recognition rate. More seriously, engineers rely on laboratory indicators such as PSNR to screen parameters, but it is found in actual measurements that a certain group of "high-score parameters" oversharpen the image edges, leading to a 15% increase in the misjudgment rate of the lane line detection model, exposing the disconnection between the indicators and the real task objectives. This parameter solidification strategy takes up to several weeks, and each adjustment step can only be probed at a granularity of 5%, making it difficult to capture the non-linear coupling effect between white balance and noise reduction.

[0003] The optimization based on the surrogate model attempts to bypass hardware limitations but falls into a new dilemma. A certain manufacturer uses a deep neural network to simulate the ISP pipeline. When training, one hundred thousand groups of laboratory data are used. However, it is found in actual deployment that the model amplifies the face brightness prediction error by 2.3 times in backlight scenarios because the surrogate model cannot reproduce the three-stage tone mapping curve inside the ISP chip. What's more troublesome is that when trying to use Bayesian optimization to search for parameters, the PSNR prediction error of the surrogate model fluctuates by ±3.5dB under low light conditions, resulting in frequent misjudgments of the search direction. This is due to the learning bias of the surrogate model towards the sensor noise pattern - the laboratory data does not cover the thermal noise characteristics of the CMOS chip as the temperature rises, resulting in periodic color block artifacts in the night video stream and a 22% increase in the target tracking ID switching rate.

[0004] The real-time defect of the single-agent architecture is particularly fatal in high-speed scenarios. A certain UAV video transmission system adopts a serial parameter optimization strategy. It takes 280ms to complete the iteration of 12 core ISP parameters, resulting in a smear of 6 meters displacement in the images taken by the aircraft at a speed of 20m / s. Test data shows that when the agent has just optimized the exposure parameter, the scene lighting conditions have changed three times, and the parameter lag causes a 41% loss of dynamic range. A more subtle problem is the parameter game - in a certain optimization, the agent increases the gamma value by 0.2 to improve the dark details, but it causes a ringing effect in the sharpening module, increasing the misjudgment rate of traffic sign recognition from 5% to 19%. This kind of one-sided optimization process makes the system fall into a local optimum in complex scenarios and cannot escape. The measured convergence speed is 4.8 times slower than that of the parallel optimization system; Therefore, the present invention proposes a proxy-free high-speed optimization method and system for ISP hyperparameters to solve the deficiencies of the prior art. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention provides a proxy-free high-speed optimization method and system for ISP hyperparameters, which solves the problems of poor adaptability in traditional offline parameter tuning scenarios, distorted proxy model modeling, and lag in single-agent response, and realizes high-speed collaborative optimization of ISP parameters and stable imaging in a dynamic environment.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A proxy-free high-speed optimization method for ISP hyperparameters includes the following steps: S1. Collect original RAW data through an image sensor, perform black level correction on the RAW data, and generate preprocessed RAW data; S2. Initialize the hyperparameters in the ISP hardware register, and establish a mapping relationship between the parameter index and the register address; S3. Jointly encode the RGB image of the current frame and the ISP parameter vector to construct a multi-dimensional state vector; S4. Based on the global image features extracted by the main agent, multiple slave agents parallelly generate discrete adjustment amounts for the corresponding ISP parameters to form an action vector; S5. Calculate the register parameter increment according to the action vector, and update the ISP hardware register through non-linear transformation and quantization operations; S6. Calculate the composite reward signal for the current parameter adjustment based on the object detection result and the image quality index; S7. Store the state transition data in the prioritized experience pool, and update the agent network parameters based on the temporal difference error.

[0007] Preferably, the construction method of the multi-dimensional state vector is as follows: Input the RGB image into a feature extraction network including an atrous convolution layer, and output the compressed spatial features; Perform temporal encoding on the current ISP parameter and the historical parameter difference sequence to generate parameter evolution features; Concatenate the spatial features and the parameter evolution features to obtain the multi-dimensional state vector.

[0008] Preferably, the dilation rate of the atrous convolution layer increases exponentially according to the layer number k, and the dilation rate of the k-th layer is 2 k .

[0009] Preferably, the interaction method between the main agent and the slave agents is as follows: The main agent extracts the global features of the RGB image through depthwise separable convolution; Each receives the current values of the global features and the corresponding ISP parameters from the agent and outputs Q values with seven discrete adjustment amplitudes; The action vector is composed of the adjustment amounts corresponding to the maximum Q values selected by each slave agent.

[0010] Preferably, the seven discrete adjustment amplitudes are ±30%, ±10%, ±3% of the current parameter value, and a zero adjustment amount.

[0011] Preferably, the calculation method of the register parameter increment is as follows: Perform a non-linear square root transformation on the adjustment amount of each parameter to obtain an intermediate adjustment amount; Perform dynamic range quantization on the intermediate adjustment amount according to the register bit width to generate a register write value.

[0012] Preferably, the non-linear square root transformation is expressed as: ; Where: is the adjustment amount written to the register; is the parameter adjustment amount; is the sign function, indicating the sign of the parameter adjustment amount; is the square root of the absolute value of the parameter adjustment amount.

[0013] Preferably, the calculation method of the composite reward signal is as follows: Generate a spatial Gaussian weight matrix based on the target detection box coordinates; Calculate the weighted peak signal-to-noise ratio and the weighted structural similarity index according to the weight matrix; Fuse the mean average precision of target detection to generate a final reward value.

[0014] Preferably, the sampling probability distribution method of the prioritized experience replay pool is as follows: Calculate the priority according to the absolute value of the temporal difference error of the state transition data; Adopt a double Q-network architecture to update the policy network parameters, and the target network parameters are updated synchronously at a fixed ratio Synchronously update.

[0015] The present invention also provides a proxy-free ISP hyperparameter high-speed optimization system, including: A sensor module for collecting and preprocessing RAW data; An ISP hardware processing unit including programmable registers and an image processing pipeline; A master-slave agent decision-making module composed of a master agent feature extraction network and multiple slave agent parameter optimization networks; A reward calculation module integrating an image quality evaluation unit and a target detection unit; The register mapping module realizes the non - linear conversion from the action vector to the register value; The experience replay module manages the storage and sampling queue of the state transition data.

[0016] The present invention provides a method and system for high - speed optimization of ISP hyperparameters without an agent, having the following beneficial effects: 1. The present invention adopts the hardware - in - the - loop design combined with deep reinforcement learning, directly interacting with the ISP hardware in real - time, achieving the optimal dynamic image quality. Compared with the existing technology that relies on the parameter tuning of the offline proxy model, it solves the optimization deviation problem caused by the model distortion. For example, in a low - light scene, the system can adaptively enhance the noise reduction intensity by real - time sensing the noise distribution, reducing the target detection box positioning error by 40%.

[0017] 2. The present invention realizes multi - parameter parallel decision - making and fast response through the master - slave multi - agent collaborative architecture. Compared with the traditional single - agent serial parameter tuning method, it solves the problems of high computational delay and inability to cope with sudden environmental changes. For example, in an autonomous driving scene with sudden strong light, the system can complete the linked adjustment of exposure and tone parameters in a short time, avoiding image over - exposure caused by frame - to - frame delay in the traditional method.

[0018] 3. The present invention is based on hardware - in - the - loop and self - supervised learning, directly extracting the optimization signal from real imaging data, achieving the effect of unsupervised and efficient training. Compared with the existing technology that relies on training the proxy model with a large amount of labeled data, it solves the defects of high data acquisition cost and insufficient scene coverage. For example, only 100 groups of unlabeled road surveillance videos are required to complete model convergence, while the traditional method requires 5000 groups of calibrated data.

[0019] 4. The present invention uses the non - linear conversion and dynamic compensation mechanism to realize the smooth transition of register parameters. Compared with the existing technology of direct linear mapping, it solves the problems of image flicker and hardware deadlock caused by parameter mutation. For example, through square - root transformation and temperature compensation, the fluctuation range of the parameter adjustment amount is reduced to 1 / 3 of the traditional method, and the abnormal trigger rate of the register is reduced from 8% to 0.2%. Brief Description of the Drawings

[0020] Figure 1 is the flowchart of the method of the present invention; Figure 2 is the system architecture diagram of the present invention; Figure 3 is the data processing flowchart of the present invention; Figure 4 is the schematic diagram of the neural network of the present invention. Detailed Embodiments

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.

[0022] Please refer to Figure 1 , the embodiments of the present invention provide a method for high-speed optimization of ISP hyperparameters without an agent, including the following steps: S1. Collect the original RAW data through an image sensor, perform black level correction on the RAW data, and generate preprocessed RAW data; In the method for optimizing ISP hyperparameters without an agent, step S1, as the initial data acquisition stage of the entire technical process, needs to form a stable data interface with the subsequent parameter optimization module. The execution quality of this step directly affects the construction accuracy of the subsequent state vector. It is necessary to ensure the standardized input format and noise suppression effect of the RAW data to provide a reliable data basis for reinforcement learning decision-making.

[0023] In this embodiment, the specific implementation method for collecting the original RAW data through an image sensor includes the following technical features: In some embodiments, the image sensor uses a CMOS image sensor arranged in a Bayer array, and its output RAW data format is 12-bit linear data. The sensor is connected to the ISP hardware processing unit through a MIPI CSI-2 interface, and the transmission protocol is configured as a four-channel LVDS differential signal, and the clock frequency is set to 1.5 GHz. At this time, the transmission delay of a single-frame RAW data with a resolution of 1920×1080 is controlled within 2.3 ms. Specifically, the effective photosensitive area of the sensor needs to shield the optical black region through register configuration to avoid invalid data mixing.

[0024] For the black level correction operation, its mathematical expression form is: ; Among them, is the original pixel value output by the sensor; represents the black level compensation matrix related to the pixel position; is the gain calibration coefficient matrix.

[0025] In a possible implementation manner, the construction method of the black level compensation matrix is: continuously collect frames of RAW data under the condition of no light, and calculate the average value of each pixel position: ; Among them, is the number of dark field acquisition frames, and the value range is 32 - 128; represents the th pixel value at the coordinate in the dark field data of the

[0026] Specifically, the gain calibration coefficient matrix is determined by the following method: ; Among them, is the output value of the sensor under the standard uniform light source; is the ideal response value, which is obtained from the gray scale card calibration experiment.

[0027] The color temperature of the standard uniform light source needs to be stable within the range of 5500K ± 100K, and the illuminance is set to 1000 lux.

[0028] Furthermore, the preprocessed RAW data needs to meet the following technical requirements: Data bit width normalization processing: linearly map the 12-bit RAW data output by the sensor to 16-bit fixed-point number representation, and the mapping relationship is: ; Among them, is the 16-bit normalized pixel value; is the maximum value of the 12-bit data ; is the maximum value of the 16-bit data ; represents the floor operation.

[0029] Ring point correction: replace abnormal pixel values based on the mean filtering algorithm of adjacent pixels, and its determination condition is: ; Among them, represents the mean value of the 3×3 neighborhood around the current pixel; is the neighborhood standard deviation; 3 is the empirical value threshold coefficient.

[0030] The calculation method of the corrected pixel value is: ; Among them, represents 8 pixels except the center point in the 3×3 neighborhood; is the neighborhood pixel validity flag (normal pixel is 1, bad pixel is 0).

[0031] As an extended technical solution, in some embodiments, a non-linear correction factor is introduced to adapt to the sensor response characteristics: ; Among them, is the normalization coefficient, satisfying ; are the fitting parameters of the sensor response curve, obtained by fitting the experimental data through the least squares method.

[0032] At the hardware implementation level, the data interaction relationship between the black level correction module and the ISP register is ensured by the following methods: Dual-buffer storage area: Set two independent DDR3 storage areas (Buffer-A and Buffer-B). When the current frame correction data is written to Buffer-A, the ISP processing unit reads the previous frame data from Buffer-B to eliminate access conflicts through ping-pong operation; Timing synchronization mechanism: Use the line synchronization (HSYNC) and frame synchronization (VSYNC) signals output by the sensor to trigger the DMA transfer to ensure the alignment of RAW data blocks. Specifically, when the rising edge of VSYNC is triggered, the DMA controller batches transfers the current frame data from the sensor FIFO to the preprocessing module; the HSYNC signal is used to verify the integrity of each line of data. If 3 consecutive HSYNC pulses are lost, the data retransmission mechanism is triggered.

[0033] It should be noted that when different models of image sensors are used, the black level compensation matrix and the gain calibration coefficient need to be dynamically updated through the online calibration process. The online calibration process is automatically executed during the ISP initialization phase and includes the following sub-steps: Dark field data acquisition: Control the mechanical shutter to completely block the incident light and continuously acquire at least 32 frames of RAW data; Statistic calculation: Calculate the mean value of the dark current pixel by pixel and variance , filter out abnormal pixels with variance greater than the set threshold (such as ); Compensation table burning: Write the calibrated and into the non-volatile memory (NAND-Flash) of the ISP hardware and establish a register address mapping table.

[0034] During the dark field data acquisition process, it is necessary to monitor the temperature of the sensor chip through a temperature sensor (such as DS18B20) and adjust the rotation speed of the cooling fan through the PID control algorithm to maintain the working temperature within the range of 25±2°C to suppress thermal noise interference. If the temperature fluctuation exceeds ±5°C, the calibration process is triggered to be executed again.

[0035] S2. Initialize the hyperparameters in the ISP hardware register and establish the mapping relationship between the parameter index and the register address; In the proxy-free ISP hyperparameter optimization method, step S2 serves as the hardware interaction foundation for the parameter optimization process and needs to form a stable parameter control link with the preprocessed RAW data output by step S1. The initialization operation ensures that the ISP processing unit is in a tunable state and provides an operable parameter space for subsequent reinforcement learning decisions.

[0036] In this embodiment, the construction methods of ISP hardware register initialization and parameter mapping relationships include the following technical features: In some embodiments, the initialization parameter set of the ISP hardware register is preloaded through a non-volatile memory. Specifically, the manufacturer's preset parameter table is read from a specified sector of NANDFlash, and this parameter table contains the initial register values of 12 core parameters such as brightness, contrast, and white balance. For example, the default register address for the contrast parameter is 0x1A3F0004, and the initial value is 0x1FF (10-bit register bit width).

[0037] For the mapping relationship between the parameter index and the register address, its mathematical expression form is: ; Where, is the mapping function from the parameter index to the register address, bit width, and range; represents the total number of adjustable parameters. In this embodiment ; is the physical address space of the register, defined as the 32-bit address range 0x1A3F0000 - 0x1A3F00FF; is the register bit width set, including three types of bit widths; represents the real number interval of the physical range of the parameter .

[0038] As an option, the mapping relationship is implemented through a lookup table, and its data structure includes the following fields: parameter index, register physical address, bit width, lower range limit, upper range limit, and non-linear compensation flag bit. In a possible implementation, the mapping entry for the white balance R gain parameter is: parameter index 3, address 0x1A3F0008, bit width 14, range [0.8, 1.2], non-linear compensation flag bit enabled.

[0039] Specifically, the quantization conversion method of the register parameter is: ; Where, is the adjustment amount written to the register; is the adjustment amount of the th parameter, normalized to the interval [-1, 1]; is the The register bit width corresponding to a parameter; and are respectively the upper and lower limits of the physical measurement range of the parameter; indicates rounding operation.

[0040] Taking the white balance R gain parameter as an example, its physical measurement range is [0.8, 1.2], and the register bit width is 14 bits. Then the quantization scale factor is calculated as: ; When the adjustment amount is .

[0041] In a possible implementation manner, the initialization process includes the following steps: Send a hardware reset instruction (0xFE) through the I2C bus to restore the ISP register to the default state; Read the parameter index, register address, and initial value item by item from the preset parameter table; Write to the register using a double-buffer mechanism: write the initial value to the shadow register group, and atomically switch to the main register group after being triggered by the vertical blanking signal (VSYNC) to avoid image output tearing.

[0042] As an extended technical solution, in some embodiments, a non-linear quantization compensation strategy is introduced to adapt to the response characteristics of the ISP chip: ; wherein, is the non-linear compensation coefficient of the parameter, which is determined through a calibration experiment, and the typical value range is [0.8, 1.5]; is the exponential factor used to adjust the non-linearity degree of the adjustment amount, and the default value is set to 0.5; is the sign function that returns the positive and negative polarity of the adjustment amount.

[0043] For example, the brightness parameter is set to in the low adjustment amount region ( ) to enhance the fine-tuning sensitivity; and is set to in the high adjustment amount region ( ) to prevent register value overflow.

[0044] For the update of the dynamic parameter mapping table, when detecting a change in the ISP hardware model, perform the following operations: Read the device ID (16-bit vendor ID + 16-bit device ID) through the PCIe configuration space and match the pre-stored register mapping template library; If there is no matching item in the template library, start the automatic detection process: within the preset address range (0x1A3F0000 - 0x1A3F00FF), write test values (such as 0xAAAA, 0x5555) in sequence, and verify the parameter validity through the change of the image processing result; Create a new mapping table entry and store it in the non-volatile memory, and record the detected register bit width and range at the same time.

[0045] It should be noted that the register writing operation needs to meet strict timing constraints. Specifically, the parameter update is completed within the vertical blanking period of each frame, and its timing window is calculated as: ; Among them, is the total time consumption for parameter update; is the duration of the vertical blanking period, and the typical value is 1.2ms (at a refresh rate of 120Hz); is the time consumption for single register writing, about 50ns (based on the I2C 400kHz clock rate); is the number of registers to be updated in a single frame. In this example, .

[0046] The calculated total update time , is much less than , meeting the real-time requirement. If it is detected that the update time exceeds the limit, trigger the degradation strategy: only update the top 6 parameters with the highest priority, and the remaining parameters are delayed until the next frame for processing.

[0047] S3. Jointly encode the RGB image of the current frame and the ISP parameter vector to construct a multi-dimensional state vector; In the proxy-free ISP hyperparameter optimization method, step S3, as the core of state perception for reinforcement learning decision-making, needs to form a spatio-temporal association with the register parameters initialized in step S2 and the preprocessed RAW data generated in step S1. The construction of the multi-dimensional state vector integrates the current image features and the parameter evolution trend, providing an interpretable decision-making basis for the intelligent agent.

[0048] In this embodiment, the construction method of the multi-dimensional state vector includes the following technical features: In some embodiments, the feature extraction of the RGB image is implemented by a hierarchical dilated convolutional network. Specifically, the RGB image (resolution 1920×1080×3) output in step S1 is input into a feature extraction network including 6 convolutional layers. The size of the convolutional kernel in each layer is fixed at 3×3, and the number of channels decreases layer by layer to 64, 32, 16, 8, 4, 2. As an option, the dilation rate of the dilated convolutional layer increases exponentially according to the layer number k, and its mathematical relationship is expressed as: ; Among them, is the dilation rate of the th layer; is the layer number.

[0049] For example, the dilation rate of the first layer is 1 (standard convolution), the dilation rate of the third layer increases to 4, and the dilation rate of the sixth layer reaches 32. The dilation rate is designed such that the shallow network captures local details and the deep network perceives large - scale context associations.

[0050] For the generation of the parameter evolution feature, its mathematical expression is: ; Among them, is the parameter evolution feature vector at the current moment (dimension 128×1); represents the parameter difference vector between adjacent frames (dimension 12×1); is the historical window length, set to 5 frames; is the network parameter, including the weight matrices and bias terms of the input gate, forget gate, and output gate.

[0051] The network has a hidden layer dimension of 128. The initial state vector is initialized to zero, and a gradient clipping strategy (threshold 1.0) is adopted to prevent training divergence.

[0052] In a possible implementation, the concatenation operation of the spatial feature and the parameter evolution feature is defined as: ; Among them, is the final multi - dimensional state vector (dimension 640×1); is the spatial feature output by the dilated convolution network (dimension 512×1); represents the concatenation operation along the feature dimension.

[0053] Before concatenation, the feature vectors need to be dimension - aligned: the spatial feature is compressed to 512 dimensions through a fully - connected layer, the parameter evolution feature remains 128 dimensions, and finally a 640 - dimensional state vector is generated by concatenation.

[0054] As an extended technical solution, in some embodiments, a feature normalization strategy is introduced to improve training stability: ; ; Among them, and are the mean and standard deviation of the spatial feature, calculated through moving average (window length 1000 frames); and is a statistic of the parameter evolution feature, and its calculation method is the same as the above; , , , are learnable scaling and offset parameters; is a small constant to prevent division-by-zero errors, and is set to 1e-5.

[0055] It should be noted that the acceleration implementation of the dilated convolutional network on the hardware side is completed through the following methods: Computation optimization: Use the Winograd algorithm to optimize the 3×3 convolution calculation, reducing the number of multiplication operations to 2.25 times that of the standard convolution; Memory management: Utilize the GPU shared memory to cache the input feature map, reducing the number of global memory accesses and increasing the throughput by 40%; Sparse processing: Enable sparse matrix compression storage for convolutional layers with dilation rates greater than 8 (such as the 5th and 6th layers), only retaining the coordinates and values of non-zero weights, and reducing the memory occupancy to 30% of the original size.

[0056] For scenarios with strict real-time requirements, the timing encoding process of the LSTM network is optimized through the following methods: Pre-computation mechanism: Immediately calculate the current parameter difference after the register update in step S5 , and store it in a circular buffer (FIFO queue) with a capacity of 5; Parallel processing: Expand the calculation of the time steps of the LSTM and allocate them to the multi-core CPU thread pool, with each time step assigned an independent thread, and the single-inference time is compressed from 3.2ms to 1.1ms; Quantization compression: Quantize the output 128-dimensional floating-point feature vector into 8-bit integers, and the scaling factor , and the dynamic range loss is controlled within 2%.

[0057] S4. Based on the global image features extracted by the main agent, multiple slave agents generate discrete adjustment amounts corresponding to the ISP parameters in parallel to form an action vector; In the agent-free ISP hyperparameter optimization method, step S4, as the core of action generation in reinforcement learning, needs to form a dynamic interaction with the multi-dimensional state vector constructed in step S3. The master-slave agent architecture realizes parameter collaborative optimization under the guidance of global features, ensuring the physical consistency of multi-dimensional adjustment amounts.

[0058] In this embodiment, the interaction method between the main agent and the slave agents includes the following technical features: In some embodiments, the main agent uses a depthwise separable convolutional network to extract global image features. Specifically, the RGB image (resolution 1920×1080×3) generated in step S3 is input, and through three layers of depthwise separable convolutional layers, it is downsampled step by step to a feature map of 120×67×256; the size of the convolutional kernel in each layer is 3×3, the stride is 2, and the activation function is LeakyReLU (negative slope 0.01). The mathematical expression of the depthwise separable convolution is: ; where, is the input RGB image( ); is the per-channel convolutional kernel weight, and each input channel is independently convolved; is the pointwise convolutional kernel weight for channel fusion; is the bias term; is the output global feature tensor.

[0059] As an extended technical solution, the decision-making logic of the slave agent is implemented in the following way: Parameter feature injection: Each slave agent receives the current value of the corresponding ISP parameter and its historical mean , and generates a parameter feature vector through a fully connected layer; Feature fusion: Flatten the global feature output by the main agent into a vector and compress it to 256 dimensions through a fully connected layer, and concatenate it with the parameter feature vector to form ; Q-value calculation: Input a two-layer fully connected network and output the Q-values corresponding to seven discrete adjustment amplitudes: ; where, are the weights and biases of the first layer; are the weights and biases of the second layer; is the discrete adjustment action.

[0060] In a possible implementation, the physical mapping relationship of the seven discrete adjustment amplitudes is: ; where, corresponds to the seven adjustment amplitudes; and are the upper and lower limits of the physical measurement range of the th parameter.

[0061] It should be noted that the action selection mechanism ensures the exploration-exploitation balance in the following way: - Greedy strategy: Initial exploration rate , linearly decays every 1000 training steps , with a lower bound ; Action masking: When the parameter is near the physical range boundary (such as ), the forward adjustment action is masked ; Prioritized sampling: Adjust the action selection probability according to the temporal difference error . If , then select the sub-optimal action (the action with the second highest Q-value) with a 50% probability.

[0062] For the hardware acceleration implementation, the parallel execution of the slave agents is optimized in the following ways: Kernel function grouping: The Q-value calculations of 12 slave agents are distributed to 12 Streaming Multiprocessors (SMs) of the GPU. Each SM independently processes one parameter, and the thread block is configured with 256 threads / block; Memory sharing: The global features are stored in the Constant Memory, reducing duplicate data transmission and increasing the bandwidth utilization rate to 92%; Register reuse: Different threads within the same SM reuse the intermediate calculation results (such as the output of the activation function), increasing the calculation throughput by 2.8 times and reducing the inference time of a single agent to 0.8 ms.

[0063] S5. Calculate the register parameter increment according to the action vector, and update the ISP hardware register through non-linear transformation and quantization operations; In the agentless ISP hyperparameter optimization method, step S5, as the final execution stage of parameter adjustment, needs to convert the action vector generated in step S4 into a register operation instruction recognizable by the hardware. The update process needs to take into account the physical range constraints of the parameters and the timing characteristics of the hardware interface to ensure the stability and real-time performance of parameter adjustment.

[0064] In this embodiment, the specific implementation methods of the register parameter increment calculation and hardware update include the following technical features: In some instances, the conversion process from the action vector to the register parameter is achieved through non-linear square root transformation and dynamic range limitation. Specifically, for the action adjustment amount of the first parameter, its corresponding register increment calculation is divided into two stages: Stage 1: Non-linear square root transformation The mathematical form of the transformation is: ; Where: Adjustment value written to the register; Parameter adjustment value; Is a sign function, representing the sign of the parameter adjustment value; Is the square root of the absolute value of the parameter adjustment value.

[0065] Stage 2: Dynamic range limitation and quantization Map the non - linear transformation result to the register bit - width and apply physical range constraints: ; Among them, Represents the rounding operation; Is a clipping function, constraining the output within the interval ; Is the register bit - width corresponding to the th parameter (for example, 14 bits correspond to the maximum register value 16383).

[0066] In a possible implementation, the non - linear transformation further introduces a scaling factor to adapt to the hardware response characteristics: ; Among them, Is a scaling coefficient, determined through a calibration experiment; the calibration method is: adjust the value under a uniform grayscale image to maximize the PSNR value of the output image.

[0067] As an extended technical solution, the register update operation ensures timing security in the following ways: Shadow register group: Set a storage area that is a complete mirror of the main register group. When updating, first write to the shadow register to avoid intermediate states during parameter writing; Vertical blanking synchronization: Use the rising edge of the vertical blanking signal (VSYNC) output by the sensor to trigger an atomic switch, and batch - copy the content of the shadow register to the main register group, controlling the switching time within 200 ns; Abnormal detection and rollback: If it is detected that the updated register value exceeds the physical range (such as ), then immediately trigger a hardware interrupt and restore from the backup register group to the first valid parameter.

[0068] For hardware interface optimization, the register write operation adopts the following strategy: Batch DMA transfer: Pack the register addresses and data of 12 parameters into a continuous memory block (the data structure contains a 32 - bit address + 32 - bit value), and directly write to the ISP hardware memory through the DMA control logic, compressing the single - transfer time from about 120 μs to 8 μs; Temperature Compensation Mechanism: According to the readings of the on-board temperature sensor , dynamically adjust the scaling factor : ; Among them, is the temperature compensation coefficient, compensating for the change of the sensor dark current with temperature; is the reference temperature; is read from the internal register 0x01D0 of the sensor through the I2C port, with an accuracy of ±1°C.

[0069] It should be noted that the register address mapping table is read through the PCIe configuration space during the system initialization phase, and its data structure includes the following fields: Base Address: The starting address of the ISP register group (such as Ox1A3F000); Address Offset: The offset of the register address of each parameter relative to the base address (such as the contrast parameter offset is 0x0004); Bit Width Identification: 4-bit binary encoding represents the register bit width (for example, 0010 represents 10 bits, and 1100 represents 14 bits); Range Metadata: Two 16-bit floating-point numbers respectively store the lower limit of the physical range and the upper limit .

[0070] S6. Based on the object detection results and image quality metrics, calculate the composite reward signal for the current parameter adjustment; In the agent-free ISP hyperparameter optimization method, step S6, as the core of the reinforcement learning feedback mechanism, needs to form a closed-loop association with the ISP parameters updated in step S5 and the RGB image output in step S3. The composite reward signal fuses the object detection accuracy and image quality metrics, providing a multi-dimensional optimization guidance for the agent.

[0071] In this embodiment, the calculation method of the composite reward signal includes the following technical features: In some examples, the generation method of the spatial Gaussian weight matrix is: based on the center coordinates of the object detection box and the box size to construct a two-dimensional Gaussian distribution weight map. Specifically, for an image with a size of , the weight matrix The calculation formula is: ; Among them, is the center coordinate of the object detection box, output by the object detection model (such as YOLOv5); is the Gaussian kernel at and The standard deviation in the direction has a linear relationship with the size of the detection box; and are the width and height of the detection box respectively, in pixels; is the pixel coordinate index.

[0072] For the calculation of the weighted peak signal-to-noise ratio (WPSNR), its mathematical expression is: ; where, is the weighted peak signal-to-noise ratio (unit: dB); is the reference image, collected in the laboratory calibration environment; is the output image after the current ISP processing; is the normalization coefficient; is the maximum pixel value of the reference image, which is 255 for an 8-bit image.

[0073] In a possible implementation, the calculation method of the weighted structural similarity index (WSSIM) is: ; where, is the weighted structural similarity index (dimensionless); and are respectively and at the position The local mean at, calculated through a Gaussian window, with a Gaussian kernel standard deviation of 1.5; and are the local standard deviations; is the local covariance; , is the stability constant, used to avoid division-by-zero errors.

[0074] As an extended technical solution, the fusion method of the mean average precision (mAP) of object detection is: ; where, is the final reward value (range [0,1]); is the normalization function, mapping each index to the interval; The normalization range of is set to , ; The normalization range of is set to , ; The normalization range of is set to , ; , , is a weighting coefficient, which is optimized and determined on the validation set through grid search.

[0075] It should be noted that the real-time guarantee of the object detection model is achieved through the following methods: Model lightweighting: The channel pruning technique is used to reduce the number of convolutional kernels of YOLOv5 by 40%. The number of parameters is compressed from 7.5M to 4.3M, and the inference speed is increased from 120ms / frame to 65ms / frame; Hardware acceleration: The model is converted to FP16 precision through TensorRT and deployed on NVIDIA Jetson AGX Xavier, and the inference time is further reduced to 28ms / frame; Asynchronous computing: When the ISP processes the current frame, the object detection model processes the previous frame image in parallel, and the processing delay is hidden within 5ms through the double-buffer mechanism.

[0076] S7. Store the state transition data in the prioritized experience replay pool and update the agent network parameters based on the temporal difference error; In the agent-free ISP hyperparameter optimization method, step S7, as the core of training and optimization in reinforcement learning, needs to form a data closed-loop with the composite reward signal calculated in step S6 and the action vector generated in step S4. The experience replay pool storage mechanism and network update strategy achieve efficient sample utilization and policy iteration optimization.

[0077] In this embodiment, the construction of the prioritized experience replay pool and the agent network update method include the following technical features: In some embodiments, the sampling probability distribution method of the prioritized experience replay pool is based on the dynamic priority calculation of the temporal difference error (TD-Error). Specifically, for each state transition data , its priority The calculation formula is: ; where is the temporal difference error; is the discount factor, and the default value is 0.99; is the Q-value function of the target network; is the Q-value function of the policy network; is a small constant to prevent zero priority.

[0078] In a possible implementation manner, the update strategy of the double Q-network architecture includes the following operations: Update of policy network parameters: The gradient descent method is used to minimize the weighted mean square error loss function: ; Among them, is the importance sampling weight; is the experience pool capacity; is the annealing coefficient, with an initial value of 0.4, linearly increasing by 0.002 every 1000 training steps; is the number of mini - batch samples.

[0079] Target network parameter synchronization: The policy network and target network parameters are proportionally mixed in a relatively updated manner: ; Among them, is the synchronization ratio coefficient; are the policy network parameters; are the target network parameters.

[0080] As an extended technical solution, the storage and sampling mechanism of the prioritized experience pool is optimized as follows: Segmented storage structure: The experience pool is divided into 8 storage segments (Segments), with each segment having a capacity of 6250 pieces of data. When writing, the old data is replaced according to the circular overwrite strategy; Hierarchical sampling: The data is divided into three levels of high (Top - 10%), medium (Next - 30%), and low (Remaining - 60%) according to priority, and the sampling ratio is set to 5:3:2 to ensure the utilization rate of high - value samples; Priority decay: Apply an exponential decay factor to the data stored for more than frames, and the update formula is: ; Among them, is the data storage duration (number of frames); is the current training step; is the training step when the data is stored in the experience pool.

[0081] It should be noted that the action selection strategy of the target network achieves stability guarantee through the following methods: ; ; Among them, selects the optimal action by the policy network; calculates the target Q - value by the target network.

[0082] For hardware - accelerated implementation, the following optimization measures are adopted for the gradient calculation process: Mixed-precision training: Store network parameters in FP16 format, retain FP32 precision for activation values and loss calculations, reduce memory occupancy by 50%, and increase the calculation speed by 1.8 times; Asynchronous data loading: While the GPU calculates gradients, preload the next batch of sample data into the video memory through the DMA engine, and increase the training throughput to 12,000 frames per second; Gradient clipping: Apply L2 norm constraint to the gradients of the policy network ; ; wherein, is the gradient clipping threshold; is the L2 norm of the gradient vector.

[0083] Please refer to Figure 2 , the present invention also provides a proxy-free ISP hyperparameter high-speed optimization system, including: A sensor module for collecting and preprocessing RAW data; An ISP hardware processing unit, including programmable registers and an image processing pipeline; A master-slave agent decision-making module, composed of a master agent feature extraction network and multiple slave agent parameter optimization networks; A reward calculation module, integrating an image quality evaluation unit and an object detection unit; A register mapping module for realizing the non-linear conversion from the action vector to the register value; An experience replay module for managing the storage and sampling queue of state transition data.

[0084] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A proxy-free ISP hyperparameter high-speed optimization method, characterized in that, It includes the following steps: S1. Collect the original RAW data through an image sensor, perform black level correction on the RAW data, and generate preprocessed RAW data; S2. Initialize the hyperparameters in the ISP hardware register, and establish a mapping relationship between the parameter index and the register address; S3. Jointly encode the RGB image of the current frame with the ISP parameter vector to construct a multi-dimensional state vector; S4. Based on the global image features extracted by the main agent, multiple slave agents parallelly generate discrete adjustment amounts corresponding to the ISP parameters to form an action vector; S5. Calculate the register parameter increment according to the action vector, and update the ISP hardware register through non-linear transformation and quantization operations; S6. Calculate the composite reward signal for the current parameter adjustment based on the target detection result and the image quality index; S7. Store the state transition data in the prioritized experience replay pool, and update the agent network parameters based on the temporal difference error.

2. The ISP hyperparameter high-speed optimization method without an agent according to claim 1, characterized in that, The construction method of the multi-dimensional state vector is as follows: Input the RGB image into a feature extraction network including atrous convolutional layers to output compressed spatial features; Perform temporal encoding on the difference sequence between the current ISP parameters and the historical parameters to generate parameter evolution features; Concatenate the spatial features and the parameter evolution features to obtain the multi-dimensional state vector.

3. The method for high-speed optimization of ISP hyperparameters without an agent according to claim 2, wherein The dilation rate of the dilated convolutional layer increases exponentially according to the layer number k, and the dilation rate of the k-th layer is 2 k .

4. A proxy-free ISP hyperparameter high-speed optimization method according to claim 1, characterized in that The interaction method between the main agent and the slave agents is as follows: The main agent extracts the global features of the RGB image through depthwise separable convolution; Each slave agent receives the global features and the current values of the corresponding ISP parameters, and outputs Q values of seven discrete adjustment amplitudes; The action vector is composed of the adjustment amounts corresponding to the maximum Q values selected by each slave agent.

5. The method for high-speed optimization of ISP hyperparameters without an agent according to claim 4, characterized in that The seven discrete adjustment amplitudes are ±30%, ±10%, ±3% of the current parameter value and zero adjustment amount.

6. The agentless ISP hyperparameter high-speed optimization method according to claim 1, characterized in that The calculation method of the register parameter increment is as follows: Perform non-linear square root transformation on the adjustment amount of each parameter to obtain an intermediate adjustment amount; Perform dynamic range quantization on the intermediate adjustment amount according to the register bit width to generate a register write value.

7. A proxy-free ISP hyperparameter high-speed optimization method according to claim 6, characterized in that The non-linear square root transformation is expressed as: ; Wherein: is the adjustment amount written to the register; is the parameter adjustment amount; is the sign function, representing the sign of the parameter adjustment amount; is the square root of the absolute value of the parameter adjustment amount.

8. An agentless ISP hyperparameter high-speed optimization method according to claim 1, characterized in that, The calculation method of the composite reward signal is as follows: Generate a spatial Gaussian weight matrix based on the target detection box coordinates; Calculate the weighted peak signal-to-noise ratio and the weighted structural similarity index according to the weight matrix; Fuse the mean average precision of target detection to generate a final reward value.

9. An agentless ISP hyperparameter high-speed optimization method according to claim 1, characterized in that The sampling probability distribution method of the prioritized experience replay pool is as follows: Calculate the priority according to the absolute value of the temporal difference error of the state transition data; The double Q-network architecture is adopted to update the parameters of the policy network, and the target network parameters are updated synchronously at a fixed ratio ​ 10. A proxy-free ISP hyperparameter high-speed optimization system, which is applied to a proxy-free ISP hyperparameter high-speed optimization method according to any one of claims 1-9, and is characterized in that, It includes: A sensor module for collecting and preprocessing RAW data; An ISP hardware processing unit including programmable registers and an image processing pipeline; A master-slave agent decision-making module composed of a main agent feature extraction network and multiple slave agent parameter optimization networks; A reward calculation module integrating an image quality evaluation unit and a target detection unit; A register mapping module for realizing the non-linear conversion from the action vector to the register value; An experience replay module for managing the storage and sampling queue of state transition data.

Citation Information

Cited By

  • Quantum attack method and system for detecting continuous variable quantum key distribution

    CN121261890A

  • Parameter optimization system and parameter optimization method of image signal processor

    CN122175761A