Methods, apparatus and storage media for determining hyperparameters in image signal processing

By constructing a proxy network model of the ISP model and jointly training it with the machine vision model, the hyperparameters of the ISP model are optimized, solving the problem that the ISP model cannot adapt to machine vision tasks and improving the accuracy of machine vision tasks.

CN116416482BActive Publication Date: 2025-10-28CAMBRICON TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111637882.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-10-28
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The traditional image signal processing (ISP) model cannot adapt to machine vision tasks, resulting in limited accuracy. This is because the ISP model aims to reproduce the visual effects of the human eye, while machine vision differs from human vision.

Method used

A proxy network model for the ISP model is constructed and jointly trained with the machine vision model. The hyperparameters of each processing module are optimized to make the ISP model suitable for machine vision tasks.

Benefits of technology

By constructing a proxy network model and jointly training a machine vision model, we obtain ISP model hyperparameters suitable for machine vision tasks, thereby improving the accuracy of machine vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416482B_ABST
    Figure CN116416482B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, and storage medium for determining hyperparameters in image signal processing. The method involves storing computer-executable instructions in a memory, and having at least one processor execute these instructions. The at least one processor then executes the hyperparameter determination method for image signal processing, obtaining hyperparameters suitable for an ISP model to perform machine vision tasks. This ensures that the ISP model's image processing results meet the requirements of machine vision tasks, thereby improving the accuracy of machine vision tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of image processing technology and artificial intelligence technology, and in particular to a method, apparatus and storage medium for determining hyperparameters in image signal processing. Background Technology

[0002] The goal of traditional ISP (Image Signal Processing) is to reproduce the real scene on the display or to make the image pleasing to the human eye.

[0003] With the development of machine vision technology, many products no longer need displays, such as robots, autonomous vehicles in industrial parks, and smart cameras. However, because machine vision differs from human vision, and traditional image processing units (ISPs) adjust parameters with the goal of replicating human vision, this limits their performance in machine vision tasks and affects their accuracy. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for determining hyperparameters of image signal processing, used to determine suitable ISP hyperparameters so that the ISP can be adapted to machine vision tasks.

[0005] In a first aspect, embodiments of this application provide a method for determining hyperparameters in image signal processing, including:

[0006] Construct and train a proxy network model for an image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model.

[0007] The proxy network model and the machine vision model are jointly trained, and the proxy sub-network model is optimized to obtain the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0008] Secondly, embodiments of this application provide a hyperparameter determination apparatus for image signal processing, comprising:

[0009] The training module is used to construct and train a proxy network model of the image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model.

[0010] The joint module is used to jointly train the proxy network model and the machine vision model, optimize the proxy sub-network model, and obtain the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0011] Thirdly, embodiments of this application provide a hyperparameter determination device for image signal processing, comprising: at least one processor and a memory;

[0012] The memory stores computer-executable instructions;

[0013] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in the first aspect.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by at least one processor, implements the method described in the first aspect.

[0015] The image signal processing hyperparameter determination method, apparatus, and storage medium provided in this application construct and train a proxy network model for the image signal processing (ISP) model. The proxy network model includes at least one proxy sub-network model that implements the functions of the processing modules in the ISP model, and the hyperparameters of the processing modules serve as input features of the corresponding proxy sub-network models. The proxy network model is jointly trained with a machine vision model, and the proxy sub-network models are optimized to obtain hyperparameters suitable for performing machine vision tasks. By constructing a proxy network model to act as an agent for the ISP model and jointly training it with the machine vision model, this application can obtain hyperparameters suitable for performing machine vision tasks, ensuring that the image processing results of the ISP model meet the requirements of machine vision tasks and improving the accuracy of machine vision tasks. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0017] Figure 1 This is a schematic diagram of a proxy network model for constructing an ISP model provided in one embodiment of this application;

[0018] Figure 2 A flowchart illustrating a hyperparameter determination method for image signal processing provided in one embodiment of this application;

[0019] Figure 3 A flowchart illustrating a hyperparameter determination method for image signal processing provided in another embodiment of this application;

[0020] Figure 4 A flowchart illustrating a hyperparameter determination method for image signal processing provided in another embodiment of this application;

[0021] Figure 5 This is a schematic diagram illustrating the separate training of a proxy sub-network model according to an embodiment of this application;

[0022] Figure 6 A flowchart illustrating a hyperparameter determination method for image signal processing provided in another embodiment of this application;

[0023] Figure 7 This is a schematic diagram illustrating the overall training of the proxy network model of the ISP model according to an embodiment of this application;

[0024] Figure 8 A flowchart illustrating a hyperparameter determination method for image signal processing provided in another embodiment of this application;

[0025] Figure 9 This is a schematic diagram illustrating the joint training of a proxy network model and a machine vision model, provided as an embodiment of this application.

[0026] Figure 10 This is a schematic diagram of the structure of a hyperparameter determination device for image signal processing provided in one embodiment of this application;

[0027] Figure 11 This is a schematic diagram of the structure of a hyperparameter determination device for image signal processing provided in another embodiment of this application;

[0028] Figure 12 This is a structural diagram of a board according to an embodiment of this application;

[0029] Figure 13 This is a structural diagram illustrating a combined processing apparatus according to an embodiment of this application;

[0030] Figure 14 This is a schematic diagram illustrating the internal structure of a single-core computing device according to an embodiment of this application.

[0031] Figure 15 This is a schematic diagram showing the internal structure of a multi-core computing device according to an embodiment of this application;

[0032] Figure 16 This is a schematic diagram illustrating the internal structure of a processor core according to an embodiment of this application.

[0033] The above drawings illustrate specific embodiments of the present disclosure, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the present disclosure in any way, but rather to illustrate the concepts of the present disclosure to those skilled in the art by reference to specific embodiments. Detailed Implementation

[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0035] To clearly understand the technical solution of this application, the solutions of the prior art will be described in detail first.

[0036] The goal of traditional image ISPs is to reproduce realistic scenes on a display or to make images pleasing to the human eye. With the development of machine vision technology, many products no longer require displays, such as robots, autonomous vehicles in industrial parks, and smart cameras. Because machine vision differs from human vision—for example, machine vision primarily aims to improve accuracy by adjusting image color difference and brightness, which differs from human vision in these aspects—traditional ISPs, which adjust parameters to reproduce human visual effects, limit their performance in machine vision tasks and affect their accuracy.

[0037] To address the aforementioned technical issues, this application constructs a proxy network model for the ISP and performs joint optimization with a machine vision model to determine the optimal values ​​of the target hyperparameters for each processing module of the ISP, making them suitable for machine vision tasks. Considering that the ISP includes multiple processing modules, each of which may have one or more hyperparameters, the complexity of hyperparameter sampling increases exponentially with the number of hyperparameters. If a proxy network model is used to proxy the entire ISP processing process, the proxy network model will face the problem of too many hyperparameters, making sampling and optimization difficult, resulting in poor proxy performance and poor joint optimization performance with the machine vision model. Therefore, this application constructs a proxy network model for each processing module of the ISP. In comparison, the proxy network model for each processing module involves fewer hyperparameters, facilitating independent sampling and network proxying, effectively improving the joint optimization performance with the machine vision model, and determining the optimal hyperparameters suitable for machine vision tasks.

[0038] Specifically, such as Figure 1As shown, this embodiment of the application constructs and trains a proxy network model for the ISP model. For at least one processing module in the ISP model, a proxy sub-network model capable of implementing the corresponding processing module's function is constructed, and the hyperparameters of the processing module serve as the input features of the corresponding proxy sub-network model. After training the proxy network model of the ISP model, the proxy network model is jointly trained with a machine vision model to optimize the proxy sub-network model, obtaining hyperparameters suitable for performing machine vision tasks. It should be noted that... Figure 1 The ISP model shown here is for illustrative purposes only and is not limited to the processing modules depicted.

[0039] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0040] Figure 2 This is a flowchart of a hyperparameter determination method for image signal processing provided in one embodiment of this application. The executing entity of this embodiment can be any electronic device. Figure 2 As shown, the hyperparameter determination method for image signal processing provided in this embodiment includes the following steps:

[0041] S201. Construct and train a proxy network model for the image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model.

[0042] In this embodiment, the ISP model can be implemented in hardware, such as an ISP chip or ISP board, or in software. The ISP model includes multiple processing modules, such as a white balance and gain processing module, a Bayer threshold noise reduction processing module, a color interpolation processing module, a color correction matrix processing module, a tone mapping processing module, a gamma correction processing module, a color space conversion processing module, a sharpening processing module, and a contrast processing module. Each processing module can use a corresponding algorithm to implement its respective processing function. Different processing modules have different hyperparameters. By adjusting the values ​​of each hyperparameter, the processing effect of each module can be adjusted. For example, the color correction matrix processing module uses a 3×3 color change matrix for color correction; the nine values ​​in the color change matrix are the hyperparameters of the color correction matrix processing module. Similarly, the gamma value of the gamma correction processing module is its hyperparameter. To make the ISP model suitable for machine vision tasks, the hyperparameters of each processing module need to be adjusted to be suitable for machine vision tasks.

[0043] Since the algorithms of each processing module in the ISP model are usually fixed, conventional methods for adjusting the hyperparameters of each module typically require trial and error by setting different values ​​for each hyperparameter based on the original algorithm. This is inefficient and requires human intervention. Therefore, in this embodiment, a proxy network model is created for the ISP model. The proxy network model includes at least one proxy sub-network model to implement the functions of the processing modules in the ISP model. The input parameters of each proxy sub-network model include the input image, and the hyperparameters of the processing modules are used as the input features of the corresponding proxy sub-network model. The output of each proxy sub-network model is the processed output image. For example, the proxy sub-network model corresponding to the color correction matrix processing module can implement the functions of the color correction matrix processing module through a machine learning model. Its input parameters include the input image and a set of 3×3 color change matrices, and the output is the color-corrected output image. The proxy sub-network model does not use the original algorithm of the color correction matrix processing module, but achieves the same color correction processing effect through machine learning.

[0044] In this embodiment, any model can be used to construct the proxy sub-network model corresponding to the processing module; no restrictions are imposed. Preferably, model information for the corresponding proxy sub-network model can be determined based on the characteristics of different processing models, including at least one of model type, model structure, model features, and model parameters, and then the proxy sub-network model is generated based on the model information. Processing modules of different complexities can use proxy sub-network models of different complexities; for example, more complex processing modules use more complex proxy sub-network models.

[0045] Optionally, the proxy sub-network corresponding to the tone mapping processing module is constructed using a UNet network model; and / or, the proxy sub-network corresponding to the gamma correction processing module is constructed using a multi-layer convolutional network model; and / or, the proxy sub-network corresponding to the Bayer threshold noise reduction processing module is constructed using a residual network model; and / or, the proxy sub-network corresponding to the sharpening processing module and / or the contrast processing module is constructed using a residual UNet network model in the Y channel.

[0046] Optionally, at least two adjacent processing modules with similar functions and / or computational methods can be merged to construct a single proxy sub-network model, thereby simplifying the structure of the proxy network model, for example... Figure 1 The sharpening and contrast processing modules are merged into a single proxy sub-network model.

[0047] Optionally, for processing modules without target hyperparameters, they can be added to the proxy network model in a differentiable manner. The ISP model may contain processing modules that do not require hyperparameter optimization, or may not contain any hyperparameter-optimized processing modules at all, for example... Figure 1 The white balance and gain processing module, the Bayer threshold noise reduction module, and other components are incorporated into the proxy network model in a differentiable manner. Differentiable processing is used to ensure that the proxy network model can achieve the complete processing of the ISP model. Differentiable methods include specific algorithms, formulas, and simple convolution processing, which will not be listed here.

[0048] Furthermore, for processing modules involving image statistical information, the image statistical information can be incorporated as an intrinsic parameter into the corresponding proxy sub-network of the processing module.

[0049] After constructing the proxy network model of the ISP model, the proxy network model can be trained to achieve the same or similar processing effect as the ISP model. In particular, the proxy network model can achieve the same or similar processing effect as the ISP model under different hyperparameters.

[0050] S202. Jointly train the proxy network model and the machine vision model, optimize the proxy sub-network model, and obtain the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0051] In this embodiment, the model parameters of the proxy network model are frozen, while the hyperparameters of the processing modules corresponding to the proxy sub-network models are used as the parameters that need to be tuned in the proxy network model. These hyperparameters are then jointly trained with the machine vision model to optimize the hyperparameters, resulting in hyperparameters suitable for performing machine vision tasks in the ISP model. Furthermore, the hyperparameters of the ISP model can be written into the ISP model itself; for example, for a hardware-implemented ISP model, the hyperparameters can be written into the corresponding registers.

[0052] The machine vision model in this embodiment can be any machine vision model such as a perception model, image segmentation model, or semantic segmentation model; no restrictions are imposed in this embodiment.

[0053] The hyperparameter determination method for image signal processing provided in this embodiment constructs and trains a proxy network model for the image signal processing (ISP) model. The proxy network model includes at least one proxy sub-network model that implements the functions of the processing modules in the ISP model, and the hyperparameters of the processing modules serve as input features for the corresponding proxy sub-network models. The proxy network model is jointly trained with a machine vision model, and the proxy sub-network models are optimized to obtain hyperparameters suitable for performing machine vision tasks. This embodiment, by constructing a proxy network model to act as an agent for the ISP model and jointly training it with the machine vision model, can obtain hyperparameters suitable for performing machine vision tasks, ensuring that the image processing results of the ISP model meet the requirements of machine vision tasks and improving the accuracy of machine vision tasks.

[0054] Based on any of the above embodiments, such as Figure 3 As shown in S201, the proxy network model for constructing and training the image signal processing (ISP) model can specifically include:

[0055] S301. Based on the different sampling values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to the different sampling values ​​of the target processing module, the proxy sub-network model corresponding to the target processing module is trained separately, wherein the target processing module is the processing module in any of the ISP models.

[0056] In this embodiment, since the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, each proxy sub-network model needs to be trained separately in order to enable each proxy sub-network model to implement the functions of the corresponding processing module.

[0057] When training a proxy sub-network model corresponding to a specific target processing module, it is necessary to first obtain training data. The training data includes different sampled values ​​of the hyperparameters of the target processing module, as well as the input and output images of the target processing module corresponding to different sampled values.

[0058] Optional, such as Figure 4 As shown, S301 may specifically include the following steps:

[0059] S3011. Sample each hyperparameter of the target processing module to obtain multiple different first target hyperparameter groups; wherein each first target hyperparameter group includes a sampled value of each hyperparameter of the target processing module.

[0060] In this embodiment, for example, suppose the target processing module has two hyperparameters, A and B. For any hyperparameter, its value range is obtained, and at least two sampled values ​​are taken within that range. For example, for hyperparameter A, sampled values ​​a1, a2, ... are taken within its value range; for hyperparameter B, sampled values ​​b1, b2, ... are taken within its value range. The different sampled values ​​of each hyperparameter of the target processing module are combined to obtain multiple different hyperparameter groups, denoted as the first target hyperparameter group. For example, hyperparameter A is a1 and hyperparameter B is b1, resulting in one first target hyperparameter group; similarly, hyperparameter A is a1 and hyperparameter B is b2, resulting in another first target hyperparameter group; and so on. Optionally, when taking at least two sampled values ​​within the value range of a hyperparameter, the Latin hypercube sampling method can be used to take at least two values ​​within the value range, resulting in more uniform sampling. Of course, other sampling methods can also be used.

[0061] S3012. Obtain the input image and output image corresponding to each of the first target hyperparameter groups by the target processing module, and use them as the input image and output image corresponding to the first target hyperparameter group.

[0062] In this embodiment, the hyperparameters of the target processing module can be set according to each first target hyperparameter group in the ISP model. When setting according to each first target hyperparameter group, one or more input images can be input into the target processing module. After processing by the target processing module, the corresponding output image is obtained. Then, the input image and output image at this time are used as the input image and output image corresponding to the first target hyperparameter group. Through the above process, input images and output images corresponding to different first target hyperparameter groups can be obtained.

[0063] Optionally, hyperparameters can be set for multiple processing modules at once, and one or more images can be input into the ISP model. Each processing module processes the images sequentially and obtains the input and output images of each processing module as the input and output images corresponding to the current first target hyperparameter group of that processing module.

[0064] S3013. Based on the first target hyperparameter group, and the input image and output image corresponding to the first target hyperparameter group, train the proxy sub-network model corresponding to the target processing module separately.

[0065] In this embodiment, each first target hyperparameter group, as well as the input and output images corresponding to each first target hyperparameter group, are used as training data. Based on the first target hyperparameter group and the input and output images corresponding to the first target hyperparameter group, the proxy sub-network model corresponding to the target processing module is trained separately.

[0066] When training the proxy sub-network model corresponding to the target processing module, such as Figure 5 As shown, the input image corresponding to the first target hyperparameter group is used as the image input of the surrogate sub-network model, and the corresponding first target hyperparameter group is also used as auxiliary features to input into the surrogate sub-network model. The output image is the image processed by the model. Based on the image processed by the model and the output image corresponding to the first target hyperparameter group, the loss value is calculated. Backpropagation is performed based on the loss value to optimize the model parameters of the surrogate sub-network model. The loss value can be L1 or L2 loss.

[0067] S302. Fine-tune each of the individually trained agent sub-network models as a whole to obtain the agent network model.

[0068] In this embodiment, after training each agent sub-network model, in order to further improve the model accuracy, the individually trained agent sub-network models can be fine-tuned as a whole model, that is, the agent network model of the ISP model is fine-tuned as a whole.

[0069] Specifically, such as Figure 6 As shown, S302 may specifically include the following steps:

[0070] S3021. The hyperparameters of each of the processing modules are sampled to obtain multiple different second target hyperparameter groups; each second target hyperparameter group includes a sampled value of each hyperparameter of each of the processing modules.

[0071] In this embodiment, the hyperparameters of all processing modules of the ISP model are sampled separately, and multiple different sets of second target hyperparameters are obtained by combining them. For example, the hyperparameter A of the first processing module is a1, the hyperparameter B is b1, the hyperparameter C of the second processing module is c1, the hyperparameter D is d1, and so on, to obtain a set of second target hyperparameters. Similarly, the hyperparameter A of the first processing module is a1, the hyperparameter B is b2, the hyperparameter C of the second processing module is c2, the hyperparameter D is d1, and so on, to obtain another set of second target hyperparameters.

[0072] S3022. Obtain the input image and output image of the ISP model when using each of the second target hyperparameter groups, and use them as the input image and output image of the second target hyperparameter group.

[0073] In this embodiment, the hyperparameters of each processing module can be set according to each second target hyperparameter group in the ISP model. When setting according to each second target hyperparameter group, one or more input images (raw images in RAW format) are input into the ISP model. After being processed by each processing module in sequence, the final output image is obtained. The input image and output image at this time are used as the input image and output image corresponding to the second target hyperparameter group. Through the above process, the input image and output image corresponding to different second target hyperparameter groups can be obtained.

[0074] S3023. Based on the second target hyperparameter group and the corresponding input and output images of the second target hyperparameter group, the model composed of each agent sub-network model is fine-tuned to obtain the agent network model.

[0075] In this embodiment, each second target hyperparameter group, as well as the input and output images corresponding to each second target hyperparameter group, are used as training data. Based on the second target hyperparameter groups and the input and output images corresponding to the second target hyperparameter groups, the proxy network model of the ISP model is trained as a whole.

[0076] When training the proxy network model of the ISP model as a whole, such as Figure 7As shown, the input image corresponding to the second target hyperparameter group is used as the image input of the surrogate sub-network model. The corresponding second target hyperparameter group is also used as auxiliary features and input into the corresponding surrogate sub-network model. The input image is processed by each surrogate sub-network model in sequence, and the finally output image is the image processed by the model. Based on the image processed by the model and the output image corresponding to the second target hyperparameter group, the loss value is calculated. Backpropagation is performed based on the loss value to optimize the model parameters of the surrogate sub-network model. Regularization can be performed by weighting the loss of the surrogate sub-network model of each processing module to obtain the regularization term. Then, backpropagation is performed based on the loss value and the regularization term to optimize the model parameters of the surrogate sub-model of each processing module.

[0077] Based on any of the above embodiments, such as Figure 8 As shown in S202, the joint training of the proxy network model and the machine vision model, and the optimization of the proxy network model to obtain the hyperparameters of the ISP model suitable for performing machine vision tasks, may specifically include:

[0078] S2021. Obtain training data for joint training, wherein the training data includes the original input image and the annotation data of the original input image.

[0079] In this embodiment, during joint training, the input image is processed by the proxy network model of the ISP model and then input into the machine vision model to perform the machine vision task, obtaining the final machine vision task result. Therefore, it is necessary to obtain training data for joint training. This training data includes the input image and the annotation data of the original input image. The input image is the input image of the ISP model, that is, the original input image (the original image in RAW format). The annotation data of the original input image is the annotation data according to the machine vision task. For example, if the machine vision task is to identify the position of a face in an image, then it is necessary to obtain the position coordinates of the face included in the original input image as the annotation data of the original input image. Of course, the annotation data will be different for different machine vision tasks, which will not be elaborated here.

[0080] S2022. Input the original input image into the proxy network model, input the output of the proxy network model into the machine vision model, and obtain the machine vision processing result.

[0081] In this embodiment, the model parameters of the proxy network model are frozen during joint training, while the hyperparameters of the processing module corresponding to the proxy sub-network model are used as the parameters that the proxy network model needs to be tuned. That is, at this time, it is not necessary to input any second target hyperparameter set.

[0082] When using certain training data for joint training, such as Figure 9As shown, the original input image is input into the proxy network model for processing, and then the output of the proxy network model is input into the machine vision model to perform machine vision tasks and obtain machine vision processing results.

[0083] S2023. Based on the machine vision processing results and the labeled data of the original input image, obtain the loss value, and optimize the hyperparameters of the processing modules corresponding to each of the proxy sub-network models based on the loss value to obtain the optimal hyperparameters of the processing modules corresponding to each of the proxy sub-network models.

[0084] In this embodiment, as Figure 9 As shown, based on the machine vision processing results and the labeled data of the original input images in the training data, the loss value is obtained, and the hyperparameters of the corresponding processing modules of each proxy sub-network model are optimized by backpropagation based on the loss value, and finally the optimal hyperparameters of the corresponding processing modules of each proxy sub-network model are obtained.

[0085] S2024. Determine the optimal hyperparameters of the processing modules corresponding to each of the agent sub-network models as the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0086] In this embodiment, after joint training is completed, the optimal hyperparameters of the processing modules corresponding to each agent sub-network model are determined as the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0087] Figure 10 This is a schematic diagram of the structure of a hyperparameter determination device for image signal processing provided in one embodiment of this application, as shown below. Figure 10 As shown, the hyperparameter determination device 400 for image signal processing provided in this embodiment includes a training module 401 and a joint module 402.

[0088] Training module 401 is used to construct and train a proxy network model of the image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the function of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model.

[0089] The joint module 402 is used to jointly train the proxy network model and the machine vision model, optimize the proxy sub-network model, and obtain the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0090] In one or more embodiments of this application, the training module 401, when constructing and training the proxy network model of the image signal processing (ISP) model, is used for:

[0091] Based on the different sampling values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to the different sampling values ​​of the target processing module, the proxy sub-network model corresponding to the target processing module is trained separately, and the target processing module is the processing module in any of the ISP models;

[0092] The individual agent sub-network models, trained separately, are fine-tuned as a whole to obtain the agent network model.

[0093] In one or more embodiments of this application, when the training module 401 trains the proxy sub-network model corresponding to the target processing module separately based on different sampled values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to the different sampled values ​​of the target processing module, it is used to:

[0094] Each hyperparameter of the target processing module is sampled to obtain multiple different first target hyperparameter groups; each first target hyperparameter group includes a sampled value of each hyperparameter of the target processing module.

[0095] The input image and output image of the target processing module when using each of the first target hyperparameter groups are obtained and used as the input image and output image of the first target hyperparameter group.

[0096] Based on the first target hyperparameter set, and the input and output images corresponding to the first target hyperparameter set, the proxy sub-network model corresponding to the target processing module is trained separately.

[0097] In one or more embodiments of this application, when the training module 401 fine-tunes each individually trained agent sub-network model as a whole to obtain the agent network model, it is used to:

[0098] The hyperparameters of each of the processing modules are sampled to obtain multiple different sets of second target hyperparameters; each set of second target hyperparameters includes a sampled value of each hyperparameter of each of the processing modules.

[0099] The input and output images of the ISP model when each of the second target hyperparameter groups are adopted are obtained as the input and output images of the second target hyperparameter groups.

[0100] Based on the second target hyperparameter set, and the corresponding input and output images of the second target hyperparameter set, the model composed of each agent sub-network model is fine-tuned to obtain the agent network model.

[0101] In one or more embodiments of this application, when the joint module 402 jointly trains the proxy network model and the machine vision model, optimizes the proxy network model, and obtains hyperparameters of an ISP model suitable for performing machine vision tasks, it is used to:

[0102] Acquire training data for joint training, the training data including the original input image and the annotation data of the original input image;

[0103] The original input image is input into the proxy network model, and the output of the proxy network model is input into the machine vision model to obtain the machine vision processing result.

[0104] Based on the machine vision processing results and the labeled data of the original input image, a loss value is obtained. The hyperparameters of the processing modules corresponding to each of the proxy sub-network models are optimized based on the loss value to obtain the optimal hyperparameters of the processing modules corresponding to each of the proxy sub-network models.

[0105] The optimal hyperparameters of the processing modules corresponding to each of the aforementioned agent sub-network models are determined as the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0106] In one or more embodiments of this application, when the training module 401 samples each hyperparameter of the target processing module to obtain multiple different first target hyperparameter sets, it is used to:

[0107] For any hyperparameter of the target processing module, obtain the value range of the hyperparameter, and take at least two sampled values ​​within the value range;

[0108] By combining the different sampled values ​​of each hyperparameter of the target processing module, multiple different first target hyperparameter groups are obtained.

[0109] In one or more embodiments of this application, when the training module 401 takes at least two sampled values ​​within the value range, it is used to:

[0110] The Latin hypercube sampling method is used to select at least two values ​​within the range.

[0111] In one or more embodiments of this application, at least two adjacent processing modules with similar functions and / or computational methods are merged to form a proxy subnetwork model.

[0112] In one or more embodiments of this application, for processing modules without target hyperparameters, a differentiable approach is adopted to incorporate them into the proxy network model.

[0113] In one or more embodiments of this application, the proxy sub-network corresponding to the tone mapping processing module is constructed using the UNet network model; and / or

[0114] The proxy subnetwork corresponding to the gamma correction processing module is constructed using a multi-layer convolutional layer network model; and / or

[0115] The proxy subnetwork corresponding to the Bayer threshold noise reduction module is constructed using a residual network model; and / or

[0116] The proxy subnetworks corresponding to the sharpening and / or contrast processing modules are constructed using a residual UNet network model in the Y channel.

[0117] The hyperparameter determination device for image signal processing provided in this embodiment can perform... Figure 2-4 , Figure 6 and Figure 8 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.

[0118] Figure 11 This is a schematic diagram of the structure of a hyperparameter determination device for image signal processing provided in another embodiment of this application, as shown below. Figure 11 As shown, the hyperparameter determination device 50 for image signal processing provided in this application embodiment includes: at least one processor 51 and a memory 52;

[0119] Memory 52 stores instructions executed by the computer;

[0120] At least one processor 51 executes computer execution instructions stored in memory 52, causing at least one processor to execute... Figure 2-4 , Figure 6 and Figure 8 Any embodiment provides a method for determining hyperparameters in image signal processing.

[0121] In one possible implementation, a computer-readable storage medium is also disclosed, in which a computer program is stored, which, when executed by at least one processor, implements... Figure 2-4 , Figure 6 and Figure 8 Any embodiment provides a method for determining hyperparameters in image signal processing.

[0122] In one possible implementation, a board is also disclosed, which can be a device-side board. Figure 12 This diagram illustrates the structure of a board 60 according to an embodiment of this application. Figure 12As shown, board 60 includes chip 601, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 60 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities.

[0123] Chip 601 is connected to external device 603 via external interface device 602. External device 603 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 603 to chip 601 via external interface device 602. The calculation results from chip 601 can be transmitted back to external device 603 via external interface device 602. Depending on the application scenario, external interface device 602 may have different interface forms, such as a PCIe interface.

[0124] The board 60 also includes a storage device 604 for storing data, which includes one or more memory cells 605. The storage device 604 is connected to and transmits data with the controller 606 and the chip 601 via a bus. The controller 606 in the board 60 is configured to regulate the state of the chip 601. Therefore, in one application scenario, the controller 606 may include a microcontroller (MCU).

[0125] In one possible implementation, a combined processing device is also provided. Figure 13 This is a structural diagram illustrating the combined processing device in chip 601 of this embodiment. (As shown...) Figure 13 As shown, the combined processing device 70 includes a computing device 701, an interface device 702, a processing device 703, and a storage device 704.

[0126] The computing device 701 is configured to perform user-specified operations. It is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 703 through the interface device 702 to jointly complete the user-specified operations.

[0127] Interface device 702 is used to transmit data and control commands between computing device 701 and processing device 703. For example, computing device 701 can obtain input data from processing device 703 via interface device 702 and write it to on-chip storage device of computing device 701. Further, computing device 701 can obtain control commands from processing device 703 via interface device 702 and write them to on-chip control cache of computing device 701. Alternatively or optionally, interface device 702 can also read data from storage device of computing device 701 and transmit it to processing device 703.

[0128] The processing device 703, as a general-purpose processing device, performs basic controls including but not limited to data transfer and starting / stopping the computing device 701. Depending on the implementation, the processing device 703 can be one or more types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), or other general-purpose and / or special-purpose processors. These processors include, but are not limited to, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, the computing device 701 of this application can be considered as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 701 and the processing device 703 are considered together, they are regarded as forming a heterogeneous multi-core structure.

[0129] The storage device 704 is used to store the data to be processed. It may be DRAM 704, which is DDR memory, typically 16G or larger in size, and is used to store the data of the computing device 701 and / or the processing device 703.

[0130] Figure 14 The diagram shows the internal structure of a single-core computing device 701. The single-core computing device 801 is used to process input data from computer vision, speech, natural language processing, data mining, etc. The single-core computing device 801 includes three main modules: a control module 81, a processing module 82, and a storage module 83.

[0131] The control module 81 coordinates and controls the operation of the computation module 82 and the storage module 83 to complete the deep learning task. It includes an instruction fetch unit (IFU) 811 and an instruction decode unit (IDU) 812. The instruction fetch unit 811 fetches instructions from the processing device 1203, and the instruction decode unit 812 decodes the fetched instructions and sends the decoding result as control information to the computation module 82 and the storage module 83.

[0132] The computation module 82 includes a vector operation unit 821 and a matrix operation unit 822. The vector operation unit 821 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformations; the matrix operation unit 822 is responsible for the core computations of deep learning algorithms, namely matrix multiplication and convolution.

[0133] Storage module 83 is used to store or move relevant data, including neuron RAM (NRAM) 831, weight RAM (WRAM) 832, and direct memory access (DMA) module 833. NRAM 831 is used to store input neurons, output neurons, and intermediate results after computation; WRAM 832 is used to store the convolution kernels of the deep learning network, i.e., the weights; DMA 833 is connected to DRAM 704 through bus 84 and is responsible for data transfer between the single-core computing device 801 and DRAM 704.

[0134] Figure 15 A schematic diagram of the internal structure of the computing device 701 as a multi-core is shown. The multi-core computing device 901 adopts a hierarchical structure design. As a system-on-a-chip, the multi-core computing device 901 includes at least one cluster, and each cluster includes multiple processor cores. In other words, the multi-core computing device 901 is constructed in a hierarchical structure of system-on-a-chip, cluster, and processor core.

[0135] From the perspective of system-on-a-chip hierarchy, such as Figure 15 As shown, the multi-core computing device 901 includes an external storage controller 901, a peripheral communication module 902, an on-chip interconnect module 903, a synchronization module 904, and multiple clusters 905.

[0136] There can be multiple external storage controllers 901; two are shown as an example in the figure. These controllers are used to respond to access requests from the processor core to access external storage devices, such as… Figure 13The DRAM 704 in the chip allows data to be read from or written to external memory. The peripheral communication module 902 receives control signals from the processing unit 703 via the interface device 702, initiating the computing unit 701 to execute tasks. The on-chip interconnect module 903 connects the external memory controller 901, the peripheral communication module 902, and multiple clusters 905 to transmit data and control signals between modules. The synchronization module 904 is a global barrier controller (GBC) used to coordinate the working progress of each cluster and ensure information synchronization. The multiple clusters 905 are the computing cores of the multi-core computing device 901; four are exemplarily shown in the figure, forming a structure like... Figure 1 The four quadrants are shown. With hardware advancements, the multi-core computing device 901 of this application can also include clusters 905 with 8, 16, 64, or even more cores. Clusters 905 are used to efficiently execute deep learning algorithms.

[0137] From the perspective of cluster hierarchy, such as Figure 15 As shown, each cluster 905 includes multiple processor cores (IPU cores) 906 and one memory core (MEM core) 907. For example, each cluster 905 includes four processor cores and one memory, which can be DRAM 704. Each processor core is equivalent to... Figure 1 One of the processing units, each memory is equivalent to Figure 1 One of the storage units.

[0138] Four processor cores 906 are shown in the figure as an example; this application does not limit the number of processor cores 906. Its internal architecture is as follows: Figure 16 As shown. Each processor core 906 is similar to Figure 14The single-core computing device 801 also includes three main modules: a control module 1001, an arithmetic module 1002, and a storage module 1003. The functions and structures of the control module 1001, arithmetic module 1002, and storage module 1003 are largely the same as those of the control module 81, arithmetic module 82, and storage module 83. The control module 1001 includes an instruction fetch unit 10011 and an instruction decode unit 10012. The arithmetic module 1002 includes a vector operation unit 10021 and a matrix operation unit 10022. Further details are omitted. It should be noted that the storage module 1003 includes an input / output direct memory access (IODMA) module 10033 and a move direct memory access (MVDMA) module 10034. IODMA10033 controls memory access of NRAM 10031 / WRAM10032 and DRAM 704 via broadcast bus 909; MVDMA 10034 is used to control memory access of NRAM 10031 / WRAM 10032 and SRAM 908.

[0139] Back Figure 13 The storage core 907 is primarily used for storage and communication, namely storing shared data or intermediate results among processor cores 906, and performing communication between cluster 905 and DRAM 704, communication between clusters 905, and communication between processor cores 906. In other embodiments, the storage core 907 has scalar operation capabilities and is used to perform scalar operations.

[0140] The storage core 907 includes an SRAM 908, a broadcast bus 909, a cluster direct memory access (CDMA) module 910, and a global direct memory access (GDMA) module 911. The SRAM 908 acts as a high-performance data relay station. Data multiplexed between different processor cores 906 within the same cluster 905 does not need to be obtained from the DRAM 704 by each processor core 906. Instead, it is relayed between processor cores 906 via the SRAM 908. The storage core 907 only needs to quickly distribute the multiplexed data from the SRAM 908 to multiple processor cores 906 to improve inter-core communication efficiency and greatly reduce on-chip and off-chip I / O access.

[0141] Broadcast bus 909, CDMA 910, and GDMA 911 are used to perform communication between processor cores 906, communication between clusters 905, and data transfer between cluster 905 and DRAM 704, respectively. These will be explained separately below.

[0142] The broadcast bus 909 is used to complete high-speed communication between the processor cores 906 within the cluster 905. In this embodiment, the broadcast bus 909 supports inter-core communication methods including unicast, multicast, and broadcast. Unicast refers to point-to-point (e.g., data transmission from one processor core to another) data transmission. Multicast is a communication method that transmits data from SRAM 908 to several specific processor cores 906. Broadcast is a communication method that transmits data from SRAM 908 to all processor cores 906, and is a special case of multicast.

[0143] CDMA 910 is used to control SRAM 908 access between different clusters 905 within the same computing device 701.

[0144] The GDMA 911 works in conjunction with the external memory controller 901 to control memory access from the SRAM 908 to the DRAM 704 in the cluster 905, or to read data from the DRAM 704 into the SRAM 908. As mentioned above, communication between the DRAM 704 and the NRAM 10031 or WRAM 10032 can be achieved through two channels. The first channel is a direct connection between the DRAM 704 and the NRAM 10031 or WRAM 10032 via the IODAM 10033; the second channel involves first transmitting data between the DRAM 704 and SRAM 908 via the GDMA 911, and then transmitting data between the SRAM 908 and the NRAM 10031 or WRAM 10032 via the MVDMA 10034. Although the second channel appears to require more components and has a longer data flow, in some embodiments, the bandwidth of the second channel is actually much greater than that of the first channel. Therefore, communication between DRAM 704 and NRAM 10031 or WRAM 10032 may be more efficient through the second channel. Embodiments of this application may select the data transmission channel based on their hardware capabilities.

[0145] In other embodiments, the functions of GDMA 911 and IODMA 10033 can be integrated into the same component. For ease of description, this application treats GDMA 911 and IODMA 10033 as different components. For those skilled in the art, as long as the functions implemented and the technical effects achieved are similar to those of this application, they fall within the protection scope of this application. Furthermore, the functions of GDMA 911, IODMA 10033, CDMA 910, and MVDMA 10034 can also be implemented by the same component.

[0146] The foregoing may be better understood in view of the following clauses:

[0147] Clause 1. A method for determining hyperparameters in image signal processing, comprising:

[0148] Construct and train a proxy network model for an image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model.

[0149] The proxy network model and the machine vision model are jointly trained, and the proxy sub-network model is optimized to obtain the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0150] Clause 2. The proxy network model for constructing and training the image signal processing ISP model according to the method described in Clause 1 includes:

[0151] Based on the different sampling values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to the different sampling values ​​of the target processing module, the proxy sub-network model corresponding to the target processing module is trained separately, and the target processing module is the processing module in any of the ISP models;

[0152] The individual agent sub-network models, trained separately, are fine-tuned as a whole to obtain the agent network model.

[0153] Clause 3. According to the method described in Clause 2, the step of separately training the proxy sub-network model corresponding to the target processing module based on different sampled values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to the different sampled values ​​of the target processing module, includes:

[0154] Each hyperparameter of the target processing module is sampled to obtain multiple different first target hyperparameter groups; each first target hyperparameter group includes a sampled value of each hyperparameter of the target processing module.

[0155] The input image and output image of the target processing module when using each of the first target hyperparameter groups are obtained and used as the input image and output image of the first target hyperparameter group.

[0156] Based on the first target hyperparameter set, and the input and output images corresponding to the first target hyperparameter set, the proxy sub-network model corresponding to the target processing module is trained separately.

[0157] Clause 4. According to the method described in Clause 2, the step of fine-tuning each individually trained agent sub-network model as a whole to obtain the agent network model includes:

[0158] The hyperparameters of each of the processing modules are sampled to obtain multiple different sets of second target hyperparameters; each set of second target hyperparameters includes a sampled value of each hyperparameter of each of the processing modules.

[0159] The input and output images of the ISP model when each of the second target hyperparameter groups are adopted are obtained as the input and output images of the second target hyperparameter groups.

[0160] Based on the second target hyperparameter set, and the corresponding input and output images of the second target hyperparameter set, the model composed of each agent sub-network model is fine-tuned to obtain the agent network model.

[0161] Clause 5. According to the method described in Clause 4, the joint training of the proxy network model and the machine vision model, and the optimization of the proxy network model to obtain hyperparameters suitable for performing machine vision tasks, includes:

[0162] Acquire training data for joint training, the training data including the original input image and the annotation data of the original input image;

[0163] The original input image is input into the proxy network model, and the output of the proxy network model is input into the machine vision model to obtain the machine vision processing result.

[0164] Based on the machine vision processing results and the labeled data of the original input image, a loss value is obtained. The hyperparameters of the processing modules corresponding to each of the proxy sub-network models are optimized based on the loss value to obtain the optimal hyperparameters of the processing modules corresponding to each of the proxy sub-network models.

[0165] The optimal hyperparameters of the processing modules corresponding to each of the aforementioned agent sub-network models are determined as the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0166] Clause 6. According to the method described in Clause 3, the sampling of each hyperparameter of the target processing module to obtain multiple different first target hyperparameter groups includes:

[0167] For any hyperparameter of the target processing module, obtain the value range of the hyperparameter, and take at least two sampled values ​​within the value range;

[0168] By combining the different sampled values ​​of each hyperparameter of the target processing module, multiple different first target hyperparameter groups are obtained.

[0169] Clause 7. According to the method described in Clause 6, taking at least two sampled values ​​within the value range includes:

[0170] The Latin hypercube sampling method is used to select at least two values ​​within the range.

[0171] Clause 8. According to the method described in Clause 1, at least two adjacent processing modules with similar functions and / or computational methods shall be merged to form a proxy subnetwork model.

[0172] Clause 9. According to the method described in Clause 1, for processing modules without target hyperparameters, a differentiable approach shall be adopted to incorporate them into the agent network model.

[0173] Clause 10. The method described in Clause 1,

[0174] The proxy sub-network corresponding to the tone mapping processing module is constructed using the UNet network model; and / or

[0175] The proxy subnetwork corresponding to the gamma correction processing module is constructed using a multi-layer convolutional layer network model; and / or

[0176] The proxy subnetwork corresponding to the Bayer threshold noise reduction module is constructed using a residual network model; and / or

[0177] The proxy subnetworks corresponding to the sharpening and / or contrast processing modules are constructed using a residual UNet network model in the Y channel.

[0178] Clause 11. A hyperparameter determination apparatus for image signal processing, comprising:

[0179] A training module is used to construct and train a proxy network model of an image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model.

[0180] The joint module is used to jointly train the proxy network model and the machine vision model, optimize the proxy sub-network model, and obtain the hyperparameters of the ISP model suitable for performing machine vision tasks.

[0181] Clause 12. A hyperparameter determination apparatus for image signal processing, comprising: at least one processor and a memory;

[0182] The memory stores computer-executed instructions;

[0183] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any one of clauses 1-10.

[0184] Clause 13. A computer-readable storage medium storing a computer program that, when executed by at least one processor, implements the method as described in any one of Clauses 1-10.

[0185] Clause 14. A computer program product comprising a computer program that, when executed by at least one processor, implements the method as described in any one of Clauses 1-10.

[0186] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0187] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0188] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0189] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0190] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, an AI processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, storage units can be any suitable magnetic or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0191] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0192] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

Claims

1. A method for determining hyperparameters in image signal processing, characterized in that, include: Construct and train a proxy network model for an image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model. The proxy network model and the machine vision model are jointly trained, and the proxy sub-network model is optimized to obtain the hyperparameters of the ISP model suitable for performing machine vision tasks. The proxy network model for constructing and training the image signal processing (ISP) model includes: Based on the different sampling values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to the different sampling values ​​of the target processing module, the proxy sub-network model corresponding to the target processing module is trained separately, and the target processing module is the processing module in any of the ISP models; The individual agent sub-network models, trained separately, are fine-tuned as a whole to obtain the agent network model.

2. The method according to claim 1, characterized in that, The step of training the proxy sub-network model corresponding to the target processing module separately based on different sampled values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to different sampled values ​​of the target processing module, includes: Each hyperparameter of the target processing module is sampled to obtain multiple different first target hyperparameter groups; each first target hyperparameter group includes a sampled value of each hyperparameter of the target processing module. The input image and output image of the target processing module when using each of the first target hyperparameter groups are obtained and used as the input image and output image of the first target hyperparameter group. Based on the first target hyperparameter set, and the input and output images corresponding to the first target hyperparameter set, the proxy sub-network model corresponding to the target processing module is trained separately.

3. The method according to claim 1, characterized in that, The process of fine-tuning the individually trained agent sub-network models as a whole to obtain the agent network model includes: The hyperparameters of each of the processing modules are sampled to obtain multiple different sets of second target hyperparameters; each set of second target hyperparameters includes a sampled value of each hyperparameter of each of the processing modules. The input and output images of the ISP model when each of the second target hyperparameter groups are adopted are obtained as the input and output images of the second target hyperparameter groups. Based on the second target hyperparameter set, and the corresponding input and output images of the second target hyperparameter set, the model composed of each agent sub-network model is fine-tuned to obtain the agent network model.

4. The method according to claim 3, characterized in that The step of jointly training the proxy network model and the machine vision model, and optimizing the proxy network model to obtain hyperparameters of the ISP model suitable for performing machine vision tasks, includes: Acquire training data for joint training, the training data including the original input image and the annotation data of the original input image; The original input image is input into the proxy network model, and the output of the proxy network model is input into the machine vision model to obtain the machine vision processing result. Based on the machine vision processing results and the labeled data of the original input image, a loss value is obtained. The hyperparameters of the processing modules corresponding to each of the proxy sub-network models are optimized based on the loss value to obtain the optimal hyperparameters of the processing modules corresponding to each of the proxy sub-network models. The optimal hyperparameters of the processing modules corresponding to each of the aforementioned agent sub-network models are determined as the hyperparameters of the ISP model suitable for performing machine vision tasks.

5. The method according to claim 2, characterized in that, The sampling of each hyperparameter of the target processing module yields multiple different sets of first target hyperparameters, including: For any hyperparameter of the target processing module, obtain the value range of the hyperparameter, and take at least two sampled values ​​within the value range; By combining the different sampled values ​​of each hyperparameter of the target processing module, multiple different first target hyperparameter groups are obtained.

6. The method according to claim 5, characterized in that, Taking at least two sampled values ​​within the value range includes: The Latin hypercube sampling method is used to select at least two values ​​within the range.

7. The method according to claim 1, characterized in that, At least two adjacent processing modules with similar functions and / or computational methods are merged to form a proxy subnetwork model.

8. The method according to claim 1, characterized in that For processing modules without target hyperparameters, a differentiable approach is used to incorporate them into the proxy network model.

9. The method according to claim 1, characterized in that, The proxy sub-network corresponding to the tone mapping processing module is constructed using the UNet network model; and / or The proxy subnetwork corresponding to the gamma correction processing module is constructed using a multi-layer convolutional layer network model; and / or The proxy subnetwork corresponding to the Bayer threshold noise reduction module is constructed using a residual network model; and / or The proxy subnetworks corresponding to the sharpening and / or contrast processing modules are constructed using a residual UNet network model in the Y channel.

10. A hyperparameter determination device for image signal processing, characterized in that, include: A training module is used to construct and train a proxy network model of an image signal processing (ISP) model, wherein the proxy network model includes at least one proxy sub-network model for implementing the functions of the processing module in the ISP model, and the hyperparameters of the processing module are used as input features of the corresponding proxy sub-network model. The joint module is used to jointly train the proxy network model and the machine vision model, optimize the proxy sub-network model, and obtain the hyperparameters of the ISP model suitable for performing machine vision tasks. The training module is specifically used for: Based on the different sampling values ​​of the hyperparameters of the target processing module, and the input and output images corresponding to the different sampling values ​​of the target processing module, the proxy sub-network model corresponding to the target processing module is trained separately, and the target processing module is the processing module in any of the ISP models; The individual agent sub-network models, trained separately, are fine-tuned as a whole to obtain the agent network model.

11. A hyperparameter determination device for image signal processing, characterized in that, include: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by at least one processor, implements the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Method and system for joint optimization of ISP and vision tasks, medium and electronic device

    CN113628124A