Information processing apparatus, information processing method, program, imaging device, and image processing method

By training a neural network with hardware constraints, optimized parameters are generated for image and inference processing units, enhancing accuracy and reducing hardware modification costs in image processing systems.

JP2025139298APending Publication Date: 2025-09-26SONY SEMICON SOLUTIONS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024038149
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing image processing methods using AI models are not optimized for hardware constraints, leading to inaccurate inference processing and potential hardware design changes, which increase costs.

Method used

A neural network architecture is trained with hardware constraints to generate optimized parameters for image processing units and inference processing units, allowing direct application to specific hardware while maintaining accuracy and flexibility.

Benefits of technology

This approach enables high-speed, accurate inference processing by aligning image processing parameters with hardware constraints, reducing the need for costly hardware modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025139298000001_ABST
    Figure 2025139298000001_ABST
Patent Text Reader

Abstract

To improve accuracy of inference processing using an AI model.SOLUTION: An information processing apparatus comprises a model generation unit configured to generate a learned model by training a neural network architecture to which a constraint on hardware on which the learned model is deployed is applied as a constraint condition.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to the technical fields of an information processing device, an information processing method, a program, an imaging device, and an image processing method that use machine learning to optimize image processing. [Background technology]

[0002] There are an increasing number of cases where problems are solved by performing inference using trained AI models. This is also the case in the field of image processing. For example, consider a case where a certain image data is used as input data and a certain inference process is performed. In this case, an AI model is used to perform the specified inference process on the subject contained in the image data and obtain the inference result.

[0003] Input image data for an AI model is obtained by performing predetermined image processing using an ISP (Image Signal Processor) installed in an imaging device such as a camera. However, image processing by an ISP often applies parameters to make the image easier for humans to see. Therefore, image data obtained by image processing using such parameters is not necessarily appropriate as input image data for an AI model.

[0004] In other words, in order to perform highly accurate inference processing using an AI model, it is important to optimize the parameters used for image processing in the ISP.

[0005] AI models can also be used for tuning ISP parameters to optimize the input image data for the AI ​​model.

[0006] For example, it is possible to adopt a configuration in which an upstream AI model is used to optimize image processing parameters in an image processing unit such as an ISP, and a downstream AI model is used to perform highly accurate inference about the subject contained in the input image data.

[0007] In the following Non-Patent Document 1, for example, an AI model is used to optimize parameters for various image processing operations performed in an image processing unit and to perform inference using image data obtained after image processing. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] Haina Qin and 6 others, "Attention-aware Learning for Hyperparameter Prediction in Image Processing Pipelines," European Conference on Computer Vision (ECCV), 2022, [Retrieved January 30, 2024], Internet<URL:https: / / www.ecva.net / papers / eccv_2022 / papers_ECCV / papers / 136790265.pdf> Summary of the Invention [Problem to be solved by the invention]

[0009] Non-Patent Document 1 proposes an AI model that performs parameter optimization and inference processing at the same time. However, there is still room to further improve the accuracy of inference processing using AI models.

[0010] This technology was developed in light of these problems, and aims to improve the accuracy of inference processing using AI models. [Means for solving the problem]

[0011] The information processing device of the present technology includes a model generation unit that generates the trained model by training a neural network architecture to which constraints on the hardware on which the trained model is deployed are assigned as constraint conditions. A trained model may be generated by training a neural network architecture and used for a specific process. In this case, the trained model with optimized parameters can be used to perform highly accurate processing. Incidentally, processing using a trained model can sometimes be achieved by applying the parameters calculated during the generation process of the trained model to relatively inflexible hardware, instead of using relatively flexible software. This allows for high speed processing at the expense of flexibility. However, in many cases, the parameters calculated during the generation process of a trained model cannot be applied to hardware as is, and hardware design changes are required, which leads to increased costs. In this configuration, training is performed while incorporating the constraints of the hardware to which the trained model is to be deployed. The parameters obtained through this training can be applied directly to the target hardware. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a block diagram showing a schematic configuration of an information processing system according to an embodiment of the present invention. [Figure 2] FIG. 1 illustrates an example of the configuration of an imaging device. [Figure 3] FIG. 1 is a block diagram showing an example of the configuration of a computer device such as a server device. [Figure 4] FIG. 1 is a diagram illustrating a schematic configuration of a neural network architecture. [Figure 5] FIG. 10 is a diagram for explaining that parameters obtained in the training of the neural network architecture are applied to an imaging device. [Figure 6] FIG. 10 is a diagram illustrating a process for obtaining a processing order parameter. [Figure 7] FIG. 7 is a diagram illustrating a process for obtaining processing order parameters following FIG. 6. [Figure 8]FIG. 8 is a diagram illustrating a process for obtaining processing order parameters following FIG. 7. [Figure 9] FIG. 9 is a diagram illustrating a process for obtaining processing order parameters following FIG. 8. [Figure 10] FIG. 2 is a diagram illustrating an example of a hardware configuration of an image processing unit. [Figure 11] FIG. 10 is a diagram illustrating another example of the hardware configuration of the image processing unit. [Figure 12] FIG. 10 is a diagram illustrating yet another example of the hardware configuration of the image processing unit. [Figure 13] FIG. 10 is a diagram illustrating another example of the hardware configuration of the image processing unit. [Figure 14] FIG. 10 is a diagram illustrating yet another example of the hardware configuration of the image processing unit. [Figure 15] FIG. 2 is a diagram illustrating an example of the configuration of an AWB processing block. [Figure 16] FIG. 10 is a diagram illustrating an example of the configuration of an imaging device according to a second embodiment. [Figure 17] FIG. 10 is a diagram illustrating an example of a data flow between a pre-determination processing unit, an image processing unit, and an inference processing unit. [Figure 18] FIG. 10 is a diagram illustrating another example of data flow between the preliminary judgment processing unit, the image processing unit, and the inference processing unit. [Figure 19] FIG. 10 is a diagram illustrating yet another example of the data flow between the preliminary judgment processing unit, the image processing unit, and the inference processing unit. [Figure 20] 10 is a flowchart illustrating an example of a flow of processing executed by the imaging device. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, with reference to the accompanying drawings, an embodiment of an information processing device according to the present technology will be described in the following order. <1. Information processing system configuration> <1-1. Configuration of imaging device> <1-2. Server device configuration> <2. Parameter search in server device> <3. Hardware configuration of the image processing unit> <3-1. Other configuration examples of the image processing unit> 4. Second Embodiment <5.Other> <6. Summary> <7. This Technology>

[0014] <1. Information processing system configuration> An information processing system 1A according to this embodiment will be described with reference to the accompanying drawings.

[0015] 1, the information processing system 1A includes an imaging device 2A and a server device 3. A plurality of imaging devices 2A and a plurality of server devices 3 may be provided.

[0016] The imaging device 2A and the server device 3 are connected to a communication network NW such as the Internet, so that they can communicate with each other.

[0017] <1-1. Configuration of imaging device> An example of the configuration of the imaging device 2A is shown in FIG. The imaging device 2A includes an imaging optical system 21, an image sensor 22, a control unit 23, an image processing unit 24A, and an inference processing unit 25A. The imaging device 2A also includes necessary units (not shown), such as a display unit that displays captured image data, an operation unit used for various operations, a power supply unit that supplies drive power to various components, and a storage unit that stores image data.

[0018] The imaging optical system 21 includes, for example, lenses such as a cover lens, a zoom lens, and a focus lens, and an iris mechanism. The imaging optical system 21 guides light (incident light) from a subject and collects the light on the light receiving surface of the image sensor 22.

[0019] The image sensor 22 is configured with a pixel array section 22a in which pixels having photoelectric conversion elements such as photodiodes are arranged two-dimensionally, and a readout circuit that reads out the charges accumulated in each pixel of the pixel array section 22a.

[0020] The image sensor 22 performs A / D (Analog / Digital) conversion processing and the like on the read-out captured image signal. Note that the image sensor 22 may also perform CDS (Correlated Double Sampling) processing and AGC (Automatic Gain Control) processing.

[0021] The image data output from the image sensor 22 is input as RAW image data Gr to the image processing unit 24A at the subsequent stage.

[0022] The control unit 23 is configured to include a microcomputer having, for example, a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), and the like.

[0023] The control unit 23 performs overall control of the imaging device 2A by causing the CPU to execute various processes in accordance with programs stored in the ROM or programs loaded into the RAM.

[0024] The control unit 23 issues drive instructions to an optical system drive unit (not shown) to drive the zoom lens, focus lens, diaphragm mechanism, etc. In response to these drive instructions, the optical system drive unit moves the focus lens and zoom lens and opens and closes the diaphragm blades of the diaphragm mechanism, i.e., drives the imaging optical system 21.

[0025] The control unit 23 controls writing and reading of various data to and from a storage unit (not shown). The storage unit is a non-volatile storage device such as a hard disk drive (HDD) or a flash memory device, and is used to save and record image data output from the image sensor 22.

[0026] The control unit 23 performs various data communications with external devices such as the server device 3 via a communication unit (not shown).

[0027] The image processing unit 24A performs predetermined image processing on the RAW image data Gr output from the image sensor 22. The image data obtained by performing the predetermined image processing by the image processing unit 24A is referred to as "processed image data Gp."

[0028] The predetermined image processing is realized by sequentially executing a plurality of image processing processes, such as AWB (Auto White Balance) processing, DMS (Demossaic) processing, CCM (Color Correction Matrix) processing, GC (Gamma Correction) processing, denoising processing, HSC (Hue Saturation Control) processing, and BCC (Brightness Contrast Control) processing. Note that these processes are merely examples.

[0029] In the image processing unit 24A, various types of image processing are normally performed in a predetermined order as predetermined image processing, thereby generating processed image data Gp.

[0030] On the other hand, in the present technology, the processing order of each image process as predetermined image processing executed by the image processing unit 24A is variable. That is, the image processing unit 24A has a structure in which the processing order of the processing blocks that perform each image process is variable.

[0031] Examples of hardware that functions as the image processing unit 24A include a CPU, a DSP (Digital Signal Processor), a programmable accelerator, a dedicated hardware accelerator, an ISP (Image Signal Processor), etc. Each of these hardware components has lower flexibility and higher processing speed in order of appearance.

[0032] The image processing unit 24A is realized by, for example, a programmable accelerator, which is more flexible than a dedicated hardware accelerator and requires less sacrifice of speed.

[0033] The configuration of the image processing unit 24A for varying the execution order of the processing blocks for realizing predetermined image processing will be described later.

[0034] The control unit 23 may drive the imaging optical system 21 and control the image sensor 22 based on the processing result of the image processing unit 24A. For example, the control unit 23 may adjust the exposure based on a histogram of brightness obtained as a result of the processing of the image processing unit 24A.

[0035] The processed image data Gp output from the image processing unit 24A is input to the inference processing unit 25A.

[0036] The inference processing unit 25A realizes a predetermined inference process by developing a trained model M1 obtained by training a predetermined neural network architecture NNA. The inference processing unit 25A is realized by, for example, a GPU (Graphics Processing Unit) or a DSP.

[0037] For example, the inference processing unit 25A performs processing such as detecting people appearing in the input processed image data Gp, estimating the number of people, and performing face authentication on people.

[0038] The inference result Di obtained in the inference processing unit 25A may be presented to the user or transmitted to another device via the control unit 23. Alternatively, as shown in Fig. 2, the inference result Di may be directly output from the inference processing unit 25A to the outside of the imaging device 2A.

[0039] The image processing unit 24A and the inference processing unit 25A of the imaging device 2A are set as a production environment or a product environment for obtaining a predetermined inference result Di based on the RAW image data Gr. That is, by applying appropriate parameters to the image processing unit 24A and the inference processing unit 25A, they can perform highly accurate inference processing on the RAW image data Gr output from the image sensor 22.

[0040] The parameters assigned to the image processing unit 24A include a processing order parameter PMo for determining the processing order of the above-mentioned processing blocks, and a processing content parameter PMp, which is a parameter for each processing block and changes the processing content of the processing block. Moreover, the parameters of the trained model M1 that are given to the inference processing unit 25A are referred to as model parameters PMm.

[0041] The processing order parameter PMo and the processing content parameter PMp of the image processing unit 24A and the inference processing unit 25A are calculated in, for example, a test environment or a learning environment. For example, the learning phase for the neural network architecture NNA requires a large amount of calculations, so it is preferable to implement it using an information processing device with abundant computer resources.

[0042] In this embodiment, as an example, the server device 3 searches for various parameters to be applied to the image processing unit 24A and the inference processing unit 25A of the imaging device 2A.

[0043] <1-2. Server device configuration> An example of the configuration of the server device 3 is shown in FIG. The server device 3 includes a CPU 71. The CPU 71 functions as an arithmetic processing unit that performs the various processes described above, and executes the various processes in accordance with a program stored in a ROM 72 or a nonvolatile memory unit 74 such as an EEPROM (Electrically Erasable Programmable Read-Only Memory), or a program loaded from a storage unit 79 to a RAM 73. The RAM 73 also stores data necessary for the CPU 71 to execute the various processes, as appropriate.

[0044] The CPU 71, ROM 72, RAM 73, and nonvolatile memory unit 74 are interconnected via a bus 83. To this bus 83, an input / output interface (I / F) 75 is also connected.

[0045] The input / output interface 75 is connected to an input unit 76 that includes an operator and an operation device. For example, the input unit 76 may be various types of operators or operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, or a remote controller. An operation by the user is detected by the input unit 76, and a signal corresponding to the input operation is interpreted by the CPU 71.

[0046] Furthermore, the input / output interface 75 is connected integrally or separately to a display unit 77 such as an LCD or organic EL panel, and an audio output unit 78 such as a speaker. The display unit 77 is a display unit that displays various information, and is configured, for example, by a display device provided in the housing of the computer device, or a separate display device connected to the computer device.

[0047] The display unit 77 displays images for various image processing, moving images to be processed, etc. on the display screen based on instructions from the CPU 71. Furthermore, the display unit 77 displays various operation menus, icons, messages, etc., i.e., GUI (Graphical User Interface), based on instructions from the CPU 71.

[0048] The input / output interface 75 may be connected to a storage unit 79 configured with a hard disk or solid-state memory, or a communication unit 80 configured with a modem or the like.

[0049] The communication unit 80 performs communication processing via a transmission path such as the Internet, and communication with various devices via wired / wireless communication, bus communication, and the like.

[0050] A drive 81 is also connected to the input / output interface 75 as required, and a removable storage medium 82 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory is appropriately mounted thereon.

[0051] The drive 81 can read data files such as programs used in various processes from a removable storage medium 82. The read data files are stored in a storage unit 79, and images and sounds contained in the data files are output on a display unit 77 and an audio output unit 78. Furthermore, the computer programs and the like read from the removable storage medium 82 are installed in the storage unit 79 as needed.

[0052] In this computer device, for example, software for the processing of this embodiment can be installed via network communication by the communication unit 80 or via a removable storage medium 82. Alternatively, the software may be stored in advance in the ROM 72, the storage unit 79, etc. Furthermore, the computer device may store the RAW image data Gr received from the image capture device 2A in a removable storage medium 82 via the storage unit 79 or drive 81.

[0053] The CPU 71 performs processing operations based on various programs, thereby realizing a parameter search function for acquiring various parameters in the process of inputting RAW image data Gr into the neural network architecture NNA and performing learning.

[0054] The communication unit 80 in the server device 3 functions as a receiving unit that receives the RAW image data Gr from the imaging device 2A. The communication unit 80 in the server device 3 also functions as a transmitting unit that transmits the parameters acquired by the parameter search function to the imaging device 2A.

[0055] The server device 3 is not limited to being configured as a single computer device as shown in Fig. 3, but may be configured as a system of multiple computer devices. The multiple computer devices may be systemized using a LAN (Local Area Network) or the like, or may be located in a remote location using a VPN (Virtual Private Network) or the like using the Internet or the like. The multiple computer devices may include computer devices as a server group (cloud) available through a cloud computing service.

[0056] <2. Parameter search in server device> The following describes the parameter search function in the server device 3. In the parameter search function in the server device 3, the above-mentioned processing order parameter PMo, processing content parameter PMp, and model parameter PMm are obtained during or after learning for the neural network architecture NNA.

[0057] In training the neural network architecture NNA, for example, RAW image data Gr is input as predetermined image data, and the neural network architecture NNA outputs an inference result Di obtained as a result of performing some kind of inference processing on the subject included in the RAW image data Gr.

[0058] Here, the neural network architecture NNA used for learning is one in which hardware constraints are set to realize the image processing unit 24A of the imaging device 2A.

[0059] For example, when the image processing unit 24A is realized using a programmable accelerator, learning is performed using a neural network architecture NNA that incorporates the constraints of the programmable accelerator as constraint conditions.

[0060] Hardware constraints include, for example, the range of values ​​and the number of bits handled by each processing block.

[0061] Specifically, the predetermined image processing includes, for example, AWB processing. Although the AWB processing implemented on the programmable accelerator can only handle 24-bit numerical values, if learning is performed in a state where the neural network architecture NNA handled by the server device 3 can handle 64-bit numerical values, the processing content parameter PMp obtained thereby may not be appropriate for the AWB processing block that can only handle 24-bit numerical values.

[0062] That is, the parameters explored in the learning environment are not necessarily optimized in the production environment.

[0063] Furthermore, the processing of each processing block such as the AWB processing block is indivisible for each processing block, and this can be considered as a constraint.

[0064] Specifically, if a certain processing block A consists of processing A1 and processing A2, and another processing block B consists of processing B1 and processing B2, there are only two possible processing orders for processing block A and processing block B, which must be executed inseparably: either execute processing A1, A2, B1, B2, or execute processing B1, B2, A1, A2.

[0065] However, if the parameters are searched without considering such constraints in the learning environment of the neural network architecture NNA, the processes A1, B2, and B1 The search result may be a processing order parameter PMo that executes the processing in the order shown in A2. Such a processing order parameter PMo cannot be applied directly to the programmable accelerator in the product environment, and this may result in design changes to the programmable accelerator.

[0066] In this embodiment, the neural network architecture NNA is constrained such that the processes A1 and A2 are executed inseparably and that the order of the processes A1 and A2 cannot be changed. Similar constraints are also imposed on the neural network architecture NNA for the processes B1 and B2.

[0067] By carrying out learning with such constraints imposed on the neural network architecture NNA, the parameters obtained as a result of learning can be suitably applied to the hardware used in the product environment.

[0068] The configuration of the neural network architecture NNA used for parameter search is shown in FIG.

[0069] The neural network architecture NNA consists of the image processing unit 24A P1, which is the first half, and the inference processing unit P2, which is the second half.

[0070] The parameters obtained by training such a neural network architecture NNA are divided into parameters obtained for the image processing part P1 and parameters obtained for the inference processing part P2.

[0071] The parameters obtained for the image processing part P1 are defined as processing order parameters PMo and processing content parameters PMp, while the parameters obtained for the inference processing part P2 are defined as model parameters PMm (see FIG. 5).

[0072] The processing order parameter PMo and the processing content parameter PMp are applied to the image processing unit 24A configured by a programmable accelerator, etc. At this time, since the hardware constraints of the image processing unit 24A are given as constraint conditions to the neural network architecture NNA, it is possible to implement it in actual hardware while maintaining the accuracy obtained during learning.

[0073] The model parameters PMm are applied to an inference processing unit 25A configured by an ISP and a DSP. That is, at least a part of a trained model M1 obtained by training the neural network architecture NNA is deployed in the inference processing unit 25A. The trained model M1 deployed in the inference processing unit 25A is a model obtained in the inference processing part P2 in the neural network architecture NNA. Therefore, the parameter search function realized by the CPU 71 of the server device 3 can also be regarded as a model generation function that generates the trained model M1 deployed in the inference processing unit 25A. That is, the CPU 71 functions as a model generation unit.

[0074] The process of obtaining the processing order parameter PMo indicating the processing order by the neural network architecture NNA is shown in FIG. 6 and other figures.

[0075] Each numerical value in the table shown in Figure 6 indicates the accuracy of the inference result Di output from the inference processing part P2. Note that the numerical values ​​in the table do not directly represent the accuracy of the inference result Di, but are merely relative evaluation values. Specifically, the higher the numerical value in the table, the more accurate the inference result Di that can be obtained.

[0076] The numbers following "Stage" indicate the order in which the processing blocks are selected. That is, the "Stage 1" row indicates the evaluation value used to determine the processing block to be selected as the first image processing. The "Stage 1" row is the evaluation value used to determine the first image processing to be performed on the RAW image data Gr.

[0077] Also, "bypass" in the figure means that the processing of the inference processing part P2 starts without performing additional image processing. For example, when "bypass" is selected in "Stage 3," this means that the processed image data Gp obtained by sequentially performing the processing selected in Stage 1 and the processing selected in Stage 2 is input to the inference processing part P2.

[0078] The inference accuracy when AWB processing is performed on the RAW image data Gr is higher than the inference accuracy when bypass is selected. Therefore, AWB processing is selected in Stage 1 (see the right side of Figure 6).

[0079] Next, the CCM processing, GC processing, denoising processing, HSC processing, and BCC processing are all processes that are assumed to be performed on image data after DMS processing.

[0080] Therefore, DMS processing is automatically selected in Stage 2 (see Figure 7).

[0081] In Stage 3, it is assumed that the highest inference accuracy is achieved when GC processing is selected for the image data after demosaic processing among CCM processing, GC processing, denoising processing, HSC processing, BCC processing, and bypass. Therefore, GC processing is selected in Stage 3 (see Figure 8).

[0082] The result of image processing selection in each stage is shown in Fig. 9. In the example shown, the inference accuracy in the inference processing part P2 can be improved by executing each processing block as the predetermined image processing in the following order: AWB processing, DMS processing, GC processing, denoising processing, HSC processing, denoising processing, and BCC processing.

[0083] Note that one processing block may be selected multiple times. In the example shown in Figure 9, denoising processing is selected in both Stage 4 and Stage 6.

[0084] Note that the examples shown in Figures 6 to 9 are intended to clearly explain the function of determining the processing order of the processing blocks, and when actually training the neural network architecture NNA, it is not necessary to determine the processing contents in order from the previous stage in the image processing part P1.

[0085] During learning, not only the processing order of the processing blocks but also the processing content parameter PMp for each processing block is optimized at the same time.

[0086] <3. Hardware configuration of the image processing unit> FIG. 10 shows an example of a hardware configuration when the image processing unit 24A of the imaging device 2A is realized by hardware.

[0087] Here, the processing block that performs AWB processing is designated as AWB processing block 24a. Similarly, the processing block that performs DMS processing is designated as DMS processing block 24b, the processing block that performs CCM processing is designated as CCM processing block 24c, the processing block that performs GC processing is designated as GC processing block 24d, the processing block that performs denoising processing is designated as denoising processing block 24e, the processing block that performs HSC processing is designated as HSC processing block 24f, and the processing block that performs BCC processing is designated as BCC processing block 24g.

[0088] The AWB processing block 24a, DMS processing block 24b, CCM processing block 24c, GC processing block 24d, denoising processing block 24e, HSC processing block 24f, and BCC processing block 24g are fabric-connected to each other so that the processing order can be changed.

[0089] As an example, in FIG. 10, the AWB processing block 24a, DMS processing block 24b, CCM processing block 24c, GC processing block 24d, denoising processing block 24e, HSC processing block 24f and BCC processing block 24g are each connected to the bus 26.

[0090] Each processing block is capable of transmitting the image data resulting from the processing to other processing blocks via a bus 26 .

[0091] The image sensor 22 and the inference processing unit 25A are also connected to the bus 26.

[0092] That is, in the image processing unit 24A, one of the processing blocks receives the RAW image data Gr from the image sensor 22. Then, the image processing in that processing block and the process of transmitting the image data as the processing result to other processing blocks are repeatedly executed. The processing block that has completed the final image processing outputs the processed image data Gp as input data via the bus 26 to the trained model M1 deployed in the inference processing unit 25A.

[0093] As described above, by adopting a configuration in which the processing blocks in the image processing unit 24A are fabric-connected, the processing order of the processing blocks can be changed as appropriate.

[0094] Note that one processing block may be configured to be able to execute the same image processing consecutively. For example, the predetermined image processing may be a processing order that includes executing denoising processing twice consecutively. In that case, the processing block that performs the denoising processing may be configured to be able to execute the same denoising processing as image processing multiple times consecutively without using the bus 26.

[0095] <3-1. Other configuration examples of the image processing unit> Another example of the hardware configuration of the image processing unit 24A is shown in FIG. Among the processing blocks included in the image processing unit 24A, there may be processing blocks that are likely to be executed in a predetermined order. A dedicated data transmission line Tr may be provided between such processing blocks.

[0096] Specifically, as shown in FIG. 11, the AWB processing block 24a has a dedicated data transmission line Tr for the DMS processing block 24b.

[0097] Similarly, the DMS processing block 24b has a dedicated data transmission line Tr for the CCM processing block 24c.

[0098] The CCM processing block 24c includes a dedicated data transmission line Tr for the GC processing block 24d.

[0099] The GC processing block 24d includes a dedicated data transmission line Tr for the denoising processing block 24e.

[0100] The denoising processing block 24e has a dedicated data transmission line Tr for the HSC processing block 24f.

[0101] The HSC processing block 24f includes a dedicated data transmission line Tr for the BCC processing block 24g.

[0102] When a dedicated data transmission line Tr is provided from the processing block that is the source of image data transmission to the processing block that is the destination of transmission, the image data is transmitted using the dedicated data transmission line Tr.

[0103] On the other hand, if a dedicated data transmission line Tr is not provided from the processing block that is the source of image data transmission to the processing block that is the destination of transmission, the image data is transmitted via the bus 26.

[0104] As a result, image data transmission between processing blocks that are closely connected is performed using a dedicated data transmission line Tr, thereby improving transmission speed. Also, when data is sent to a common bus 26, processing is required to match the data to a common input / output format on the bus 26. On the other hand, when image data is transmitted using a dedicated data transmission line Tr, processing to match the input / output format is not required.

[0105] Therefore, the execution speed of the predetermined image processing in the image processing unit 24A can be increased.

[0106] For example, when the above-described image processing is performed for each image frame captured by the image sensor 22, it is possible to accommodate even an increased number of image frames.

[0107] FIG. 12 shows yet another example of the hardware configuration of the image processing unit 24A. In the example shown in FIG. 12, an AWB processing block 24a and a DMS processing block 24b, which perform AWB processing and DMS processing, respectively, on image data before demosaicing, are connected to a first bus 26a.

[0108] In addition, the CCM processing block 24c, GC processing block 24d, denoising processing block 24e, HSC processing block 24f, and BCC processing block 24g, which perform CCM processing, GC processing, denoising processing, HSC processing, denoising processing, and BCC processing, which are processes on RGB image data after demosaicing, are connected to the second bus 26b.

[0109] The first bus 26a and the second bus 26b are configured to allow data transfer between them.

[0110] 12, image processing is performed in a selected order between the AWB processing block 24a and the DMS processing block 24b connected to the first bus 26a. The processed image data is sent to the second bus 26b. This image data is subjected to image processing in the selected order in each processing block connected to the second bus 26b, and the generated processed image data Gp is ​​provided to the trained model M1 of the inference processing unit 25A via the second bus 26b.

[0111] Another example of the hardware configuration of the image processing unit 24A is shown in FIG.

[0112] It is considered that changing the order of the AWB processing and demosaic processing, which are performed on image data before demosaic, does not affect the accuracy of the inference result Di. In this case, the processing of the AWB processing block 24a and the DMS processing block 24b may be fixed, and the order of the processing of each processing block on the RGB image data after demosaic may be configured to be changeable.

[0113] For example, as shown in FIG. 13, only the CCM processing block 24c, GC processing block 24d, denoising processing block 24e, HSC processing block 24f, and BCC processing block 24g, which perform processing on demosaiced RGB image data, are connected to the bus 26.

[0114] In this case, as shown in FIG. 13, it is desirable that image data be transmitted between the AWB processing block 24a and the DMS processing block 24b via a dedicated data transmission line Tr. This allows for faster data transfer of image data before demosaicing.

[0115] In each of the configurations shown in FIGS. 12 and 13, similarly to FIG. 11, a dedicated data transmission line Tr and bus 26 may be used in combination for transmitting image data between processing blocks that are closely connected.

[0116] FIG. 14 shows a hardware configuration of the image processing unit 24A that more actively uses a dedicated data transmission line Tr.

[0117] In the example shown in FIG. 14, only the image sensor 22, the AWB processing block 24a, the BCC processing block 24g, and the inference processing unit 25A are connected to the bus .

[0118] Between the AWB processing block 24a and the BCC processing block 24g, the DMS processing block 24b, the CCM processing block 24c, the GC processing block 24d, the denoising processing block 24e, and the HSC processing block 24f are connected in series via dedicated data transmission lines Tr, respectively.

[0119] The RAW image data Gr output from the image sensor 22 is first input to the AWB processing block 24a. Each of the seven processing blocks from the AWB processing block 24a to the BCC processing block 24g has a path for executing a predetermined process and a path for bypassing the predetermined process.

[0120] Specifically, the configuration of the AWB processing block 24a is shown in FIG.

[0121] The AWB processing block 24a is internally provided with a switch 27. The switch 27 is provided with an AWB processing unit 24a1 configured with a circuit that executes AWB processing, and an avoidance path 28 that avoids the AWB processing.

[0122] That is, the AWB processing block 24a can select whether or not to execute the AWB processing section 24a1 by operating the switch 27.

[0123] Each processing block other than the AWB processing block 24a is similarly configured and includes a processing unit and an avoidance path 28. For example, the DMS processing block 24b includes a switch 27, a DMS processing unit 24b1 that performs DMS processing, and an avoidance path 28.

[0124] The image data input to the AWB processing block 24 a undergoes image processing in one of the processing blocks before being output from the BCC processing block 24 g to the bus 26 .

[0125] For example, when image processing in the processing blocks is applied in the processing order shown in FIG. 9, in one loop from when image data is input from the bus 26 to the AWB processing block 24a to when the image data is output from the BCC processing block 24g to the bus 26, the switch 27 is controlled to the processing unit side only in the AWB processing block 24a, while the switch 27 is controlled to the avoidance path 28 side in the other processing blocks.

[0126] In the subsequent second loop, the switch 27 is controlled to the processing section side only in the DMS processing block 24b.

[0127] In the third loop, the switch 27 is controlled to the processing section side only in the GC processing block 24d.

[0128] After the predetermined image processing is completed, the processed image data is output to the inference processing unit 25A via the bus 26 and subjected to the predetermined inference processing.

[0129] 14 can be realized by simply modifying a part of the hardware configuration that performs conventional image processing in a fixed processing order. In other words, the configuration from the AWB processing block 24a to the BCC processing block 24g in the image processing unit 24A can be reused as is. Therefore, it is possible to realize at low cost a configuration for changing the processing order of each image process in a predetermined image process.

[0130] 4. Second Embodiment The information processing system 1B in the second embodiment is an example in which multiple sets of processing content parameters PMp and processing order parameters PMo are used by switching between them. Specifically, the image processing unit 24B of the imaging device 2B in the information processing system 1B selects one set of parameters from multiple sets of processing content parameters PMp and processing order parameters PMo prepared according to the situation and uses it for predetermined image processing.

[0131] 1, the information processing system 1B includes an imaging device 2B and a server device 3. A plurality of imaging devices 2B and a plurality of server devices 3 may be provided.

[0132] The imaging device 2B and the server device 3 are connected to a communication network NW such as the Internet, so that they can communicate with each other.

[0133] As shown in FIG. 16, the imaging device 2B of the information processing system 1B includes an imaging optical system 21, an image sensor 22, a preliminary determination processing unit 29, a control unit 23, an image processing unit 24B, and an inference processing unit 25B.

[0134] The preliminary judgment processing unit 29 receives the RAW image data Gr from the image sensor 22. A model trained by machine learning is deployed in the preliminary judgment processing unit 29. The trained model deployed in the preliminary judgment processing unit 29 is referred to as a "predictor M2." The predictor M2 that is a trained model is obtained by, for example, a training process in the server device 3.

[0135] The predictor M2 selects a combination of parameters to be adopted by the image processing unit 24B based on the characteristics of the RAW image data Gr before the RAW image data Gr is input to the image processing unit 24B. That is, the predictor M2 is a model acquired by learning, and outputs a combination of parameters to be adopted by the image processing unit 24B using the RAW image data Gr as input data.

[0136] In the following description, a group of parameters including at least some of the processing order parameters PMo, the processing content parameters PMp, and the model parameters PMm will be referred to as a "parameter set PS."

[0137] When the parameter set PS includes the processing order parameter PMo, the image processing unit 24B has the above-described configuration that can change the processing order of various image processes included in the predetermined image processing.

[0138] For example, when RAW image data Gr taken during the day is input, the predictor M2 outputs the parameter set PS1 itself or data specifying the parameter set PS1. When RAW image data Gr taken at night is input, the predictor M2 outputs the parameter set PS2 itself or data specifying the parameter set PS2.

[0139] The parameter set PS1 is a group of parameters obtained by executing the learning phase of the neural network architecture NNA described above using only the RAW image data Gr captured during the day.

[0140] The parameter set PS2 is a group of parameters obtained by executing the learning phase of the neural network architecture NNA using only the RAW image data Gr captured at night.

[0141] The parameter sets PS1 and PS2 are merely examples. The parameter sets PS1 and PS2 are used to switch the parameter sets PS depending on the time of day, but other parameter sets PS may also be prepared depending on the season, such as summer or winter.

[0142] Alternatively, parameter sets PS may be prepared according to the subject. For example, a parameter set PSa may be prepared when the subject is a human, and a parameter set PSb may be prepared when the subject is a vehicle. In this case, for example, the parameter set PSa when the subject is a human may include model parameters PMm that perform predetermined inference processing for humans, such as estimating posture. And the parameter set PSb when the subject is a vehicle may include model parameters PMm that perform inference processing different from that when the subject is a human, such as obtaining a vehicle license plate number.

[0143] In other words, the parameter set PS selected as a result of performing inference on the RAW image data Gr using the predictor M2 may not only change the manner of a predetermined image processing in the image processing unit 24B, but may also change the inference processing in the inference processing unit 25B to something completely different in nature.

[0144] By changing the parameter set PS through the selection of the predictor M2, the inference processing to be performed can be changed according to the captured RAW image data Gr, and a predetermined image processing can be realized to obtain processed image data Gp that is optimal for realizing that inference processing.

[0145] An example of a specific data flow between the preliminary judgment processing unit 29, the image processing unit 24B, and the inference processing unit 25B is shown in FIGS.

[0146] The preliminary judgment processing unit 29 is configured by, for example, a DSP, etc. The image processing unit 24B is configured by, for example, a programmable accelerator, etc. Furthermore, the inference processing unit 25B is configured by, for example, a DSP, etc.

[0147] The pre-determination processing unit 29 has a predictor M2 deployed. Based on the estimation result of the predictor M2, the pre-determination processing unit 29 selects a parameter set PS corresponding to the estimation result from among multiple parameter sets PS and applies the selected parameter set PS to the image processing unit 24B and the inference processing unit 25B. At this time, the pre-determination processing unit 29 may acquire the parameter set PS itself and transmit it to the image processing unit 24B and the inference processing unit 25B. Alternatively, the pre-determination processing unit 29 may transmit selection information of the parameter set PS to the image processing unit 24B and the inference processing unit 25B, and have the image processing unit 24B and the inference processing unit 25B acquire and apply the parameters.

[0148] The application of the parameters by the pre-determination processing unit 29 to the image processing unit 24B and the inference processing unit 25B may be performed via the control unit 23.

[0149] The preliminary determination processing unit 29 further transmits the RAW image data Gr input from the preceding image sensor 22 as is to the image processing unit 24B, whereby the RAW image data Gr output from the image sensor 22 is input to the image processing unit 24B.

[0150] The imaging device 2B may be configured so that the RAW image data Gr is transmitted directly from the image sensor 22 to the image processing unit 24B without going through the preliminary determination processing unit 29.

[0151] The inference processing unit 25B applies parameters based on the results of the estimation processing in the preliminary judgment processing unit 29 to the expanded trained model M1, thereby performing highly accurate inference processing in line with the purpose.

[0152] If the content of the inference process is to be significantly changed, it may be possible to change the structure of the trained model M1 itself. In that case, as shown in Figure 18, a configuration may be adopted in which multiple trained models M1a and M1b with different network structures are deployed in 25B.

[0153] The pre-judgment processing unit 29 selects the trained model M1 to be used from the trained models M1a, M1b, ... according to the inference result of the predictor M2, and further applies the selected parameters to the selected trained model M1. That is, the preliminary judgment processing unit 29 may have a function of selecting the trained model M1.

[0154] The function of the preliminary judgment processing unit 29 can also be included in the inference processing unit 25C.

[0155] Specifically, as shown in Fig. 19, both the trained model M1 and the predictor M2 may be deployed in the inference processing unit 25C. Furthermore, the trained model M1 deployed in the inference processing unit 25C may be multiple models, such as trained model M1a and trained model M1b.

[0156] The RAW image data Gr is input from the image sensor 22 not only to the image processing unit 24C but also to the inference processing unit 25C.

[0157] In the inference processing unit 25C, the RAW image data Gr is input to the predictor M2, and the parameter set PS selected according to the inference result is applied to the image processing unit 24C and the trained model M1.

[0158] The image processing unit 24C performs predetermined image processing according to the selected parameter set PS on the RAW image data Gr input from the image sensor 22, and supplies the processed image data Gp to the trained model M1 of the inference processing unit 25C.

[0159] That is, the inference processing unit 25C performs a predetermined inference process while appropriately switching between the trained model M1 and the predictor M2.

[0160] This allows the number of chips, etc. that make up the processing unit used for inference processing to be reduced, thereby reducing costs and size.

[0161] An example of the flow of processing executed in the imaging device 2B of the information processing system 1B in the second embodiment is shown in Fig. 20. As described above, each processing executed in the preliminary judgment processing unit 29 may be realized by the inference processing unit 25C in which the predictor M2 is deployed.

[0162] In step S101, the imaging device 2B executes a frame imaging process using the image sensor 22. This process may be realized by the control unit 23 controlling the image sensor 22.

[0163] In step S102, the imaging device 2B performs inference processing using the predictor M2 by the preliminary determination processing unit 29. This processing is processing for obtaining selection information for the parameter set PS based on the feature amount for the RAW image data Gr.

[0164] In step S103, the pre-determination processing unit 29 of the imaging device 2B determines whether or not a parameter change is required. Specifically, the pre-determination processing unit 29 determines whether or not the parameter set PS selected in step S102 is currently being set, and if it is determined that the parameter set PS is currently being set, it determines "No" in step S103, and if it is determined that the parameter set PS is not currently being set, it determines "Yes" in step S103.

[0165] If the answer is "Yes" in step S103, the pre-determination processing unit 29 of the imaging device 2B applies the newly selected parameter set PS to the image processing unit 24B and the inference processing unit 25B in step S104.

[0166] On the other hand, if the determination in step S103 is "No", the image capture device 2B avoids the process of step S104.

[0167] In step S105, the inference processing unit 25B of the imaging device 2B performs inference processing using the trained model M1.

[0168] Subsequently, the inference processing unit 25B of the imaging device 2B outputs the inference result Di in step S106.

[0169] The imaging device 2B repeatedly executes the processes from step S101 to step S106 until a predetermined termination condition is met.

[0170] 20, parameters can be changed for each frame, but the present invention is not limited to this, and may be configured to execute the processing from step S102 to step S106 every few frames, or to execute the processing from step S102 to step S106 when a predetermined condition is met.

[0171] The case where a predetermined condition is met is, for example, the case where a change in the feature amount calculated for the RAW image data Gr exceeds a threshold, and the feature amount may be a moving average of several frames.

[0172] <5.Other> In the following description, when the information processing system 1A, the information processing system 1B, etc. are collectively referred to, they will be referred to as "information processing system 1" excluding the capital alphabet added to the end of the reference numeral. The same applies to other reference numerals.

[0173] In the learning phase of the neural network architecture NNA, an example has been described in which parameters to be applied to the image processing unit 24 and the inference processing unit 25 are searched for. In the learning phase of this neural network architecture NNA, furthermore, the configuration may be such that setting information for the imaging optical system 21, the image sensor 22 and the imaging operation thereof is obtained as a search result.

[0174] For example, in the learning phase of the neural network architecture NNA, setting information such as the shutter speed of the image sensor 22 and the control positions of the various lenses of the imaging optical system 21 may be obtained. The setting information obtained as a result of the search may include information to be set for the control unit 23, for example.

[0175] <6. Summary> As explained using the above examples, the server device 3 as an information processing device of the present technology is equipped with a model generation unit (CPU 71) that generates a trained model M1 by training a neural network architecture NNA to which hardware constraints on which the trained model M1 (M1a, M1b) is deployed are assigned as constraint conditions. A trained model M1 may be generated by training the neural network architecture NNA and used for a predetermined process. In this case, highly accurate processing can be performed by using the trained model M1 to which optimized parameters (model parameters PMm) are applied. Incidentally, processing using the trained model M1 can sometimes be achieved by incorporating the parameters (model parameters PMm) calculated during the generation process of the trained model M1 into relatively inflexible hardware, instead of using relatively flexible software. This allows for high speed processing at the expense of flexibility. However, in many cases, the parameters calculated during the generation process of the trained model M1 cannot be applied to hardware as is, and hardware design changes are required, which leads to increased costs. In this configuration, training is performed while incorporating the constraints of the hardware to which the trained model M1 is deployed. The parameters obtained through this training can be applied directly to the target hardware. Therefore, optimum processing can be achieved at high speed using limited hardware.

[0176] As explained with reference to Figures 4 and 5, in the server device 3 as an information processing device, the constraints imposed on the neural network architecture NNA may be constraints on the hardware that functions as the image processing unit 24 that obtains image data by image processing. The predetermined processing may be, for example, predetermined image processing of image data, etc. Specifically, this is the case when image data suitable for subsequent inference processing is obtained from the RAW image data Gr. Such image processing can be performed at a higher speed by using an image processing unit (such as an ISP) as relatively inflexible hardware rather than by using relatively flexible software. The parameters obtained when generating the trained model M1 are parameters obtained taking into account hardware constraints such as ISP, so they can be applied to the hardware as is. Therefore, optimal image processing can be performed using hardware, and high-speed processing can be achieved while ensuring inference accuracy.

[0177] As explained with reference to Figures 4 and 5, in the server device 3 as an information processing device, the neural network architecture NNA performs predetermined image processing to obtain processed image data Gp from predetermined image data input as input data, and the model generation unit (CPU 71) may search for parameters related to the predetermined image processing (processing order parameter PMo, processing content parameter PMp) in learning. For example, the predetermined image processing may be AWB processing, CCM processing, denoising processing, or the like. In such image processing, the processed image data Gp obtained after processing varies depending on the adjustment of the processing content parameter PMp. The optimum processed image data Gp may further vary depending on the content of the inference processing at the subsequent stage. The model generation unit (CPU 71) searches for parameters used in these image processes during learning for the neural network architecture NNA. In hardware to which such parameters are applied, image processing is realized to obtain processed image data Gp suitable for subsequent processing. Therefore, the parameter search results can be used efficiently.

[0178] As explained with reference to each of Figures 4 to 9, in the server device 3 as an information processing device, the predetermined image processing includes multiple types of image processing, and the model generation unit (CPU 71) may search for a parameter (processing order parameter PMo) that determines the processing order of the multiple types of image processing as a parameter related to the predetermined image processing. Examples of the multiple types of image processing include AWB processing, CCM processing, GC processing, denoising processing, HSC processing, and BCC processing. These processes are usually performed in an image processing unit 24 such as an ISP provided within the image sensor 22 or at a stage subsequent to the image sensor 22 . Changing the order of these image processes changes the processed image data Gp that is generated, which means that by changing the order of the processes, it is possible to obtain processed image data Gp that is more suitable for subsequent processing. According to this configuration, when learning the neural network architecture NNA, the processing order of multiple types of image processing to be executed by the image processing unit 24 is searched for as a processing order parameter PMo, thereby making it possible to generate processed image data Gp suitable for subsequent processing using hardware such as an ISP.

[0179] As explained with reference to Figures 4 and 5, etc., in the server device 3 as an information processing device, the parameters related to the specified image processing may include parameters (processing content parameters PMp) corresponding to the processing content for each of multiple types of image processing. The processing units (AWB processing block 24a, etc.) that perform various image processing such as AWB processing can adjust the processing content by adjusting the parameters. That is, by adjusting the processing content parameters PMp of each image processing included in a predetermined image processing, the output processed image data Gp can be adjusted. According to this configuration, during training of the neural network architecture NNA, a processing content parameter PMp for adjusting the processing content is searched for for each of the multiple types of image processing executed by the image processing unit 24. Therefore, it becomes possible to generate processed image data Gp that is more suitable for subsequent processing using hardware such as an ISP.

[0180] As explained with reference to Figures 4 and 5, in the server device 3 as an information processing device, the neural network architecture NNA outputs the inference results obtained by performing a predetermined inference process on the processed image data Gp as output data, and the model generation unit (CPU 71) may search for parameters (model parameters PMm) related to the predetermined inference process during learning. For example, the neural network architecture NNA is a model that receives image data as input, performs inference processing, and outputs the inference results. The neural network architecture NNA is equipped with a structure for realizing image processing, which is processing of the input image data, and inference processing targeting the processed image data Gp. The image processing and inference processing are executed as a single unit, for example, from the input of RAW image data Gr to the output of the inference results. Furthermore, during training of the neural network architecture NNA, both various parameters used in image processing and various parameters used in inference processing are obtained. By applying at least some of the parameters related to image processing (processing content parameter PMp, processing order parameter PMo) and parameters related to inference processing (model parameter PMm) obtained here to hardware, it is possible to obtain inference results of the inference processing with high accuracy and high speed. Also, a wide variety of modes are possible, such as realizing image processing using a programmable accelerator, which is relatively high-speed and low-flexibility hardware, and realizing inference processing using relatively low-speed and highly flexible software such as a CPU or DSP, or a software-oriented, highly flexible configuration.

[0181] As described with reference to FIGS. 4 and 5, in the server device 3 serving as an information processing device, the predetermined image data may be RAW image data Gr. As a result, the RAW image data Gr output from the image sensor 22 is subjected to image processing using hardware to which the optimized parameters have been applied. Therefore, it is possible to perform optimal and high-speed image processing, and subsequent processing can be performed favorably.

[0182] As explained with reference to each of Figures 3 to 5, the information processing device is a server device 3 equipped with a receiving unit (communication unit 80) that receives RAW image data Gr output from the image sensor 22, a model generation unit (CPU 71), and a transmitting unit (communication unit 80) that transmits the trained model M1, and the model generation unit (CPU 71) may train the neural network architecture NNA using the received RAW image data Gr. As a result, the parameters are optimized by learning the neural network architecture NNA in the server device 3. That is, the parameter search process is executed by an information processing device that is generally considered to have higher performance than an edge computer. Therefore, it is possible to efficiently perform parameter search and obtain optimized parameters in a short time.

[0183] The information processing method of the present technology includes a step of generating a trained model M1 by training a neural network architecture NNA to which constraints on the hardware on which the trained model M1 is deployed are assigned as constraint conditions.

[0184] The program of the present technology causes an information processing device to perform the function of generating a trained model M1 by training a neural network architecture NNA to which hardware constraints on which the trained model M1 is deployed are assigned as constraint conditions. Such an information processing method or program can provide the various effects described above.

[0185] Such a program can be pre-recorded on a hard disk drive (HDD) as a recording medium built into a device such as a computer, or on a ROM in a microcomputer having a CPU. Alternatively, the program can be temporarily or permanently stored (recorded) on a removable recording medium such as a flexible disk, a CD-ROM (Compact Disk Read Only Memory), a Magneto Optical (MO) disk, a Digital Versatile Disc (DVD), a Blu-ray Disc (registered trademark), a magnetic disk, a semiconductor memory, or a memory card. Such removable recording media can be provided as so-called packaged software. Such a program can be installed onto a personal computer or the like from a removable recording medium, or can be downloaded from a download site via a network such as a LAN or the Internet.

[0186] As explained using the examples above, the imaging device 2 of the present technology includes an image processing unit 24 that performs predetermined image processing on RAW image data Gr output from the image sensor 22 to obtain processed image data Gp, and the parameters used in the predetermined image processing may be parameters obtained in the process of generating a trained model M1 by training a neural network architecture NNA to which hardware constraints in the image processing unit 24 are assigned as constraint conditions. This makes it possible to apply the parameters optimized by learning of the neural network architecture NNA directly to the image processing unit 24 included in the imaging device 2. Therefore, when the image processing unit 24 is realized by relatively high-speed, low-flexibility hardware, it is possible to apply suitable parameters that take into account the low flexibility of the hardware, thereby realizing high-speed image processing.

[0187] As described with reference to FIG. 2 and other figures, the image processing unit 24 in the imaging device 2 may be configured as a programmable accelerator. For example, when various types of image processing are performed by a programmable accelerator, parameter search processing is performed taking into account the range of values ​​used for each image processing, the number of significant digits, and the like. This allows the processing accuracy on paper when using the trained model M1 to match the processing accuracy when applying the parameters to hardware, preventing a decrease in processing accuracy when implemented in hardware. Furthermore, there may be constraints that require each image processing to be executed indivisibly in hardware, such as when it is impossible to implement a configuration in which one image processing is executed in the middle of another image processing. By performing training using the neural network architecture NNA, which takes into account such hardware constraints, the obtained parameters can be implemented directly in the hardware, making it possible to achieve the desired performance.

[0188] As described with reference to FIG. 2 and the like, the imaging device 2 may include an inference processing unit 25 that performs a predetermined inference process on the processed image data Gp using the trained model M1. This makes it possible to simultaneously search for parameters to be used in the image processing unit 24 and generate a trained model M1 to be used in inference processing during the process of training the neural network architecture NNA. Therefore, it is possible to efficiently set the imaging device 2 to achieve desired processing.

[0189] As explained with reference to Figure 2 etc., the image processing unit 24 in the imaging device 2 performs multiple types of image processing as the predetermined image processing, and the parameters related to the predetermined image processing may include parameters that determine the processing order of the multiple types of image processing. This allows various types of image processing, such as AWB processing, CCM processing, and GC processing, to be performed in an optimal processing order as predetermined image processing. Furthermore, since these image processing operations take into account the constraints of the hardware that realizes the image processing unit 24, they can be applied directly to the image processing unit 24, resulting in high efficiency.

[0190] As explained with reference to each of Figures 10 to 13, the image processing unit 24 in the imaging device 2 may have a configuration (bus 26) in which processing blocks (AWB processing block 24a, DMS processing block 24b, CCM processing block 24c, GC processing block 24d, denoising processing block 24e, HSC processing block 24f, BCC processing block 24g, etc.) that perform each of the image processing included in multiple types of image processing are fabric-connected. This makes it possible to appropriately reflect in the image processing unit 24 the parameters obtained in the learning of the neural network architecture NNA and related to the processing order of various image processes.

[0191] As explained with reference to Figure 12 etc., the image processing unit 24 in the imaging device 2 may have a first fabric connection (first bus 26a) in which processing blocks among multiple processing blocks that target image data before demosaicing are each fabric-connected, and a second fabric connection (second bus 26b) in which processing blocks that target image data after demosaicing are each fabric-connected. It is considered inappropriate to switch the processing order between the processing block targeted at image data before demosaicing and the processing block targeted at image data after demosaicing. According to this configuration, the processing order of the processing blocks can be interchanged between the processing blocks before and after demosaicing. This eliminates the need for excessive flexibility due to fabric connections, and makes it possible to simplify bus wiring, for example. Furthermore, since it is possible to standardize input / output for the first fabric connection and for the second fabric connection, it is possible to reduce the difficulty of standardizing input / output data.

[0192] As described with reference to FIG. 9 and other figures, in the imaging device 2, the processing order of multiple types of image processing may include a processing order in which the same image processing is executed multiple times. This makes it possible to execute each processing block as a predetermined image processing in a more suitable order, and obtain processed image data Gp suitable for subsequent processing.

[0193] As explained with reference to each of Figures 16 to 20, the imaging device 2 is provided with a pre-determination processing unit 29 that performs pre-determination processing on the RAW image data Gr using an inference device (predictor M2), and the image processing unit 24 may select parameters to be used for predetermined image processing depending on the results of the pre-determination processing. This makes it possible to change the parameters used in a predetermined image processing for an image captured during the day and an image captured at night, for example. These daytime parameters (parameter set PS1) and nighttime parameters (parameter set PS2) are, for example, parameters that are obtained in advance during learning of the neural network architecture NNA. By appropriately changing the parameters based on the imaging conditions or the characteristics of the captured image, it is possible to generate processed image data Gp that is more suitable for subsequent processing.

[0194] The image processing method of the present technology includes a step of performing predetermined image processing on RAW image data Gr output from an image sensor 22 to obtain processed image data Gp, and a step of performing predetermined inference processing on the processed image data Gp using a trained model M1. The trained model M1 is obtained by training a neural network architecture NNA to which hardware constraints used in the step of obtaining the processed image data Gp are assigned as constraint conditions. Such an image processing method can provide the various effects described above.

[0195] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0196] Furthermore, the above-described examples may be combined in any manner, and even when various combinations are used, the various effects described above can be obtained.

[0197] <7. This Technology> (1) A model generation unit is provided that generates the trained model by training a neural network architecture to which constraints on the hardware on which the trained model is deployed are assigned as constraint conditions. Information processing device. (2) The constraints imposed on the neural network architecture are constraints on hardware that functions as an image processing unit that obtains image data by image processing. The information processing device according to (1) above. (3) The neural network architecture performs predetermined image processing to obtain processed image data from predetermined image data input as input data, The model generation unit searches for parameters related to the predetermined image processing in the learning. The information processing device according to any one of (1) to (2) above. (4) the predetermined image processing includes a plurality of types of image processing, The model generation unit searches for a parameter that determines a processing order of the plurality of types of image processing as a parameter related to the predetermined image processing. The information processing device according to (3) above. (5) The parameters relating to the predetermined image processing include parameters corresponding to the processing contents of each of the plurality of types of image processing. The information processing device according to (4) above. (6) the neural network architecture performs a predetermined inference process on the processed image data and outputs an inference result as output data; The model generation unit searches for parameters related to the predetermined inference process in the learning. An information processing device according to any one of (3) to (5) above. (7) The predetermined image data is RAW image data. An information processing device according to any one of (3) to (6) above. (8) A server device including a receiving unit that receives RAW image data output from an image sensor, the model generation unit, and a transmitting unit that transmits the trained model, The model generation unit performs the learning of the neural network architecture using the received RAW image data. An information processing device according to any one of (1) to (7) above. (9) and generating the trained model by training a neural network architecture to which constraints on the hardware on which the trained model is deployed are assigned as constraint conditions. Information processing methods. (10) The information processing device is caused to execute a function of generating a trained model by training a neural network architecture to which constraints on the hardware on which the trained model is deployed are assigned as constraint conditions. program. (11) an image processing unit that performs predetermined image processing on the RAW image data output from the image sensor to obtain processed image data; The parameters used in the predetermined image processing are parameters obtained in the process of generating a trained model by performing training on a neural network architecture to which hardware constraints in the image processing unit are assigned as constraint conditions. Imaging device. (12) The image processing unit is configured as a programmable accelerator. The information processing device according to (11) above. (13) an inference processing unit that performs a predetermined inference process on the processed image data using the trained model; The imaging device according to any one of (11) to (12) above. (14) the image processing unit executes a plurality of types of image processing as the predetermined image processing, The parameters relating to the predetermined image processing include a parameter that determines the processing order of the plurality of types of image processing. The imaging device according to any one of (11) to (13) above. (15) The image processing unit has a configuration in which processing blocks that perform each of the image processes included in the plurality of types of image processing are fabric-connected. The imaging device according to (14) above. (16) The image processing unit has a first fabric connection in which the processing blocks for image data before demosaicing among the plurality of processing blocks are fabric-connected, and a second fabric connection in which the processing blocks for image data after demosaicing are fabric-connected. The imaging device according to (15) above. (17) The processing order of the multiple types of image processing includes a processing order in which the same image processing is executed multiple times. The imaging device according to any one of (14) to (16) above. (18) a pre-determination processing unit that performs pre-determination processing on the RAW image data by using an inference unit; The image processing unit selects parameters to be used in the predetermined image processing in accordance with the result of the preliminary determination processing. The imaging device according to any one of (11) to (17) above. (19) a step of performing predetermined image processing on the RAW image data output from the image sensor to obtain processed image data; and performing a predetermined inference process on the processed image data using a trained model, The trained model is obtained by training a neural network architecture to which constraints on the hardware used in the step of obtaining the processed image data are assigned as constraint conditions. Image processing methods. [Explanation of symbols]

[0198] 3. 3A server equipment 22 Image Sensor 24, 24A, 24B, 24C Image processing unit 24a AWB processing block (processing block) 24b DMS processing block (processing block) 24c CCM Processing Block (Processing Block) 24d GC Processing Block (Processing Block) 24e Denoising Processing Block (Processing Block) 24f HSC Processing Block (Processing Block) 24g BCC Processing Block (Processing Block) 25, 25A, 25B, 25C inference processing unit 29 Pre-determination processing unit 71 CPU (model generation unit) 80 Communication unit (receiving unit, transmitting unit) Gp processed image data Gr RAW image data M1, M1a, M1b pre-trained models M2 predictor NNA Neural Network Architecture PMm model parameters (parameters) PMo processing order parameters (parameters) PMp processing parameters (parameters)

Claims

1. A model generation unit is provided that generates the trained model by training a neural network architecture to which constraints on the hardware on which the trained model is deployed are assigned as constraint conditions. Information processing device.

2. The constraints imposed on the neural network architecture are constraints on hardware that functions as an image processing unit that obtains image data by image processing. The information processing device according to claim 1 .

3. The neural network architecture performs predetermined image processing to obtain processed image data from predetermined image data input as input data, The model generation unit searches for parameters related to the predetermined image processing in the learning. The information processing device according to claim 1 .

4. the predetermined image processing includes a plurality of types of image processing, The model generation unit searches for a parameter that determines a processing order of the plurality of types of image processing as a parameter related to the predetermined image processing. The information processing device according to claim 3 .

5. The parameters relating to the predetermined image processing include parameters corresponding to the processing contents of each of the plurality of types of image processing. The information processing device according to claim 4 .

6. the neural network architecture performs a predetermined inference process on the processed image data and outputs an inference result as output data; The model generation unit searches for parameters related to the predetermined inference process in the learning. The information processing device according to claim 3 .

7. The predetermined image data is RAW image data. The information processing device according to claim 3 .

8. a server device including a receiving unit that receives RAW image data output from an image sensor, the model generation unit, and a transmitting unit that transmits the trained model; The model generation unit performs the learning of the neural network architecture using the received RAW image data. The information processing device according to claim 1 .

9. and generating the trained model by training a neural network architecture to which constraints on the hardware on which the trained model is deployed are assigned as constraint conditions. Information processing methods.

10. The information processing device is caused to execute a function of generating a trained model by training a neural network architecture to which constraints on the hardware on which the trained model is deployed are assigned as constraint conditions. program.

11. an image processing unit that performs predetermined image processing on raw image data output from the image sensor to obtain processed image data; The parameters used in the predetermined image processing are parameters obtained in the process of generating a trained model by performing training on a neural network architecture to which hardware constraints in the image processing unit are assigned as constraint conditions. Imaging device.

12. The image processing unit is configured as a programmable accelerator. The information processing device according to claim 11.

13. an inference processing unit that performs a predetermined inference process on the processed image data using the trained model; The imaging device according to claim 11.

14. the image processing unit executes a plurality of types of image processing as the predetermined image processing, The parameters relating to the predetermined image processing include a parameter that determines the processing order of the plurality of types of image processing. The imaging device according to claim 11.

15. The image processing unit has a configuration in which processing blocks that perform each of the image processes included in the plurality of types of image processing are fabric-connected. The imaging device according to claim 14.

16. The image processing unit has a first fabric connection in which the processing blocks for image data before demosaicing among the plurality of processing blocks are fabric-connected, and a second fabric connection in which the processing blocks for image data after demosaicing are fabric-connected. The imaging device according to claim 15.

17. The processing order of the multiple types of image processing includes a processing order in which the same image processing is executed multiple times. The imaging device according to claim 14.

18. a pre-determination processing unit that performs pre-determination processing on the RAW image data using an inference unit; The image processing unit selects parameters to be used in the predetermined image processing in accordance with the result of the preliminary determination processing. The imaging device according to claim 11.

19. a step of performing predetermined image processing on the raw image data output from the image sensor to obtain processed image data; and performing a predetermined inference process on the processed image data using a trained model, The trained model is obtained by training a neural network architecture to which constraints on the hardware used in the step of obtaining the processed image data are assigned as constraint conditions. Image processing methods.