Information processing device, information processing system and information processing method

By dividing the CNN into two parts for line-based processing, the method addresses the challenge of slow image processing in existing devices, achieving faster and more efficient image recognition.

WO2026063272A1PCT designated stage Publication Date: 2026-03-26SONY SEMICON SOLUTIONS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-03-26

Smart Images

  • Figure JP2025031743_26032026_PF_FP_ABST
    Figure JP2025031743_26032026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to the present disclosure includes a first processing unit for executing processing by a first network from among the first network, which includes an input layer, and a second network, which includes an output layer, the first and second networks being obtained by splitting a network that processes input image data using an n-pixel × n-line filter (where n is an integer of 2 or more) at an intermediate layer between an N-th layer and an (N+1)-th layer (where N is an integer of 2 or more). The first processing unit causes the first network to perform the processing on the input image in units of n lines, and outputs the result of the processing line by line to the second network.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing system, and information processing method

[0001] This disclosure relates to an information processing device, an information processing system, and an information processing method.

[0002] Sensors that include an image sensor and a network that performs image processing such as image recognition on the output of the image sensor have been known for some time.

[0003] Japanese Patent Publication No. 2022-186333

[0004] In information processing devices equipped with such sensors, there is a need to enable faster image processing of the output of the image sensor.

[0005] Therefore, the present disclosure aims to provide an information processing device that enables faster image processing of the output of an image sensor in a sensor configuration that includes an image sensor and a network that performs image processing on the output of the image sensor.

[0006] The information processing apparatus according to this disclosure comprises a first network including an input layer and a second network including an output layer, wherein the network that processes input image data using an n-pixel × n-line (n is an integer of 2 or more) filter is divided between the Nth layer and the (N+1)th layer (N is an integer of 2 or more) in the intermediate layer, and a first processing unit that executes the processing by the first network, wherein the first processing unit outputs the results of the processing line by line to the second network after the first network has performed the processing on the input image in units of n lines.

[0007] This is a functional block diagram of an example illustrating the functions of an information processing device according to an embodiment. This is a diagram showing an example in which the sensor according to the embodiment is formed by a two-layer stacked CIS. This is a diagram showing an example in which the sensor according to the embodiment is formed by a three-layer stacked CIS. This is a block diagram showing the configuration of an example imaging device applicable to the embodiment. This is a block diagram showing the hardware configuration of an example information processing device according to an embodiment. This is a schematic diagram for explaining the rolling shutter method. This is a schematic diagram for explaining the rolling shutter method. This is a schematic diagram for explaining the rolling shutter method. This is a schematic diagram for explaining line thinning in the rolling shutter method. This is a schematic diagram for explaining line thinning in the rolling shutter method. This is a schematic diagram for explaining line thinning in the rolling shutter method. This is a diagram schematically showing an example of another imaging method in the rolling shutter method. This is a schematic diagram for explaining another imaging method in the rolling shutter method. This is a schematic diagram for explaining the global shutter method. This is a schematic diagram for explaining the global shutter method. This is a schematic diagram for explaining the global shutter method. This is a diagram schematically showing an example of a sampling pattern that can be realized in the global shutter method. This diagram schematically shows examples of sampling patterns that can be implemented in a global shutter system. This diagram provides a schematic explanation of image recognition processing using a CNN. This diagram provides a schematic explanation of image recognition processing that obtains recognition results from a portion of the image to be recognized. This diagram shows an example of reading out all lines in an image. This diagram shows an example of reading out lines in an image after decimation. This schematic diagram explains an example of a CNN divided into partitioning layers using existing technology. This schematic diagram explains an example of a CNN processing performed in parallel with reading from an imaging device using existing technology. This schematic diagram shows the configuration of an example of a CNN applicable to the first example of existing technology. This block diagram shows the configuration of an example of an information processing device using the first example of existing technology. This flowchart shows an example of processing by an information processing device in the first example of existing technology. This sequence diagram shows an example of processing by an information processing device in the first example of existing technology.This is a schematic diagram showing the configuration of an example CNN applicable to a second example of existing technology. This is a block diagram showing the configuration of an example information processing device according to a second example of existing technology. This is a flowchart showing an example of processing by an information processing device in a second example of existing technology. This is a sequence diagram showing an example of processing by an information processing device in a first example of existing technology. This is a schematic diagram for explaining the outline of an embodiment of this disclosure. This is a schematic diagram for explaining a segmented CNN model according to an embodiment. This is a schematic diagram showing the processing according to the first embodiment in comparison with processing according to existing technology. This is a sequence diagram showing an example of processing according to the first embodiment in comparison with processing according to existing technology. This is a schematic diagram showing the configuration of an example CNN applicable to the first embodiment. This is a block diagram showing the configuration of an example information processing device according to the first embodiment. This is a flowchart showing an example of processing by an information processing device according to the first embodiment. This is a sequence diagram showing an example of processing by an information processing device according to the first embodiment. This is a schematic diagram showing an example of the size of the feature maps of each layer relative to the output of the pixel array when YOLO is used as the object recognition algorithm. This is a schematic diagram showing an example of an output format when outputting a feature map from a sensor, applicable to the embodiment. This is a schematic diagram showing the configuration of an example CNN applicable to a first modification of the first embodiment. This is a block diagram showing the configuration of an example information processing device related to a first modification of the first embodiment. This is a flowchart showing an example of processing by the information processing device according to the first embodiment. This is a sequence diagram showing an example of processing by the information processing device according to the first embodiment. This is a schematic diagram showing the configuration of an example CNN applicable to a second modification of the first embodiment. This is a block diagram showing the configuration of an example information processing device related to a second modification of the first embodiment. This is a flowchart showing an example of processing by the information processing device related to a second modification of the first embodiment. This is a sequence diagram showing an example of processing by the information processing device according to a second modification of the first embodiment. This is a schematic diagram showing the processing according to the second embodiment in comparison with processing by existing technology. This is a sequence diagram showing an example of processing according to the second embodiment in comparison with processing by existing technology. This is a schematic diagram showing the configuration of an example CNN applicable to the second embodiment.This is a block diagram showing the configuration of an example of an information processing device according to the second embodiment. This is a flowchart showing an example of processing by the information processing device according to the second embodiment. This is a sequence diagram showing an example of processing by the information processing device according to the second embodiment. This is a schematic diagram showing the configuration of an example of a CNN applicable to a modified example of the second embodiment. This is a block diagram showing the configuration of an example of an information processing device according to a modified example of the second embodiment. This is a flowchart showing an example of processing by the information processing device according to the second embodiment. This is a sequence diagram showing an example of processing by the information processing device according to the second embodiment.

[0008] The embodiments of this disclosure will be described in detail below with reference to the drawings. In the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0009] The embodiments of this disclosure will be described below in the following order. 1. Technology applicable to the embodiments of this disclosure 1-1. Configuration applicable to the embodiments of this disclosure 1-2. Processing applicable to the embodiments of this disclosure 1-2-1. Imaging method 1-2-2. DNN 1-2-3. Driving speed 2. Existing technology 2-1. Outline of existing technology 2-2. Specific examples of existing technology 2-2-1. First example of existing technology 2-2-2. Second example of existing technology 3. Outline of the embodiments of this disclosure 4. First embodiment of this disclosure 4-1-1. Outline of the first embodiment 4-1-2. More specific description of the first embodiment 4-1-3. Visibility 4-2. First modification of the first embodiment 4-3. Second modification of the first embodiment 5. Second embodiment of this disclosure 5-1-1. Outline of the second embodiment 5-1-2. More specific description of the second embodiment 5-2. Modified form of the second embodiment

[0010] (1. Outline of Embodiments of the Disclosure) First, the embodiments of the Disclosure will be described in general terms. In the embodiments of the Disclosure, for example, an image captured by an imaging device is used as the input image, and a convolution process is performed on the input image using, for example, a CNN (Convolutional Neural Network) to generate a feature map.

[0011] In this embodiment, the CNN is divided between one layer and the next layer, and the divided preceding CNN (referred to as CNN#A) performs convolution on a line basis, and the processing results are transmitted line by line to the divided succeeding CNN (referred to as CNN#B). By dividing the CNN, for example, when CNN#A and the imaging device are configured as a single sensor, it is possible to reduce the processing load on both the sensor and the subsequent CNN#B.

[0012] Hereafter, the layer immediately preceding the CNN split point will be referred to as the splitting layer.

[0013] Figure 1 is a functional block diagram of an example illustrating the functions of an information processing device according to an embodiment. In the example of Figure 1, the information processing device 1 includes a sensor 10 and an AI (Artificial Intelligence) processing unit 210. The sensor 10 also includes an imaging device 100 and an AI processing unit 200. The AI ​​processing unit 200 takes the image captured and output by the imaging device 100 as an input image and performs CNN#A processing on a line basis. The AI ​​processing unit 200 outputs the processing results of CNN#A to the AI ​​processing unit 210 on a line basis. The AI ​​processing unit 210 performs CNN#B processing on the processing results output from the AI ​​processing unit 200 on a line basis.

[0014] As described above, in the information processing device 1 according to this embodiment, when the CNN is divided and processed, the preceding CNN#A performs line-based convolution processing and outputs the processing results to the subsequent CNN#B on a line-by-line basis. This allows the preceding CNN#A to output line-by-line feature maps in parallel with the reading of pixel signals from the imaging device 100. Furthermore, the subsequent CNN#B can perform subsequent processing in a short time after the reading from the imaging device 100, thereby shortening the overall processing time for the CNN.

[0015] Furthermore, the information processing device 1 can output the captured image captured by the imaging device 100 to the AI ​​processing unit 200, and can also output directly from the sensor 10 to the AI ​​processing unit 210. In this case, the AI ​​processing unit 210 may output the captured image output from the sensor 10 to the outside as is, or it may use it for processing by CNN#B.

[0016] (1-1. Configurations Applicable to Embodiments of the Disclosure) Next, configurations applicable to embodiments of the disclosure will be described.

[0017] Each processing unit that performs the above-mentioned processing in the information processing device 1 is implemented, for example, by a circuit. When the sensor 10 is implemented by a circuit in the information processing device 1, for example, the sensor 10 can be formed on a single substrate. Alternatively, for example, the sensor 10 may be a stacked CIS (CMOS (Complementary Metal Oxide Semiconductor) Image Sensor) in which multiple semiconductor chips are stacked and formed integrally.

[0018] As an example, the sensor 10 can be formed by a two-layer structure in which two semiconductor chips are stacked. Figure 2A shows an example in which the sensor 10 according to the embodiment is formed by a two-layer stacked CIS. In the structure of Figure 2A, a pixel section 20a is formed on the first layer semiconductor chip, and a memory + logic section 20b is formed on the second layer semiconductor chip. The pixel section 20a includes at least a pixel array in the imaging device 100. The memory + logic section 20b includes, for example, a line memory, a parameter memory, a sensor control section (not shown), and an interface for communication between the sensor 10 and the outside, together with the AI ​​processing unit 200. The memory + logic section 20b further includes part or all of the drive circuit that drives the pixel array in the imaging device 100.

[0019] As shown on the right side of Figure 2A, the sensor 10 is configured as a single solid-state image sensor (image sensor) 2a by bonding the first layer semiconductor chip and the second layer semiconductor chip together while making electrical contact.

[0020] As another example, the sensor 10 can be formed by a three-layer structure in which semiconductor chips are stacked in three layers. Figure 2B shows an example in which the sensor 10 according to the embodiment is formed by a three-layer stacked CIS. In the structure of Figure 2B, a pixel section 20a is formed on the first layer semiconductor chip, a memory section 20c is formed on the second layer semiconductor chip, and a logic section 20b' is formed on the third layer semiconductor chip. In this case, the logic section 20b' includes, for example, an AI processing unit 200, a line memory, a parameter memory, a sensor control unit (not shown), and an interface for communication between the sensor 10 and the outside.

[0021] As shown on the right side of Figure 2B, the sensor 10 is configured as a single solid-state image sensor 2b by bonding together the first layer semiconductor chip, the second layer semiconductor chip, and the third layer semiconductor chip while electrically contacting them.

[0022] Furthermore, some of the processing units of the sensor 10 shown in Figure 1 may be implemented using software (programs). For example, the AI ​​processing unit 200 may be implemented by having a processor such as a CPU (Central Processing Unit) or a DSP (Digital Signal Processor) execute a program.

[0023] The program executed by the sensor 10 of this embodiment is recorded in an installable or executable file format on a computer-readable storage medium such as a CD (Compact Disk), DVD (Digital Versatile Disk), or memory card, and provided as a computer program product.

[0024] Alternatively, the program executed by the sensor 10 of the embodiment may be stored on a computer connected to a communication network such as the Internet and provided by being downloaded via the communication network. Alternatively, the program executed by the sensor 10 of the embodiment may be provided via a communication network such as the Internet without requiring downloading.

[0025] Alternatively, the program for the sensor 10 of this embodiment may be pre-installed and provided in ROM (Read Only Memory) or the like.

[0026] Furthermore, when multiple processors are used to implement each processing unit, each processor may implement one processing unit or multiple processing units.

[0027] In the following, we will refer to neural networks simply as "networks" and networks used for communication, such as the Internet, as "communication networks" to distinguish between them.

[0028] Figure 3 is a block diagram showing the configuration of an example of an imaging device 100 applicable to the embodiment. In Figure 3, the imaging device 100 includes a pixel array unit 101, a vertical scanning unit 102, an AD (Analog to Digital) conversion unit 103, a pixel signal line 106, a vertical signal line VSL, a control unit 1100, and a signal processing unit 1101. Note that in Figure 3, the control unit 1100 and the signal processing unit 1101 may be included in, for example, the sensor control unit (not shown) described above.

[0029] The pixel array section 101 includes a plurality of pixels Pix, each containing a photoelectric conversion element, such as a photodiode, which performs photoelectric conversion on received light, and a circuit for reading the charge from the photoelectric conversion element. In the pixel array section 101, the plurality of pixels Pix are arranged in a matrix arrangement in the horizontal (row) and vertical (column) directions. In the pixel array section 101, the row arrangement of pixels Pix is ​​called a line. For example, when one frame image is formed with 3840 pixels × 2160 lines, the imaging device 100 includes at least 2160 lines, each containing at least 3840 pixels Pix. The pixel signals read from the pixels Pix included in the frame form one frame image (image data).

[0030] Hereinafter, the operation of reading pixel signals from each pixel Pix included in a frame in the imaging device 100 will be described as, for example, "reading pixels from a frame." Similarly, the operation of reading pixel signals from each pixel Pix of a line included in a frame will be described as, for example, "reading lines."

[0031] Furthermore, for each row and column of each pixel Pix in the pixel array unit 101, a pixel signal line 106 is connected, and a vertical signal line VSL is connected, for each column. The ends of the pixel signal lines 106 that are not connected to the imaging device 100 are connected to the vertical scanning unit 102. The vertical scanning unit 102 transmits control signals, such as drive pulses for reading pixel signals from pixels, to the pixel array unit 101 via the pixel signal lines 106, in accordance with the control of the control unit 1100, which will be described later. The ends of the vertical signal lines VSL that are not connected to the pixel array unit 101 are connected to the AD conversion unit 103. The pixel signals read from the pixels are transmitted to the AD conversion unit 103 via the vertical signal lines VSL.

[0032] This section provides a general overview of the control of reading out pixel signals from pixels (Pix). Reading out pixel signals from pixels involves transferring the charge accumulated in the photoelectric conversion element due to exposure to a floating diffusion layer (FD), and then converting the transferred charge into a voltage in the floating diffusion layer. The voltage converted from the charge in the floating diffusion layer is output to the vertical signal line VSL via an amplifier.

[0033] More specifically, in a pixel Pix, during exposure, the connection between the photoelectric conversion element and the floating diffusion layer is kept off (open), allowing the photoelectric conversion element to accumulate charge generated in response to the incident light through photoelectric conversion. After exposure ends, the floating diffusion layer and the vertical signal line VSL are connected according to a selection signal supplied via the pixel signal line 106. Furthermore, the floating diffusion layer is briefly connected to the power supply voltage VDD or black level voltage supply line in response to a reset pulse supplied via the pixel signal line 106, thereby resetting the floating diffusion layer. The vertical signal line VSL outputs a voltage corresponding to the reset level of the floating diffusion layer (let's call it voltage A). Subsequently, a transfer pulse supplied via the pixel signal line 106 turns the connection between the photoelectric conversion element and the floating diffusion layer on (closed), transferring the charge accumulated in the photoelectric conversion element to the floating diffusion layer. The vertical signal line VSL outputs a voltage corresponding to the amount of charge in the floating diffusion layer (let's call it voltage B).

[0034] The AD conversion unit 103 includes an AD converter 107 provided for each vertical signal line VSL, a reference signal generation unit 104, and a horizontal scanning unit 105. The AD converter 107 is a column AD converter that performs AD conversion processing for each column of the pixel array unit 101. The AD converter 107 performs AD conversion processing on the pixel signal supplied from the pixel Pix via the vertical signal line VSL and generates two digital values ​​(values ​​corresponding to voltage A and voltage B, respectively) for correlated double sampling (CDS) processing to reduce noise.

[0035] The AD converter 107 supplies the two generated digital values ​​to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values ​​supplied from the AD converter 107 to generate a pixel signal (pixel data) using digital signals. The pixel data generated by the signal processing unit 1101 is output to the outside of the imaging device 100.

[0036] Based on the control signal input from the control unit 1100, the reference signal generation unit 104 generates a ramp signal used by each AD converter 107 to convert the pixel signal into two digital values as a reference signal. The ramp signal is a signal whose level (voltage value) decreases at a constant slope with respect to time, or a signal whose level decreases in a stepwise manner. The reference signal generation unit 104 supplies the generated ramp signal to each AD converter 107. The reference signal generation unit 104 is configured using, for example, a DAC (Digital to Analog Converter).

[0037] When a ramp signal whose voltage drops stepwise according to a predetermined slope is supplied from the reference signal generation unit 104, the counter starts counting according to the clock signal. The comparator compares the voltage of the pixel signal supplied from the vertical signal line VSL with the voltage of the ramp signal, and stops the counting by the counter at the timing when the voltage of the ramp signal crosses the voltage of the pixel signal. The AD converter 107 converts the pixel signal by an analog signal into a digital value by outputting a value corresponding to the count value of the time when the counting is stopped.

[0038] The AD converter 107 supplies the two generated digital values to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values supplied from the AD converter 107 and generates a pixel signal (pixel data) by a digital signal. The pixel signal by the digital signal generated by the signal processing unit 1101 is output to the outside of the imaging device 100.

[0039] Under the control of the control unit 1100, the horizontal scanning unit 105 performs selection scanning to select each AD converter 107 in a predetermined order, so that each digital value temporarily held by each AD converter 107 is sequentially output to the signal processing unit 1101. The horizontal scanning unit 105 is configured using, for example, a shift register, an address decoder, or the like.

[0040] The control unit 1100 performs drive control of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, the horizontal scanning unit 105, etc., in accordance with the imaging control signal supplied from the sensor control unit 12. The control unit 1100 generates various drive signals that serve as the basis for the operations of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, and the horizontal scanning unit 105. The control unit 1100 generates, for example, a control signal for the vertical scanning unit 102 to supply to each pixel Pix via the pixel signal line 106, based on the vertical synchronization signal or external trigger signal included in the imaging control signal and the horizontal synchronization signal. The control unit 1100 supplies the generated control signal to the vertical scanning unit 102.

[0041] Also, the control unit 1100 passes, for example, information indicating the analog gain included in the imaging control signal supplied from the sensor control unit 12 to the AD conversion unit 103. The AD conversion unit 103 controls the gain of the pixel signal input to each AD converter 107 included in the AD conversion unit 103 via the vertical signal line VSL according to the information indicating this analog gain.

[0042] Based on the control signal supplied from the control unit 1100, the vertical scanning unit 102 supplies various signals including drive pulses to the pixel signal lines 106 of the selected pixel rows of the pixel array unit 101 to each pixel Pix line by line, and causes each pixel Pix to output a pixel signal to the vertical signal line VSL. The vertical scanning unit 102 is configured using, for example, a shift register or an address decoder. Also, the vertical scanning unit 102 controls the exposure at each pixel Pix according to the information indicating the exposure supplied from the control unit 1100.

[0043] The imaging device 100 configured in this way is a column AD type CMOS image sensor in which the AD converters 107 are arranged for each column.

[0044] Figure 4 is a block diagram showing the hardware configuration of an example of an information processing device 1 according to an embodiment. In Figure 4, the information processing device 1 includes a sensor 10, a DSP 1010, a CPU 1000, a ROM 1001, a RAM (Random Access Memory) 1002, a storage device 1003, a data I / F (interface) 1004, a communication I / F 1005, and a UI (User Interface) unit 1006, and each of these units is connected to each other via a bus 1020 so as to be able to communicate with each other.

[0045] The storage device 1003 is a non-volatile storage medium such as flash memory or a hard disk drive. The CPU 1000 controls the overall operation of the information processing device 1 using the RAM 1002 as work memory, according to the programs stored in the storage device 1003 and ROM 1001.

[0046] The data interface 1004 transmits and receives data to and from external devices via wired or wireless connection. The communication interface 1005 controls communication to communication networks such as the Internet and LAN (Local Area Network). The UI unit 1006 includes an input unit that allows user input operations and a display device that can present information to the user, providing an interface to the user by this information processing device 1.

[0047] As explained with reference to Figure 1, the sensor 10 includes an imaging device 100 and an AI processing unit 200. The sensor 10 takes an image using the imaging device 100 in response to light incident through the optical system 11, and uses the image obtained from the imaging as an input image to perform CNN#A processing by the AI ​​processing unit 200. The sensor 10 transmits the processing result from the AI ​​processing unit 200 to the DSP 1010.

[0048] The DSP 1010 has the functions of the AI ​​processing unit 210 shown in Figure 1, and performs CNN#B processing on the data transmitted from the sensor 10 in accordance with the instructions of the CPU 1000.

[0049] In the information processing device 1, for example, the CPU 1000 may acquire the above-mentioned CNN#A and CNN#B via a communication network or storage medium such as a CD (Compact Disk), DVD (Digital Versatile Disk), or memory card, and store them in the storage device 1003 or RAM 1002. For example, the sensor 10 and DSP 1010 may implement the functions of the above-mentioned AI processing units 200 and 210 using the CNN#A and CNN#B stored in these storage devices 1003 and RAM 1002.

[0050] Although the AI ​​processing unit 210 is described here as being configured on the DSP 1010, this is not limited to this example. For example, the AI ​​processing unit 210 may be configured on the CPU 1000.

[0051] (1-2. Processing Applicable to Embodiments of the Disclosure) Prior to describing embodiments relating to the Disclosure, the technologies applicable to the Disclosure will be briefly described in order to facilitate understanding.

[0052] (1-2-1. Imaging Method) Two imaging methods are known for imaging using the pixel array unit 101: the rolling shutter (RS) method and the global shutter (GS) method.

[0053] (Overview of Rolling Shutter) First, the rolling shutter method will be briefly explained. Figures 5A, 5B, and 5C are schematic diagrams illustrating the rolling shutter method. In the rolling shutter method, as shown in Figure 5A, imaging is performed line by line, starting from, for example, the upper line 501 of the frame 500.

[0054] For example, in the configuration shown in Figure 3, exposure is performed simultaneously at each pixel Pix included in a single line. After exposure is complete, the pixel signal based on the charge accumulated during exposure is simultaneously transmitted at each pixel Pix in that line via the corresponding vertical signal line VSL. By sequentially performing this operation line by line, rolling shutter imaging can be achieved.

[0055] Figure 5B schematically illustrates an example of the relationship between imaging and time in a rolling shutter system. In Figure 5B, the vertical axis represents the line position, and the horizontal axis represents time. In a rolling shutter system, exposure is performed sequentially for each line, so as shown in Figure 5B, the timing of exposure for each line is shifted sequentially according to the line position. Therefore, for example, if the horizontal positional relationship between the imaging device and the subject changes rapidly, distortion occurs in the image of the captured frame 500, as illustrated in Figure 5C. In the example in Figure 5C, the image 502 corresponding to frame 500 is tilted at an angle corresponding to the speed and direction of the change in the horizontal positional relationship between the imaging device and the subject.

[0056] In the rolling shutter method, it is also possible to perform imaging by skipping lines. Figures 6A, 6B, and 6C are schematic diagrams illustrating line skipping in the rolling shutter method. As shown in Figure 6A, similar to the example in Figure 5A described above, imaging is performed line by line from line 501 at the top of frame 500 toward the bottom of frame 500. At this time, imaging is performed while skipping lines at predetermined intervals.

[0057] For the sake of explanation, we will assume that imaging is performed every other line by thinning out the lines. That is, after imaging the k-th line, imaging of the (k+2)-th line will be performed. In this case, the time from imaging the k-th line to imaging the (k+2)-th line will be assumed to be equal to the time from imaging the k-th line to imaging the (k+1)-th line if thinning is not performed.

[0058] Figure 6B schematically shows an example of the relationship between imaging and time when one line is thinned in a rolling shutter system. In Figure 6B, the vertical axis represents the line position and the horizontal axis represents time. In Figure 6B, exposure A corresponds to the exposure in Figure 5B without thinning, and exposure B shows the exposure when one line is thinned. As shown in exposure B, by thinning the lines, the timing difference of exposure at the same line position can be shortened compared to when line thinning is not performed. Therefore, as exemplified as image 503 in Figure 6C, the tilt distortion that occurs in the image of the captured frame 500 is smaller compared to the case without line thinning shown in Figure 5C. On the other hand, when line thinning is performed, the image resolution is lower compared to when line thinning is not performed.

[0059] The above describes an example in which imaging is performed line by line from the top to the bottom of frame 500 in a rolling shutter system, but this is not the only example. Figures 7A and 7B schematically show examples of other imaging methods in a rolling shutter system. For example, as shown in Figure 7A, imaging can be performed line by line from the bottom to the top of frame 500 in a rolling shutter system. In this case, the horizontal direction of distortion in image 502 is reversed compared to when imaging is performed line by line from the top to the bottom of frame 500.

[0060] Furthermore, by setting the range of the vertical signal line VSL that transmits the pixel signals, for example, it is possible to selectively read out a portion of the line. In addition, by setting the lines on which imaging is performed and the vertical signal line VSL that transmits the pixel signals, it is possible to set the lines on which imaging begins and ends to be other than the upper and lower ends of the frame 500. Figure 7B schematically shows an example in which a rectangular region 505 whose width and height are less than the width and height of the frame 500, respectively, is used as the imaging range. In the example in Figure 7B, imaging is performed line by line from line 504 at the upper end of region 505 toward the lower end of region 505.

[0061] (Overview of Global Shutter) Next, the global shutter (GS) method will be briefly explained as an imaging method when imaging is performed by the imaging device 100. Figures 8A, 8B, and 8C are schematic diagrams for explaining the global shutter method. In the global shutter method, as shown in Figure 8A, exposure is performed simultaneously on all pixels Pix included in the frame 500.

[0062] In the configuration shown in Figure 3, when implementing a global shutter system, one example is to further provide a capacitor between the photoelectric conversion element and the FD in each pixel Pix. A first switch is provided between the photoelectric conversion element and the capacitor, and a second switch is provided between the capacitor and the floating diffusion layer. The opening and closing of these first and second switches are controlled by pulses supplied via the pixel signal line 106.

[0063] In this configuration, during the exposure period, the first and second switches are opened for all pixels (Pix) in frame 500, respectively. At the end of exposure, the first switch is closed, transferring charge from the photoelectric element to the capacitor. Subsequently, the capacitor is treated as a photoelectric element, and the charge is read from the capacitor in a sequence similar to the readout operation described in the rolling shutter method. This enables simultaneous exposure for all pixels (Pix) in frame 500.

[0064] Figure 8B schematically shows an example of the relationship between imaging and time in a global shutter system. In Figure 8B, the vertical axis represents the line position, and the horizontal axis represents time. In a global shutter system, exposure is performed simultaneously for all pixels (Pix) included in the frame 500, so the exposure timing for each line can be made the same, as shown in Figure 8B. Therefore, even if, for example, the horizontal positional relationship between the imaging device and the subject changes rapidly, the image 506 of the captured frame 500 does not exhibit distortion corresponding to this change, as illustrated in Figure 8C.

[0065] In the global shutter system, the simultaneity of exposure timing can be ensured for all pixels (Pix) included in the frame 500. Therefore, by controlling the timing of each pulse supplied by the pixel signal line 106 of each line and the timing of transfer by each vertical signal line VSL, sampling (readout of pixel signals) in various patterns can be realized.

[0066] Figures 9A and 9B schematically illustrate examples of sampling patterns that can be implemented in the global shutter method. Figure 9A shows an example in which samples 508, from which pixel signals are read out, are extracted in a checkerboard pattern from each pixel Pix arranged in a matrix within the frame 500. Figure 9B shows an example in which samples 508, from which pixel signals are read out, are extracted in a grid pattern from each of the same pixels Pix. Furthermore, in the global shutter method, imaging can be performed in line sequential order, similar to the rolling shutter method described above.

[0067] (1-2-2. About DNN) Next, we will briefly describe the recognition process using a DNN (Deep Neural Network) applicable to the embodiment. As described above, in the embodiment, a CNN is used to perform recognition processing on image data. Hereinafter, "recognition processing on image data" will be referred to as "image recognition processing" or the like as appropriate.

[0068] (Overview of CNN) First, let's briefly explain CNN. Generally, CNN image recognition processing is performed based on image information, such as pixels arranged in a matrix. Figure 10 is a diagram illustrating CNN image recognition processing in general terms. Processing is performed by a predetermined trained CNN 52 on the entire pixel information 51 of an image 50 in which the object to be recognized, the number "8", is drawn. As a result, the number "8" is recognized as the recognition result 53.

[0069] In contrast, it is also possible to perform CNN processing based on the image for each line and obtain recognition results from a portion of the image to be recognized. Figure 11 is a diagram illustrating in general terms this image recognition process that obtains recognition results from a portion of the image to be recognized. In Figure 11, image 50' is a partial acquisition of the number "8", which is the object to be recognized, on a line-by-line basis. Processing by a predetermined trained CNN 52' is sequentially applied to, for example, the line-by-line pixel information 54a, 54b, and 54c that form the pixel information 51' of this image 50'.

[0070] For example, the recognition result 53a obtained by CNN 52' in the recognition process for the pixel information 54a of the first line is considered not to be a valid recognition result. Here, a valid recognition result refers to, for example, a recognition result in which the confidence score for the recognized result is above a predetermined level. Based on this recognition result 53a, CNN 52' updates its internal state 55. Next, the CNN 52', whose internal state has been updated 55 based on the previous recognition result 53a, performs recognition processing on the pixel information 54b of the second line. In Figure 11, as a result, a recognition result 53b is obtained indicating that the number to be recognized is either "8" or "9". Furthermore, based on this recognition result 53b, CNN 52' updates its internal information 55. Next, the CNN 52', whose internal state has been updated 55 based on the previous recognition result 53b, performs recognition processing on the pixel information 54c of the third line. In Figure 11, as a result, the number to be recognized is narrowed down to "8" from among "8" and "9".

[0071] In this recognition process shown in Figure 11, the internal state of the CNN is updated using the results of the previous recognition process. The updated CNN then performs recognition using the pixel information of lines adjacent to the line that was previously recognized. In other words, the recognition process shown in Figure 11 is executed line by line on the image, updating the internal state of the CNN based on the previous recognition results. Therefore, the recognition process shown in Figure 11 is a recursive process executed line by line, and can be considered to have a structure equivalent to an RNN (Recurrent Neural Network).

[0072] (1-2-3. Regarding Driving Speed) Next, the relationship between the frame driving speed and the amount of pixel signal readout will be explained using Figures 12A and 12B. Figure 12A shows an example of reading out all lines in an image. Here, the resolution of the image to be processed for recognition is assumed to be 640 horizontal pixels × 480 vertical pixels (480 lines). In this case, driving at a driving speed of 14400 (lines / second) makes it possible to output at 30 fps (frames per second).

[0073] Next, let's consider imaging with line decimation. For example, as shown in Figure 12B, imaging is performed using 1 / 2 decimation readout, where one line is skipped during reading. As a first example of 1 / 2 decimation, if the drive speed is 14400 (lines / second) as described above, the number of lines read from the image is halved, so the resolution decreases, but output at 60 fps, twice the speed of when decimation is not performed, is possible, improving the frame rate. As a second example of 1 / 2 decimation, if the drive speed is set to 7200 fps, half of the first example, the frame rate will be 30 fps, the same as when decimation is not performed, but power saving is possible.

[0074] When reading lines from an image, whether to perform decimation, perform decimation to increase the drive speed, or perform decimation while maintaining the same drive speed as when decimation is not performed can be selected depending on the purpose of the recognition processing based on the read pixel signals, for example.

[0075] (2. Existing Technology) Next, we will describe existing technology related to the embodiments of this disclosure.

[0076] (2-1. Overview of Existing Technologies) Existing technologies will be briefly explained. Figure 13 is a schematic diagram illustrating an example of a CNN divided into division layers using existing technologies. Here, the processing up to the division layer in the CNN is referred to as process #A, and the processing after the division layer is referred to as process #B.

[0077] In Figure 13, section (a) shows an example where CNN processing is performed downstream of the sensor 10 without splitting the CNN. In the example in section (a), processing #A and processing #B are performed by the CNN downstream of the sensor 10 on the readout data of one frame image from the imaging device 100 at the sensor 10.

[0078] Section (b) of Figure 13 shows an example where the CNN is divided into CNN#A and CNN#B (also referred to as divided CNN#A and divided CNN#B, respectively) by a splitting layer, CNN#A is executed in the AI ​​processing unit 200 within the sensor 10, and CNN#B is executed in the AI ​​processing unit 210 outside the sensor 10. Processing #A is executed by divided CNN#A within the sensor 10 for the readout data of one frame image from the imaging device 100. The processing result of processing #A is output from divided CNN#A as a feature map for the image and input to the subsequent divided CNN#B. Divided CNN#B executes processing #B on the input feature map and outputs the result.

[0079] In the case of the processing in section (b), if the computational performance of the preceding AI processing unit 200 is not higher than that of the succeeding AI processing unit 210, the output of the final processing result will be slower compared to the case where the CNN shown in section (a) is not divided, as shown as a delay in the figure. In other words, if the succeeding AI processing unit 210 has higher computational performance than the preceding AI processing unit 200, the overall processing will be faster if both processing #A and #B are entrusted to the succeeding AI processing unit 210.

[0080] Figure 14 is a schematic diagram illustrating an example of performing CNN processing in parallel with readout from the imaging device 100 using existing technology. The example in Figure 14 shows an example of performing CNN on a line basis. In this case, as shown in section (a), the imaging device 100 reads out image data line by line in the sensor 10 in synchronization with the horizontal synchronization signal XHS. The AI ​​processing unit 200 in the sensor 10 performs CNN processing line by line on the line-by-line image data read out from the imaging device 100 in synchronization with the horizontal synchronization signal XHS.

[0081] In this case, if the CNN processing speed by the AI ​​processing unit 200 is slow, it becomes necessary to lengthen the time interval of the horizontal synchronization signal XHS to match the speed of the CNN processing, thereby lengthening the interval for reading lines from the imaging device 100. In this case, the entire processing will be delayed.

[0082] To avoid this processing delay, for example, as shown in section (b) of Figure 14, the CNN can be made lighter, and the AI ​​processing unit 200 in the sensor 10 can perform simple tasks that require less computation. With this method, as shown in section (c) of Figure 14, it becomes possible to read out lines from the imaging device 100 and perform CNN processing on the read-out lines by the AI ​​processing unit 200 at the interval of the original horizontal synchronization signal XHS.

[0083] However, because this method reduces the complexity of CNN processing, there is a risk that the accuracy of recognition processing using CNN on image data, for example, may decrease.

[0084] (2-2. Specific Examples of Existing Technologies) Examples of CNN processing using existing technologies will be explained in more detail.

[0085] (2-2-1. First example of existing technology) The first example of existing technology is an example in which the CNN is not divided and the CNN processing is performed at a later stage, i.e., outside of the sensor 10.

[0086] Figure 15 is a schematic diagram showing the configuration of an example of a CNN applicable to the first example of existing technology. In the following, the CNN is assumed to have four intermediate layers, from layer 1 to layer 4, and each layer performs convolution processing using a 3-pixel × 3-line filter. In addition, 9-pixel × 9-line image data is assumed to be input to the input layer. In the following, "m pixels × m lines" will be written as "m × m" as appropriate.

[0087] In Figure 15, the CNN generates 7x7 image data in each of the three channels (ch) in the first layer by convolution using a 3x3 filter on the 9x9 image data input to the input layer. These three channels of image data are feature maps generated by the first layer. The CNN generates 5x5 image data in each of the five channels (5x7) in the second layer by convolution using a 3x3 filter on the 7x7 image data of the first layer. These five channels of image data are feature maps generated by the second layer. The CNN generates 3x3 image data in each of the seven channels (5x5) in the third layer by convolution using a 3x3 filter. These seven channels of image data are feature maps generated by the third layer. Next, the CNN generates 1x1 image data in each of the five channels (5x5) in the fourth layer by convolution using a 3x3 filter on the seven channels (3x3) image data of the third layer. These 5-channel image data are feature maps based on the fourth layer.

[0088] Although not shown in the diagram, the 5-channel image data (feature map) in the fourth layer is combined by the fully connected output layer, and the inference result is output from the CNN.

[0089] In Figure 15, the processing from the first to the fourth layer corresponds to the execution range of the subsequent processing on the sensor 10.

[0090] Figure 16 is a block diagram showing the configuration of an example of an information processing device 1a according to a first example of existing technology. In Figure 16, the sensor 10a includes a pixel array unit 101 and a data path 110 to which image data read from the pixel array unit 101 is transferred. In Figure 16 and subsequent similar figures, the imaging device 100 in Figure 1 is represented by the pixel array unit 101.

[0091] In the first example, the sensor 10a does not perform CNN processing. Therefore, the image data read from the pixel array unit 101 by the sensor 10a is output to the subsequent processing unit 300a via the data path 110.

[0092] The downstream processing unit 300a performs downstream processing on the sensor 10a and includes frame buffers 301a and 301b, and an AI engine 310, which are connected to each other via a bus 320 for communication. In Figure 16 and Figure 18 described later, frame buffers 301a and 301b are also shown as frame buffer #A and #B, respectively. The AI ​​engine 310 corresponds to the AI ​​processing unit 210 shown in Figure 1.

[0093] Furthermore, the operation of the parts shown in Figure 16, and later in Figures 20, 28, 34, 38, 44, and 48, is controlled by a control unit (not shown) included in the information processing device 1 (information processing devices 1a to 1g). This control unit may be configured, for example, by executing a predetermined program on the CPU 1000, or it may be configured by a hardware circuit.

[0094] Figure 17 is a flowchart illustrating an example of processing by the information processing device 1a in a first example of the existing technology. Figure 18 is a sequence diagram illustrating an example of processing by the information processing device 1a in a first example of the existing technology. Each process shown in the sequence in Figure 18 corresponds to each process in the flowchart in Figure 17.

[0095] In Figure 18, XVS represents the vertical synchronization signal, and XHS represents the horizontal synchronization signal. "Writing" indicates the period during which image data is being written to frame buffer 301a or 301b, and "Readable" indicates the period during which image data can be read from frame buffer 301a or 301b. Furthermore, "Processing" indicates the period during which processing by the AI ​​engine 310 is being performed.

[0096] In the following, we will assume that one frame consists of 9 pixels x 9 lines.

[0097] In Figure 17, in step S100, the information processing device 1a reads out one line of image data from the pixel array unit 101 (time t in Figure 18). 100). In the next step S101, the information processing device 1a writes the image data of one line read in step S100 to the frame buffer 301a. The writing of one line of image data to the frame buffer 301a is performed cumulatively, that is, without erasing or overwriting previously written data, until the writing of one frame's worth of lines is completed.

[0098] In the next step S102, the information processing device 1a determines whether or not the writing of nine lines, i.e., one frame's worth of image data, to the frame buffer 301a is complete. If the information processing device 1a determines that it is not complete (step S102, "No"), it returns to step S100 and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1a determines that the writing of nine lines' worth of image data is complete (step S102, "Yes"), it moves the process to step S110 (time t in Figure 18). 101 ).

[0099] Step S110 onward constitutes the subsequent execution range in which processing by the subsequent processing unit 300a is performed.

[0100] In step S110, the information processing device 1a reads one frame's worth of image data from the frame buffer 301a. In the next step S111, the information processing device 1a uses the AI ​​engine 310 to perform a convolution process (also described as conv in the diagram) using a 3x3 filter in the first layer on the one frame's worth of image data read from the frame buffer 301a, and outputs 7-line x 3-channel (7 pixels x 7 lines x 3 channels) image data. In the next step S112, this 7-line x 3-channel image data is written to the frame buffer 301b.

[0101] In the next step S120, the information processing device 1a reads 7-line x 3-channel image data from the frame buffer 301b. In the next step S121, the AI ​​engine 310 performs a convolution process using a second layer of 3x3 filters on the read image data, outputting 5-line x 5-channel (5 pixels x 5 lines x 5 channels) image data. In the next step S122, this 5-line x 5-channel image data is written to the frame buffer 301a.

[0102] In the next step S130, the information processing device 1a reads 5-line x 5-channel image data from the frame buffer 301a. In the next step S131, the AI ​​engine 310 performs a convolution process on the read image data using a third layer of 3x3 filters, outputting 3-line x 7-channel (3 pixels x 3 lines x 7 channels) image data. In the next step S132, this 3-line x 7-channel image data is written to the frame buffer 301b.

[0103] In the next step S140, the information processing device 1a reads 3-line x 7-channel image data from the frame buffer 301b. In the next step S141, the AI ​​engine 310 performs a convolution process using a fourth layer of 3x3 filters on the read image data, outputting 1-line x 5-channel (1 pixel x 1 line x 5 channels) image data. In the next step S142, this 1-line x 5-channel image data is written to the frame buffer 301a.

[0104] Although not shown in the diagram, the information processing device 1a combines the 1-line × 5-channel image data written to the frame buffer 301a in step S142 using the AI ​​engine 310 in the output layer and outputs it as, for example, an inference result.

[0105] (2-2-2. Second example of existing technology) The second example of existing technology is one in which the CNN is divided into a partitioned layer, the processing by the first stage CNN#A is executed by the AI ​​engine inside the sensor 10, and the processing by the second stage CNN#B is executed by the AI ​​engine outside the sensor 10.

[0106] Figure 19 is a schematic diagram showing the configuration of an example of a CNN applicable to the second example of existing technology. The CNN configuration itself is the same as the configuration described using Figure 15 above, so its explanation is omitted here. In Figure 19, the processing of the third and fourth layers constitutes the execution range of subsequent processing on the sensor 10.

[0107] Figure 20 is a block diagram showing the configuration of an example of an information processing device 1b according to a second example of existing technology. In Figure 20, the sensor 10b includes a pixel array unit 101, frame buffers 122a and 122b, an AI engine 120, and a switch 123b. The frame buffers 122a, 112b, and the AI ​​engine 120 are connected to each other via a bus 121 so that they can communicate with one another.

[0108] In Figure 20 and Figure 22 (described later), frame buffers 122a and 122b are also shown as frame buffer #a and #b, respectively. Furthermore, the AI ​​engine 120 corresponds to the AI ​​processing unit 200 shown in Figure 1.

[0109] Image data read from the pixel array unit 101 is transferred to the frame buffer 122a via the data path 110. Image data read from the frame buffer 122b is input to terminal "1" of switch 123b. Switch 123b selects terminal "1" during the vertical blanking period (V blank) and terminal "0" during other periods. The path through the selected terminal is closed (on). The output of switch 123b is transferred to the frame buffer 301a of the downstream processing unit 300a.

[0110] The configuration of the subsequent processing unit 300a is the same as that of the subsequent processing unit 300a described with reference to Figure 16, so its description is omitted here.

[0111] Figure 21 is a flowchart illustrating an example of processing by the information processing device 1b in a second example of the existing technology. Figure 22 is a sequence diagram illustrating an example of processing by the information processing device 1b in a first example of the existing technology. Each process shown in the sequence in Figure 22 corresponds to each process in the flowchart in Figure 21.

[0112] The meaning of each part in Figure 22 is the same as that of each part in Figure 18 described above, so an explanation will be omitted here. Also, as mentioned above, one frame is assumed to be 9 pixels × 9 lines.

[0113] In Figure 21, in step S100a, the information processing device 1b reads out one line of image data from the pixel array unit 101 (time t 100 ). In the next step S101a, the information processing device 1b writes the image data of one line read in step S100a to the frame buffer 122a. The writing of one line of image data to the frame buffer 122a is performed cumulatively until the writing of one frame's worth of lines is completed.

[0114] In the next step S102a, the information processing device 1b determines whether or not the writing of nine lines of image data to the frame buffer 122a is complete. If the information processing device 1b determines that it is not complete (step S102a, "No"), it returns to step S100a and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1b determines that the writing of nine lines of image data is complete (step S102a, "Yes"), it moves the process to step S110a (time t 101 ).

[0115] In step S110a, the information processing device 1b reads one frame's worth of image data from the frame buffer 122a. In the next step S111a, the information processing device 1b uses the AI ​​engine 120 to perform a convolution process (also described as conv in the diagram) using a 3x3 filter in the first layer on the one frame's worth of image data read from the frame buffer 122a, and outputs 7-line x 3-channel (7 pixels x 7 lines x 3 channels) image data. In the next step S112a, this 7-line x 3-channel image data is written to the frame buffer 122b.

[0116] In the next step S120a, the information processing device 1b reads 7x7x3ch image data from the frame buffer 122b. In the next step S121a, the AI ​​engine 120 performs a convolution process using a second layer 3x3 filter on the read image data, outputting 5-line x 5-ch (5 pixels x 5 lines x 5 channels) image data. In the next step S122a, this 5-line x 5-ch image data is written to the frame buffer 301a of the downstream processing unit 300a.

[0117] Step S130a onward constitutes the subsequent execution range in which processing by the subsequent processing unit 300a is executed.

[0118] In the next step S130a, the information processing device 1b reads 5x5x5ch image data from the frame buffer 301a. In the next step S131a, the AI ​​engine 310 performs a convolution process on the read image data using a third layer of 3x3 filters, outputting 3-line x 7-channel (3 pixels x 3 lines x 7 channels) image data. In the next step S132a, this 3-line x 7-channel image data is written to the frame buffer 301b.

[0119] In the next step S140a, the information processing device 1b reads 3-line x 7-channel image data from the frame buffer 301b. In the next step S141a, the AI ​​engine 310 performs a convolution process on the read image data using a fourth layer of 3x3 filters, outputting 1-line x 5-channel (1 pixel x 1 line x 5 channels) image data. In the next step S142, this 1-line x 5-channel image data is written to the frame buffer 301a.

[0120] Although not shown in the diagram, the information processing device 1b uses the AI ​​engine 310 to combine the 1-line × 5-channel image data written to the frame buffer 301a in step S142a at the output layer and outputs it, for example, as an inference result.

[0121] As shown in Figure 22, the processing times for steps S110a to S122a are set to be longer, assuming that the processing speed of the sensor 10b is slower than the processing speed of the subsequent processing unit 300a.

[0122] (3. Outline of Embodiments of the Disclosure) Next, embodiments of the disclosure will be described in general terms.

[0123] Figure 23 is a schematic diagram illustrating an overview of an embodiment of the present disclosure. In Figure 23, the upper section schematically shows the processing by the CNN model 60 before segmentation. In the sensor, the captured image output from the imaging device 100 (pixel array unit 101) is input to the CNN model 60 as an input image. The CNN model 60 performs predetermined processing, such as convolution processing, on the input image using predetermined parameters and outputs an inference result based on the input image.

[0124] In Figure 23, the lower section schematically shows the processing when the CNN model 60, according to the embodiment, is divided into a sensor CNN 60a and a subsequent CNN 60b at the division layer. In this case, the parameters used in the CNN model 60 up to the division layer are applied to the sensor CNN 60a. Also, the parameters used in the CNN model 60 from the division layer onward are applied to the subsequent CNN 60b.

[0125] Sensor CNN 60a is a CNN mounted on sensor 10 and corresponds to CNN#A shown in Figure 1. Furthermore, the subsequent CNN 60b is a CNN mounted in the processing unit downstream of sensor 10 and corresponds to CNN#B shown in Figure 1.

[0126] The sensor CNN 60a performs CNN processing, such as line-based convolution, on the input image and outputs a line-by-line feature map. By performing line-based CNN processing in the sensor CNN 60a, it is possible to perform the reading of image data from the imaging device 100 and the CNN processing in the sensor CNN 60a in parallel.

[0127] The subsequent CNN 60b performs CNN processing on the feature map output line by line from the sensor CNN 60a, and outputs the processing result as an inference result.

[0128] Figure 24 is a schematic diagram illustrating a segmented CNN model according to the embodiment. In Figure 24, section (a) shows an example of the CNN model 60 before segmentation. In this example, the CNN model 60 includes an input layer, a fully connected layer (output layer), and four intermediate layers. In each of the four intermediate layers (layers 1 to 4), the CNN model 60 generates a feature map based on the output of the preceding layer. The CNN model 60 combines the feature maps generated in the fourth layer in the fully connected layer and outputs an inference result based on the input image.

[0129] In Figure 24, section (b) shows an example where the CNN model 60 is divided into a sensor CNN 60a and a subsequent CNN 60b. Processing by the sensor CNN 60a takes place within the sensor 10. In this example, the CNN model 60 is divided using the second of the four intermediate layers as the dividing layer. The sensor CNN 60a includes the first and second layers. In the first layer, the sensor CNN 60a performs line-based CNN processing on the input image to generate the first layer feature map. In the second layer, the sensor CNN 60a performs line-based CNN processing on the feature map generated in the first layer to generate the second layer feature map.

[0130] The sensor CNN 60a outputs the second-layer feature map, generated by the second-layer CNN processing, on a line-by-line basis. The subsequent CNN 60b takes the second-layer feature map output from the sensor CNN 60a as its input and performs the third-layer CNN processing to generate a feature map. In the fourth layer, the subsequent CNN 60b performs line-based CNN processing on the feature map generated in the third layer to generate the fourth-layer feature map.

[0131] The subsequent CNN60b combines the feature maps generated in the fourth layer using a fully connected layer and outputs inference results based on the input image.

[0132] Since the sensor CNN 60a performs CNN processing on a line basis, it can output line-level feature maps in parallel with the reading of image data from the imaging device 100 (pixel array unit 101). Therefore, the subsequent CNN 60b can continue processing the CNN processing performed by the sensor CNN 60a immediately after the reading of image data from the imaging device 100 (pixel array unit 101).

[0133] Furthermore, the division point of the CNN model 60 may be between the N layer (including the input layer) and the (N+1) layer, provided that N is an integer of 2 or more and the total number of layers, including the input and output layers, is (N+2) or more.

[0134] (4. First Embodiment of the Disclosure) Next, a first embodiment of the Disclosure will be described. The first embodiment is an example in which, when the CNN is divided into divided CNN#A within the sensor 10 and divided CNN#B which performs CNN processing on the sensor 10 in a subsequent stage, CNN processing is performed on a line basis in divided CNN#A and CNN processing is performed on a frame basis in divided CNN#B.

[0135] (4-1-1. Outline of the First Embodiment) First, the first embodiment will be described in general terms using Figures 25 and 26. Figure 25 is a schematic diagram showing the process according to the first embodiment in comparison with the process according to existing technology. Sections (a) and (b) of Figure 25 are reproduced from sections (a) and (b) of Figure 13 described above for comparison, so a detailed explanation will be omitted here.

[0136] Section (c) of Figure 25 shows an example of processing according to the first embodiment. In the segmented CNN #A within the sensor 10, processing #A is executed on a line-based basis on the image data read from the imaging device 100 (pixel array unit 101). Therefore, it is possible to read the image data from the imaging device 100 and perform CNN processing by segmented CNN #A in parallel. On the other hand, the conventional segmented CNN #A shown in section (b) of Figure 25 performs CNN processing on a frame-based basis, so the subsequent segmented CNN #B had to wait for the completion of CNN processing for one frame by segmented CNN #A before starting its own CNN processing.

[0137] Therefore, by applying the first embodiment, it is possible to shorten the latency from the readout of the imaging device 100 to the output of the result of process #B, compared to the configuration using existing technology.

[0138] Figure 26 is an example sequence diagram showing the process according to the first embodiment in comparison with a process using existing technology.

[0139] In Figure 26, section (a) shows a sequence diagram of an example of processing using existing technology. The sequence diagram shown in section (a) of Figure 26 corresponds to Figures 16 to 18 and section (a) of Figure 25 described above. The explanation will be given with reference to Figure 16 as appropriate. The sensor 10a (pixel array section 101) is assumed to read out pixel signals using a rolling shutter method.

[0140] In section (a) of Figure 26, the reading of image data from the pixel array section 101 of the sensor 10a is started in synchronization with the vertical synchronization signal XVS, and the sensor reading data is written to the frame buffer 301a (frame buffer #A) of the downstream processing unit 300a. When one frame's worth of image data is written to the frame buffer 301a, the AI ​​engine 310 reads the image data from the frame buffer 301a and executes process A#, and then executes process #B.

[0141] Synchronized with the next vertical synchronization signal XVS, the reading of image data from the pixel array unit 101 of the sensor 10a begins. If processing #B by the AI ​​engine 310 is being executed at this point, the image data read from the pixel array unit 101 is written to the frame buffer 301b. The AI ​​engine 310 reads the image data from the frame buffer 301b, executes processing A#, and then executes processing #B.

[0142] Section (b) of Figure 26 shows a sequence diagram of an example of processing according to the first embodiment. The sensor 10 includes an AI engine #A that executes processing #A using CNN #A, and a line buffer capable of storing a number of lines corresponding to the filters used for CNN processing. The downstream processing unit after the sensor 10 includes an AI engine #B that executes processing #B using CNN #B, and a frame buffer for storing feature maps.

[0143] In section (b) of Figure 26, in synchronization with the vertical synchronization signal XVS, the reading of image data from the pixel array section 101 of the sensor 10 is started, and the read image data (sensor read data) is written to the line buffer of the sensor 10. For example, if process #A performs CNN processing using a 3x3 filter, when three lines of image data are stored in the line buffer, the AI ​​engine #A reads the three lines of image data from the line buffer and executes process #A.

[0144] AI engine #A sends the image data (feature map) resulting from process #A to the subsequent processing unit. The subsequent processing unit stores the image data sent from AI engine #A in its frame buffer. In the subsequent processing unit, when one frame's worth of image data (feature map) is stored in the frame buffer, AI engine #B executes process #B.

[0145] If this process #B can be completed by the next vertical synchronization signal XVS, the processing by the sensor 10 can be started in response to the next vertical synchronization signal XVS. Thus, in the first embodiment, if the latency from reading image data from the pixel array unit 101 to outputting the result of process #B is short, the subsequent processing unit can use only one frame buffer to process each frame.

[0146] (4-1-2. More Specific Description of the First Embodiment) Next, the first embodiment will be described in more detail.

[0147] Figure 27 is a schematic diagram showing an example configuration of a CNN applicable to the first embodiment. The configuration of the CNN itself is the same as the configuration described using Figure 15 above, so its explanation is omitted here. In Figure 27, the processing of the third and fourth layers constitutes the execution range of subsequent processing on the sensor 10.

[0148] Figure 28 is a block diagram showing the configuration of an example of an information processing device 1c according to the first embodiment. In Figure 28, the sensor 10c includes a pixel array unit 101, line buffers 124a and 124b, a frame buffer 122c, an AI engine 120, and a switch 123b. The line buffers 124a and 124b, the frame buffer 122c, and the AI ​​engine 120a are connected to each other via a bus 121 so that they can communicate with one another.

[0149] Line buffers 124a and 124b are each capable of storing at least a number of lines corresponding to the size of the filter used by the AI ​​engine 120a in CNN processing. In the following, it is assumed that the AI ​​engine 120a performs convolution processing using a 3x3 filter, and line buffers 124a and 124b are each capable of storing at least 3 lines of image data.

[0150] Furthermore, the size of the filter applicable to the AI ​​engine 120a is not limited to 3x3. For example, larger filters such as 5x5 or 7x7 may be applied to the AI ​​engine 120a. The same applies to the AI ​​engine 310 in the subsequent processing unit 300a.

[0151] In Figure 28 and Figure 30 (described later), line buffers 124a and 124b are also shown as line BF#a and line BF#b, respectively. Frame buffer 122c is also shown as frame buffer #c. The AI ​​engine 120 corresponds to the AI ​​processing unit 200 shown in Figure 1.

[0152] Image data read from the pixel array unit 101 is transferred to the line buffer 124a via the data path 110. Image data read from the line buffer 124a is transferred to the AI ​​engine 120a via the bus 121. Image data read from the frame buffer 122c is input to terminal "1" of switch 123b. Terminal "1" of switch 123b is selected during the vertical blanking period (V blank), and terminal "0" is selected during other periods. The output of switch 123b is transferred to the frame buffer 301a of the downstream processing unit 300a.

[0153] The configuration of the subsequent processing unit 300a is the same as that of the subsequent processing unit 300a described with reference to Figure 16, so its description is omitted here.

[0154] Figure 29 is a flowchart illustrating an example of processing by the information processing device 1c according to the first embodiment. Figure 30 is a sequence diagram illustrating an example of processing by the information processing device 1c according to the first embodiment. Each process shown in the sequence in Figure 30 corresponds to each process in the flowchart in Figure 29.

[0155] The meaning of each part in Figure 29 is the same as that of each part in Figure 18 described above, so an explanation will be omitted here. Also, as mentioned above, one frame is assumed to be 9 pixels × 9 lines.

[0156] In Figure 29, in step S100b, the information processing device 1c reads out one line of image data from the pixel array unit 101 (time t 100 ). In the next step S101b, the information processing device 1c writes the image data of one line read in step S100b to the line buffer 124a. The writing of one line of image data to this line buffer 124a is performed cumulatively until the writing of three lines is completed.

[0157] In Figure 30, line buffer #a(1), line buffer #a(2), and line buffer #a(3) each represent line-by-line writing to line buffer 124a.

[0158] In the next step S102b, the information processing device 1c determines whether the writing of three lines of image data to the line buffer 124a is complete. If the information processing device 1c determines that it is not complete (step S102, "No"), it returns to step S100b and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1c determines that the writing of three lines of image data is complete (step S102b, "Yes"), it moves the process to step S150 (time t 101 ).

[0159] In step S150, the information processing device 1c reads three lines of image data from the line buffer 124a using the AI ​​engine 120a. The AI ​​engine 120a then performs a convolution process (also described as conv in the diagram) using a 3x3 filter in the first layer on the read three lines of image data, outputting 1 line x 3 channel image data. This 1 line x 3 channel image data is written to the line buffer 124b in the next step S151. The writing of 1 line x 3 channel image data to the line buffer 124b is performed cumulatively until the writing of 3 lines x 3 channels is completed.

[0160] In Figure 30, line buffer #b(1), line buffer #b(2), and line buffer #b(3) each represent line-by-line writing to line buffer 124b.

[0161] In the next step, S152, the information processing device 1c determines whether the writing of 3 lines x 3 channels to the line buffer 124b has been completed. If the information processing device 1c determines that the writing has not been completed (step S152, "No"), it returns to step S100b. On the other hand, if the information processing device 1c determines that the writing has been completed (step S152, "Yes"), it moves the process to step S160.

[0162] In step S160, the information processing device 1c reads three lines of image data from the line buffer 124b using the AI ​​engine 120, and performs a convolution process using a second layer 3x3 filter on the read three lines of image data using the AI ​​engine 120, outputting 1 line x 5 channel image data. This 1 line x 5 channel image data is written to the frame buffer 122c in the next step S161. The writing of this 1 line x 5 channel image data to the frame buffer 122c is performed cumulatively until the writing of 5 lines x 5 channels is completed (second layer in Figure 27).

[0163] In the next step, S162, the information processing device 1c determines whether or not the writing of 5-line x 5-channel image data (feature map) to the frame buffer 122c is complete. If the information processing device 1c determines that the writing is not complete (step S162, "No"), it returns to step S100b. On the other hand, if the information processing device 1c determines that the writing is complete (step S162, "Yes"), it moves the process to step S163.

[0164] In step S163, the information processing device 1c waits for the vertical blanking period (V blank) in the readout process of the pixel array unit 101. When the vertical blanking period arrives, terminal "1" is selected at switch 123b, and the device becomes ready to transfer image data from frame buffer 122c to frame buffer 301a in the subsequent processing unit 300a.

[0165] In the next step S164, the information processing device 1c reads the 5-line x 5-channel image data (feature map) written in the processing up to step S162 from the frame buffer 122c and outputs it from the sensor 10c. In the next step S165, the information processing device 1c writes the 5-line x 5-channel (5 pixels x 5 lines x 5 channels) image data output from the sensor 10c to the frame buffer 301a.

[0166] After the processing in step S165, the information processing device 1c proceeds to step S130b. Steps S130b and beyond constitute the subsequent execution range in which processing by the subsequent processing unit 300a is performed.

[0167] In step S130b, the information processing device 1c reads 5-line x 5-channel image data from the frame buffer 301a. In the next step S131b, the AI ​​engine 310 performs a convolution process using a third layer of 3x3 filters on the read image data, outputting 3-line x 7-channel (3 pixels x 3 lines x 7 channels) image data. In the next step S132b, this 3-line x 7-channel image data is written to the frame buffer 301b.

[0168] In the next step S140b, the information processing device 1c reads 3-line x 7-channel image data from the frame buffer 301b. In the next step S141b, the AI ​​engine 310 performs a convolution process on the read image data using a fourth layer of 3x3 filters, outputting 1-line x 5-channel (1 pixel x 1 line x 5 channels) image data. In the next step S142b, this 1-line x 5-channel image data is written to the frame buffer 301a.

[0169] Although not shown in the diagram, the information processing device 1c combines the 1-line × 5-channel image data written to the frame buffer 301a in step S142b using the AI ​​engine 310 in the output layer and outputs it as, for example, an inference result.

[0170] (4-1-3. Visibility) The feature map output when the CNN is divided and processed is image data. Therefore, the feature map may be added to the captured image output from the imaging device 100 and transferred to the subsequent processing unit, or it may be read by the subsequent processing unit via a register or the like. In either case, since some image data is output from the sensor 10, visibility can be ensured.

[0171] In the above explanation, the input image size is assumed to be 9 pixels x 9 lines, but the actual input image and the image data (feature maps) of each layer of the CNN are much larger. Figure 31 is a schematic diagram showing an example of the size of the feature maps of each layer relative to the output of the pixel array 101 when YOLO (You Only Look Once) is used as the object recognition algorithm.

[0172] Assume that the total number of pixels in the image data 70 output from the image sensor (imaging device 100) is 2160 lines × 3840 pixels = 8,294,400 pixels. The feature map output after convolution processing using a 3x3 filter in the first layer has a total number of pixels of 112 lines × 112 pixels × 192 channels = 2,408,448 pixels. Similarly, in the second layer, the total number of pixels is 56 lines × 56 pixels × 256 channels = 802,816 pixels, in the third layer, the total number of pixels is 28 lines × 28 pixels × 512 channels = 401,408 pixels, in the fourth layer, the total number of pixels is 14 lines × 14 pixels × 1024 channels = 200,704 pixels, and in the fifth layer, the total number of pixels is 7 lines × 7 pixels × 1024 channels = 50,176 pixels.

[0173] Figure 32 is a schematic diagram showing an example of an output format for outputting a feature map from a sensor, applicable to the embodiment. Figure 32 shows an example in which the second layer feature map shown in Figure 31 is added to the image data 70 output from the image sensor (imaging device 100).

[0174] In Figure 32, section (a) shows an example of rearranging the feature map data 71a into row-oriented data and adding it to the lower end of the image data 70 for output. More specifically, the 56-line × 56-pixel × 256-ch feature map data is aligned to one line per channel and rearranged into 256-line × 3136-pixel feature map data 71a. This feature map data 71a is then added to the lower end of the image data 70.

[0175] Section (b) of Figure 32 shows an example of rearranging feature map data 71b into columnar data and distributing it to each line of image data 70. More specifically, feature map data of 6 lines × 56 pixels × 256 channels is rearranged into feature map data 71b of 2048 lines × 392 pixels (= 56 pixels × 7) so that it fits into the 2160 lines of image data 70. This feature map data 71b is added to the horizontal ranking period of image data 70.

[0176] In this way, by rearranging the feature map data into data in the column or row direction and attaching it to the edges of the image data 70, it is possible to ensure the visibility of the image data 70.

[0177] (4-2. First Modification of the First Embodiment) Next, a first modification of the first embodiment will be described.

[0178] In the first embodiment described above, the sensor 10c has two line buffers and one frame buffer, and the image data stored in the frame buffer is output to the subsequent processing unit 300a during the vertical blanking period. In contrast, in the first modification of the first embodiment, the sensor has three line buffers, and the image data stored in the line buffer used for output is output to the subsequent processing unit during the horizontal blanking period.

[0179] Figure 33 is a schematic diagram showing the configuration of an example of a CNN applicable to the first modification of the first embodiment. The configuration of the CNN itself is the same as the configuration described using Figure 15 above, so its explanation is omitted here. In Figure 33, the processing of the third and fourth layers is the execution range of the processing performed on the sensor 10 by the subsequent stage.

[0180] Figure 34 is a block diagram showing the configuration of an example of an information processing device 1d according to a first modification of the first embodiment. In Figure 34, the sensor 10d includes a pixel array unit 101, line buffers 124a, 124b, and 124c, an AI engine 120b, and a switch 123c. The line buffers 124a, 124b, and 124c, and the AI ​​engine 120b are connected to each other via a bus 121 so that they can communicate with one another.

[0181] Each of the line buffers 124a, 124b, and 124c is capable of storing at least a number of lines corresponding to the size of the filter used by the AI ​​engine 120 in CNN processing. In the following, it is assumed that the AI ​​engine 120 performs convolution processing using a 3x3 filter, and each of the line buffers 124a, 124b, and 124c is capable of storing at least three lines of image data.

[0182] In Figure 34 and Figure 36 (described later), line buffers 124a, 124b, and 124c are also shown as line BF#a, line BF#b, and line BF#c, respectively. The AI ​​engine 120 corresponds to the AI ​​processing unit 200 shown in Figure 1.

[0183] Image data read from the pixel array unit 101 is transferred to the line buffer 124a via the data path 110. Image data read from line buffers 124a and 124b is transferred to the AI ​​engine 120 via the bus 121. Image data read from line buffer 124c is input to terminal "1" of switch 123c. Terminal "1" of switch 123c is selected during the horizontal blanking period (H blank), and terminal "0" is selected during other periods. The output of switch 123c is transferred to the frame buffer 301a of the downstream processing unit 300a.

[0184] The configuration of the subsequent processing unit 300a is the same as that of the subsequent processing unit 300a described with reference to Figure 16, so its description is omitted here.

[0185] Figure 35 is a flowchart illustrating an example of processing by the information processing device 1d according to the first embodiment. Figure 36 is a sequence diagram illustrating an example of processing by the information processing device 1d according to the first embodiment. Each process shown in the sequence in Figure 36 corresponds to each process in the flowchart in Figure 35.

[0186] The meaning of each part in Figure 36 is the same as that of each part in Figure 18 described above, so an explanation will be omitted here. Also, as mentioned above, one frame is assumed to be 9 pixels × 9 lines.

[0187] In Figure 35, in step S100c, the information processing device 1d reads out one line of image data from the pixel array unit 101 (time t 100 ). In the next step S101c, the information processing device 1d writes the image data of one line read in step S100c to the line buffer 124a. The writing of one line of image data to this line buffer 124a is performed cumulatively until the writing of three lines is completed.

[0188] In Figure 36, line buffer #a(1), line buffer #a(2), and line buffer #a(3) each represent line-by-line writing to line buffer 124a.

[0189] In the next step S102c, the information processing device 1d determines whether the writing of three lines of image data to the line buffer 124a is complete. If the information processing device 1d determines that it is not complete (step S102c, "No"), it returns to step S100c and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1d determines that the writing of three lines of image data is complete (step S102c, "Yes"), it moves the process to step S150a (time t 101 ).

[0190] In step S150a, the information processing device 1d reads three lines of image data from the line buffer 124a using the AI ​​engine 120b, and performs a convolution process using a first-layer 3x3 filter on the read three lines of image data using the AI ​​engine 120b, outputting 1 line x 3 channel image data. This 1 line x 3 channel image data is written to the line buffer 124b in the next step S151a. The writing of this 1 line x 3 channel image data to the line buffer 124b is performed cumulatively until the writing of 3 lines x 3 channels is completed.

[0191] In Figure 36, line buffer #b(1), line buffer #b(2), and line buffer #b(3) each represent line-by-line writing to line buffer 124b.

[0192] In the next step S152a, the information processing device 1d determines whether the writing of 3 lines × 3 channels to the line buffer 124b has been completed. If the information processing device 1d determines that the writing has not been completed (step S152a, "No"), it returns to step S100c. On the other hand, if the information processing device 1d determines that the writing has been completed (step S152a, "Yes"), it moves the process to step S170.

[0193] In step S170, the information processing device 1d reads three lines of image data from the line buffer 124b using the AI ​​engine 120, and performs a convolution process using a second layer 3x3 filter on the read three lines of image data using the AI ​​engine 120b, outputting 1 line x 5 channel image data. This 1 line x 5 channel image data is written to the line buffer 124c in the next step S171. The writing of this 1 line x 5 channel image data to the line buffer 124c is performed cumulatively until the writing of 5 lines x 5 channels is completed (second layer in Figure 33).

[0194] In the next step S172, the information processing device 1d waits for the horizontal blanking period (H blank) in the reading process of the pixel array unit 101. When the horizontal blanking period arrives, terminal "1" is selected at switch 123c, and the device becomes ready to transfer image data from line buffer 124c to frame buffer 301a in downstream processing unit 300a.

[0195] In the next step S173, the information processing device 1d reads a 1-line x 5-channel feature map from the line buffer 124c and outputs the read feature map from the sensor 10d via the switch 123c. In the next step S174, the information processing device 1d writes the 5-line x 5-channel (5 pixels x 5 lines x 5 channels) image data output from the sensor 10d to the frame buffer 301a.

[0196] In the next step, S174, the information processing device 1d determines whether or not the writing of 5 lines x 5 channels of image data to the frame buffer 301a is complete. If the information processing device 1d determines that the writing is not complete (step S175, "No"), it returns to step S100c and reads the next line. On the other hand, if the information processing device 1d determines that the writing is complete (step S175, "Yes"), it moves the process to step S130c.

[0197] Step S130c and subsequent steps constitute the subsequent execution range in which processing by the subsequent processing unit 300a is performed.

[0198] In step S130c, the information processing device 1d reads 5-line x 5-channel image data from the frame buffer 301a. In the next step S131c, the AI ​​engine 310 performs a convolution process using a third layer of 3x3 filters on the read image data, outputting 3-line x 7-channel (3 pixels x 3 lines x 7 channels) image data. In the next step S132c, this 3-line x 7-channel image data is written to the frame buffer 301b.

[0199] In the next step S140c, the information processing device 1d reads 3-line x 7-channel image data from the frame buffer 301b. In the next step S141c, the AI ​​engine 310 performs a convolution process on the read image data using a fourth layer of 3x3 filters, outputting 1-line x 5-channel (1 pixel x 1 line x 5 channels) image data. In the next step S142c, this 1-line x 5-channel image data is written to the frame buffer 301a.

[0200] Although not shown in the diagram, the information processing device 1d uses the AI ​​engine 310 to combine the 1-line × 5-channel image data written to the frame buffer 301a in step S142c at the output layer and outputs it, for example, as an inference result.

[0201] Thus, in the first modification of the first embodiment, the sensor 10d uses three line buffers 124a to 124c and does not use a frame buffer, making it possible to reduce the size and cost of the hardware compared to the first embodiment.

[0202] (4-3. Second Modification of the First Embodiment) Next, a second modification of the first embodiment will be described.

[0203] In the first modification of the first embodiment described above, the sensor 10c has two line buffers and one frame buffer, and the image data stored in the frame buffer is output to the subsequent processing unit 300a during the vertical blanking period. In contrast, in the second modification of the first embodiment, the sensor has two line buffers, and the image data stored in the line buffer used for output is output to the subsequent processing unit.

[0204] Figure 37 is a schematic diagram showing the configuration of an example of a CNN applicable to a second modification of the first embodiment. The configuration of the CNN itself is the same as the configuration described using Figure 15 above, so its explanation is omitted here. In Figure 37, the processing of the third and fourth layers constitutes the execution range of subsequent processing on the sensor 10.

[0205] Figure 38 is a block diagram showing the configuration of an example of an information processing device 1e according to a second modification of the first embodiment. In Figure 38, the sensor 10e includes a pixel array unit 101, line buffers 124a and 124b, and an AI engine 120c. The line buffers 124a and 124b and the AI ​​engine 120c are connected to each other via a bus 121 so that they can communicate with one another.

[0206] Line buffers 124a and 124b are each capable of storing at least a number of lines corresponding to the size of the filter used by the AI ​​engine 120c in CNN processing. In the following, it is assumed that the AI ​​engine 120c performs convolution processing using a 3x3 filter, and line buffers 124a and 124b are each capable of storing at least 3 lines of image data.

[0207] In Figure 38 and Figure 40 (described later), line buffers 124a and 124b are also shown as line BF#a and line BF#b, respectively. The AI ​​engine 120c corresponds to the AI ​​processing unit 200 shown in Figure 1.

[0208] The image data read from the pixel array unit 101 is transferred to the line buffer 124a via the data path 110. Also, the image data read from the pixel array unit 101 can be directly output from the data path 110 to the subsequent processing unit 300a. The image data read from the line buffers 124a and 124b is transferred to the AI engine 120c via the bus 121. Also, the image data read from the line buffer 124b is transferred to the frame buffer 301a of the subsequent processing unit 300a.

[0209] The configuration of the subsequent processing unit 300a is equivalent to the subsequent processing unit 300a described using FIG. 16, so the description here is omitted.

[0210] FIG. 39 is a flowchart showing an example of the processing by the information processing apparatus 1e according to the second modification of the first embodiment. Also, FIG. 40 is a sequence diagram showing an example of the processing by the information processing apparatus 1e according to the second modification of the first embodiment. Each process shown in the sequence of FIG. 40 corresponds to each process of the flowchart of FIG. 39.

[0211] The meaning of each part in FIG. 40 is the same as each part of FIG. 18 described above, so the description here is omitted. Also, as described above, it is assumed that one frame is 9 pixels × 9 lines.

[0212] In FIG. 39, in step S100d, the information processing apparatus 1e reads one line of image data from the pixel array unit 101 (time t 100 ). In the next step S101d, the information processing apparatus 1e writes the one line of image data read in step S100d into the line buffer 124a. The writing of one line of image data to this line buffer 124a is performed accumulatively until the writing of three lines is completed.

[0213] Note that in FIG. 40, line buffer #a(1), line buffer #a(2), and line buffer #a(3) each indicate the writing of one line to the line buffer 124a.

[0214] In the next step S102d, the information processing device 1e determines whether the writing of three lines of image data to the line buffer 124a is complete. If the information processing device 1e determines that it is not complete (step S102d, "No"), it returns to step S100d and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1e determines that the writing of three lines of image data is complete (step S102d, "Yes"), it moves the process to step S150a (time t 101 ).

[0215] In step S150b, the information processing device 1e reads three lines of image data from the line buffer 124a using the AI ​​engine 120, and performs a convolution process using a 3x3 filter in the first layer on the read three lines of image data using the AI ​​engine 120, outputting one line x three channels of image data. This one line x three channels of image data is written to the line buffer 124b in the next step S151b. The writing of one line x three channels of image data to the line buffer 124b is performed cumulatively until the writing of three lines x three channels is completed.

[0216] In Figure 40, line buffer #b(1), line buffer #b(2), and line buffer #b(3) each represent line-by-line writing to line buffer 124b.

[0217] In the next step S152b, the information processing device 1e determines whether the writing of 3 lines × 3 channels to the line buffer 124b has been completed. If the information processing device 1e determines that the writing has not been completed (step S152b, "No"), it returns to step S100b. On the other hand, if the information processing device 1e determines that the writing has been completed (step S152b, "Yes"), it moves the process to step S180.

[0218] In step S180, the information processing device 1e reads image data for three lines from the line buffer 124b using the AI ​​engine 120, and performs a convolution process using a second layer 3x3 filter on the read image data for three lines using the AI ​​engine 120, outputting image data for 1 line x 5 channels. This image data for 1 line x 5 channels is transferred to the downstream processing unit 300a in the next step S181 and written to the frame buffer 301a. The writing of this image data for 1 line x 5 channels to the line buffer 331a is performed cumulatively until the writing of 5 lines x 5 channels is completed (second layer in Figure 37).

[0219] In the next step S182, the information processing device 1e determines whether the writing of 5 lines x 5 channels to the frame buffer 301a has been completed. If the information processing device 1e determines that the writing is not complete (step S182, "No"), it returns to step S100d and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1e determines that the writing is complete (step S182, "Yes"), it moves the process to step S130d.

[0220] Step S130d and subsequent steps constitute the subsequent execution range in which processing by the subsequent processing unit 300a is performed.

[0221] In step S130d, the information processing device 1e reads 5-line x 5-channel image data from the frame buffer 301a. In the next step S131d, the AI ​​engine 310 performs a convolution process on the read image data using a third layer of 3x3 filters, outputting 3-line x 7-channel (3 pixels x 3 lines x 7 channels) image data. In the next step S132d, this 3-line x 7-channel image data is written to the frame buffer 301b.

[0222] In the next step S140d, the information processing device 1e reads 3-line x 7-channel image data from the frame buffer 301b. In the next step S141d, the AI ​​engine 310 performs a convolution process using a fourth layer of 3x3 filters on the read image data, outputting 1-line x 5-channel image data. In the next step S142d, this 1-line x 5-channel image data is written to the frame buffer 301a.

[0223] Although not shown in the diagram, the information processing device 1e combines the 1-line × 5-channel image data written to the frame buffer 301a in step S142d using the AI ​​engine 310 in the output layer and outputs it as, for example, an inference result.

[0224] Thus, in the second modification of the first embodiment, the sensor 10e uses two line buffers 124a and 124b and does not use a frame buffer, making it possible to further reduce the size and cost of the hardware compared to the first modification of the first embodiment.

[0225] (5. Second Embodiment of the Disclosure) Next, a second embodiment of the Disclosure will be described. In the first embodiment and its various modifications described above, the subsequent processing unit executed processing #B by CNN#B on a frame basis. In contrast, the second embodiment executes processing #B by CNN#B on a line basis in the subsequent processing unit.

[0226] (5-1-1. Outline of the Second Embodiment) First, the first embodiment will be described in general terms using Figures 41 and 42. Figure 41 is a schematic diagram showing the process according to the second embodiment in comparison with the process according to existing technology. Sections (a) and (b) of Figure 41 are reproduced for comparison with sections (a) and (b) of Figure 13 described above, so a detailed explanation will be omitted here.

[0227] Section (c) of Figure 41 shows an example of processing according to the first embodiment. In the segmented CNN #A within the sensor 10, processing #A is executed on a line basis on the image data read from the imaging device 100 (pixel array section 101). Therefore, it is possible to perform the reading of image data from the imaging device 100 and the CNN processing by segmented CNN #A in parallel.

[0228] Furthermore, in the second embodiment, processing #B by segmented CNN #B is also executed on a line basis in the downstream processing unit. Therefore, it is possible to execute the CNN processing by segmented CNN #A in the sensor and the CNN processing by segmented CNN #B in the downstream processing unit in parallel. Consequently, by applying the second embodiment, it is possible to shorten the latency from the readout of the imaging device 100 to the output of the result of processing #B compared to the configuration of the first embodiment.

[0229] Figure 42 is an example sequence diagram showing the process according to the second embodiment in comparison with the process using existing technology.

[0230] In Figure 42, section (a) is a reproduction of section (a) of Figure 26 described above for comparative reference, so its explanation is omitted here.

[0231] Section (b) of Figure 42 shows a sequence diagram of an example of processing according to the second embodiment. The sensor 10 includes an AI engine #A that executes processing #A using CNN #A, and a line buffer capable of storing a number of lines corresponding to the filters used for CNN processing. The downstream processing unit after the sensor 10 includes an AI engine #B that executes processing #B using CNN #B, and a frame buffer for storing feature maps.

[0232] In section (b) of Figure 42, the reading of image data from the pixel array section 101 of the sensor 10 is started in synchronization with the vertical synchronization signal XVS, and the read image data (sensor read data) is written to the line buffer of the sensor 10. For example, if process #A performs CNN processing using a 3x3 filter, when three lines of image data are stored in the line buffer, the AI ​​engine #A reads the three lines of image data from the line buffer and executes process #A.

[0233] AI engine #A transmits the image data (feature map) resulting from processing #A to the subsequent processing unit. The subsequent processing unit stores the image data transmitted from AI engine #A in line buffers on a line-by-line basis. In the subsequent processing unit, when three lines of (feature maps) are stored in the line buffer, AI engine #B executes processing #B. Thus, in the second embodiment, the subsequent processing unit can be configured using a small-capacity line buffer for feature maps.

[0234] In the second embodiment, the results of processing #A by CNN#A in the sensor are stored in a line buffer on a line-by-line basis in the subsequent processing unit, and the subsequent processing unit performs processing #B by CNN#B on the image data stored in the line buffer. Therefore, it is possible to perform the reading of image data from the pixel array unit 101 of the sensor and processing #A by CNN#A, and the processing #B by CNN#B in the subsequent processing unit in parallel, and compared to the configuration of the first embodiment, it is possible to further shorten the latency from the reading of the image data from the imaging device 100 to the output of the result of processing #B.

[0235] Furthermore, in the second embodiment, it is possible to further reduce the latency in the subsequent processing unit. Therefore, if the same amount of time as in the first embodiment can be secured for processing in the subsequent processing unit, applying the second embodiment makes it possible to handle more complex tasks than the configuration of the first embodiment.

[0236] (5-1-2. More Specific Description of the Second Embodiment) Next, the second embodiment will be described in more detail.

[0237] Figure 43 is a schematic diagram showing an example configuration of a CNN applicable to the second embodiment. The configuration of the CNN itself is the same as the configuration described using Figure 15 above, so its explanation is omitted here. In Figure 43, the processing of the third and fourth layers constitutes the execution range of subsequent processing on the sensor 10.

[0238] Figure 44 is a block diagram showing an example configuration of an information processing device 1f according to the second embodiment. The configuration of the sensor 10d according to the second embodiment is equivalent to the sensor 10d described using Figure 34, so its description is omitted here.

[0239] The downstream processing unit 300b performs downstream processing on the sensor 10d and includes line buffers 330a, 330b, and 330c, and an AI engine 310a, which are connected to each other via a bus 320 for communication. In Figure 44 and Figure 46 described later, the line buffers 330a, 330b, and 330c are also shown as lines BF#A, BF#B, and BF#C, respectively. The AI ​​engine 310a corresponds to the AI ​​processing unit 210 shown in Figure 1.

[0240] Figure 45 is a flowchart illustrating an example of processing by the information processing device 1f according to the second embodiment. Figure 46 is a sequence diagram illustrating an example of processing by the information processing device 1f according to the second embodiment. Each process shown in the sequence in Figure 46 corresponds to each process in the flowchart in Figure 45.

[0241] The meaning of each part in Figure 46 is the same as that of each part in Figure 18 described above, so an explanation will be omitted here. Also, as mentioned above, one frame is assumed to be 9 pixels × 9 lines.

[0242] In Figure 45, in step S100e, the information processing device 1f reads out one line of image data from the pixel array unit 101 (time t 100). In the next step S101e, the information processing device 1f writes the image data of one line read in step S100e to the line buffer 124a. The writing of one line of image data to this line buffer 124a is performed cumulatively until the writing of three lines is completed.

[0243] In Figure 46, line buffer #a(1), line buffer #a(2), and line buffer #a(3) each represent line-by-line writing to line buffer 124a.

[0244] In the next step S102e, the information processing device 1f determines whether the writing of three lines of image data to the line buffer 124a is complete. If the information processing device 1f determines that it is not complete (step S102e, "No"), it returns to step S100e and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1f determines that the writing of three lines of image data is complete (step S102e, "Yes"), it moves the process to step S150c (time t 101 ).

[0245] In step S150c, the information processing device 1f reads three lines of image data from the line buffer 124a using the AI ​​engine 120b, and performs a convolution process using a 3x3 filter in the first layer on the read three lines of image data, outputting one line x three channels of image data. This one line x three channels of image data is written to the line buffer 124b in the next step S151c. The writing of this one line x three channels of image data to the line buffer 124b is performed cumulatively until the writing of three lines x three channels is completed.

[0246] In Figure 46, line buffer #b(1), line buffer #b(2), and line buffer #b(3) each represent line-by-line writing to line buffer 124b.

[0247] In the next step S152c, the information processing device 1f determines whether the writing of 3 lines × 3 channels to the line buffer 124b has been completed. If the information processing device 1f determines that the writing has not been completed (step S152c, "No"), it returns to step S100e. On the other hand, if the information processing device 1f determines that the writing has been completed (step S152c, "Yes"), it moves the process to step S170a.

[0248] In step S170a, the information processing device 1f reads three lines of image data from the line buffer 124b using the AI ​​engine 120b, and performs a convolution process using a second layer 3x3 filter on the read three lines of image data using the AI ​​engine 120b, outputting 1 line x 5 channel image data. This 1 line x 5 channel image data is written to the line buffer 124c in the next step S171a. The writing of this 1 line x 5 channel image data to the line buffer 124c is performed cumulatively until the writing of 5 lines x 5 channels is completed (second layer in Figure 43).

[0249] In the next step S172a, the information processing device 1f waits for the horizontal blanking period (H blank) in the reading process of the pixel array unit 101. When the horizontal blanking period arrives, terminal "1" is selected at switch 123c, and the device becomes ready to transfer image data from line buffer 124c to line buffer 330a in the subsequent processing unit 300a.

[0250] In the next step S180a, the information processing device 1f reads a 1-line x 5-channel feature map from the line buffer 124c and outputs the read feature map from the sensor 10f via the switch 123c. In the next step S181a, the information processing device 1f transfers the 5-line x 5-channel (5 pixels x 5 lines x 5 channels) image data output from the sensor 10f to the downstream processing unit 300b via the switch 123c and writes it to the line buffer 330a.

[0251] In the next step S182a, the information processing device 1f determines whether or not the writing of 5 lines x 5 channels of image data to the line buffer 330a is complete. If the information processing device 1f determines that the writing is not complete (step S175, "No"), it returns to step S100e and reads the next line. On the other hand, if the information processing device 1f determines that the writing is complete (step S182a, "Yes"), it moves the process to step S200.

[0252] Step S200 and subsequent steps constitute the subsequent execution range in which processing by the subsequent processing unit 300b is executed.

[0253] In step S200, the information processing device 1f reads 3-line x 5-channel image data from the line buffer 330a, and performs a convolution process using a third layer 3x3 filter on the read image data with the AI ​​engine 310a, outputting 1-line x 7-channel (3 pixels x 3 lines x 7 channels) image data. This 1-line x 7-channel image data is written to the line buffer 330b in the next step S201.

[0254] In the next step, S202, the information processing device 1f determines whether or not the writing of 3 lines x 7 channels of image data to the line buffer 330b is complete. If the information processing device 1f determines that the writing is not complete (step S202, "No"), it returns to step S100e and reads the next line. On the other hand, if the information processing device 1f determines that the writing is complete (step S202, "Yes"), it moves the process to step S203.

[0255] In step S203, the information processing device 1f reads 3 lines x 7 channels of image data from the line buffer 330b. In the next step S204, the AI ​​engine 310a performs a convolution process on the read image data using a fourth layer of 3x3 filters, outputting 1 line x 5 channels of image data. In the next step S205, this 1 line x 5 channels of image data is written to the line buffer 330c.

[0256] Although not shown in the diagram, the information processing device 1f combines the 1-line × 5-channel image data written to the line buffer 330c in step S205 using the AI ​​engine 310a in the output layer and outputs it as, for example, an inference result.

[0257] In step S201, when writing the third line to the line buffer 330b, the information processing device 1f may transfer the image data of the third line directly to the AI ​​engine 310a without writing it to the line buffer 330b. In this case, the AI ​​engine 310a uses the image data of the first and second lines written to the line buffer 330b and the image data of the third line transferred directly to perform a convolution process using a fourth layer 3x3 filter.

[0258] Thus, in the second embodiment, the downstream processing unit 333b uses three line buffers 330a to 330c and does not use a frame buffer, making it possible to reduce the size and cost of the hardware compared to the first embodiment and its various modifications.

[0259] (5-2. Modifications of the Second Embodiment) Next, modifications of the second embodiment will be described. The modifications of the second embodiment are examples that combine the second embodiment with the second modification of the first embodiment described above.

[0260] Figure 47 is a schematic diagram showing the configuration of an example of a CNN applicable to a modified version of the second embodiment. The configuration of the CNN itself is the same as the configuration described using Figure 15 above, so its explanation is omitted here. In Figure 47, the processing of the third and fourth layers constitutes the execution range of subsequent processing on the sensor 10.

[0261] Figure 48 is a block diagram showing the configuration of an example of an information processing device 1g according to a modification of the second embodiment. The configuration of the sensor 10e according to the modification of the second embodiment is equivalent to the sensor 10e described using Figure 38, so its description is omitted here. Also, the configuration of the downstream processing unit 300b according to the modification of the second embodiment is equivalent to the configuration of the downstream processing unit 300b described using Figure 44, so its description is omitted here. That is, in the information processing device 1g according to the modification of the second embodiment, the sensor 10e has two line buffers, and the downstream processing unit 300b has three line buffers.

[0262] Figure 49 is a flowchart illustrating an example of processing by the information processing device 1g according to the second embodiment. Figure 50 is a sequence diagram illustrating an example of processing by the information processing device 1g according to the second embodiment. Each process shown in the sequence in Figure 50 corresponds to each process in the flowchart in Figure 49.

[0263] The meaning of each part in Figure 50 is the same as that of each part in Figure 18 described above, so an explanation will be omitted here. Also, as mentioned above, one frame is assumed to be 9 pixels × 9 lines.

[0264] In Figure 49, in step S100f, the information processing device 1g reads out one line of image data from the pixel array unit 101 (time t 100 ). In the next step S101f, the information processing device 1g writes the image data of one line read in step S100f to the line buffer 124a. The writing of one line of image data to this line buffer 124a is performed cumulatively until the writing of three lines is completed.

[0265] In Figure 50, line buffer #a(1), line buffer #a(2), and line buffer #a(3) each represent line-by-line writing to line buffer 124a.

[0266] In the next step S102f, the information processing device 1g determines whether the writing of three lines of image data to the line buffer 124a is complete. If the information processing device 1g determines that it is not complete (step S102f, "No"), it returns to step S100f and reads the next line from the pixel array unit 101. On the other hand, if the information processing device 1g determines that the writing of three lines of image data is complete (step S102f, "Yes"), it moves the process to step S150d (time t 101 ).

[0267] In step S150d, the information processing device 1g reads three lines of image data from the line buffer 124a using the AI ​​engine 120c, and performs a convolution process using a 3x3 filter in the first layer on the read three lines of image data using the AI ​​engine 120c, outputting one line x three channels of image data. This one line x three channels of image data is written to the line buffer 124b in the next step S151d. The writing of this one line x three channels of image data to the line buffer 124b is performed cumulatively until the writing of three lines x three channels is completed.

[0268] In Figure 50, line buffer #b(1), line buffer #b(2), and line buffer #b(3) each represent line-by-line writing to line buffer 124b.

[0269] In the next step, S152d, the information processing device 1g determines whether the writing of 3 lines × 3 channels to the line buffer 124b has been completed. If the information processing device 1g determines that the writing has not been completed (step S152d, "No"), it returns to step S100f. On the other hand, if the information processing device 1g determines that the writing has been completed (step S152d, "Yes"), it moves the process to step S190.

[0270] In step S190, the information processing device 1g reads three lines of image data from the line buffer 124b using the AI ​​engine 120c. The AI ​​engine 120 then performs a convolution process using a second layer 3x3 filter on the read three lines of image data, outputting one line x five channels of image data. In the next step S191, this one line x five channels of image data is transferred to the downstream processing unit 300b and written to the line buffer 330a. The writing of this one line x five channels of image data to the line buffer 330a is performed cumulatively until the writing of three lines x five channels is completed (second layer in Figure 47).

[0271] In the next step S192, the information processing device 1g determines whether or not the writing of 3 lines × 5 channels of image data to the line buffer 330a is complete. If the information processing device 1g determines that the writing is not complete (step S192, "No"), it returns to step S100f and reads the next line. On the other hand, if the information processing device 1g determines that the writing is complete (step S192, "Yes"), it moves the process to step S200a.

[0272] Step S200a onward constitutes the subsequent execution range in which processing by the subsequent processing unit 300b is executed.

[0273] In step S200a, the information processing device 1g reads 3-line × 5-channel image data from the line buffer 330a, and performs a convolution process using a third layer 3 × 3 filter on the read image data with the AI ​​engine 310a, outputting 1-line × 7-channel (1 pixel × 1 line × 7 channels) image data. This 1-line × 7-channel image data is written to the line buffer 330b in the next step S201a. The writing of this 1-line × 5-channel image data to the line buffer 330b is performed cumulatively until the writing of 3 lines × 5 channels is completed (second layer in Figure 47).

[0274] In the next step S202a, the information processing device 1g determines whether or not the writing of 3 lines × 7 channels of image data to the line buffer 330b is complete. If the information processing device 1g determines that the writing is not complete (step S202a, "No"), it returns to step S100f and reads the next line. On the other hand, if the information processing device 1g determines that the writing is complete (step S202a, "Yes"), it moves the process to step S203a.

[0275] In step S203a, the information processing device 1g reads 3 lines × 7 channels of image data from the line buffer 330b. In the next step S204a, the AI ​​engine 310a performs a convolution process on the read image data using a fourth layer of 3 × 3 filters, outputting 1 line × 5 channels of image data. In the next step S205a, this 1 line × 5 channels of image data is written to the line buffer 330c.

[0276] Although not shown in the diagram, the information processing device 1g uses the AI ​​engine 310a to combine the 1-line × 5-channel image data written to the line buffer 330c in step S205a at the output layer and outputs it, for example, as an inference result.

[0277] In step S201a, when writing the third line to the line buffer 330b, the information processing device 1g may transfer the image data of the third line directly to the AI ​​engine 310a without writing it to the line buffer 330b. In this case, the AI ​​engine 310a uses the image data of the first and second lines written to the line buffer 330b and the image data of the third line transferred directly to perform a convolution process using a fourth layer 3x3 filter.

[0278] Thus, in this modified version of the second embodiment, the sensor 10d uses two line buffers 124a and 124b and does not use a frame buffer, making it possible to further reduce the size and cost of the hardware compared to the second embodiment.

[0279] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0280] Furthermore, this technology can also take the following configurations: (1) An information processing device comprising: a first network including an input layer and a second network including an output layer, wherein the network that processes input image data using an n-pixel × n-line (n is an integer of 2 or more) filter is divided between the Nth layer and the (N+1)th layer (N is an integer of 2 or more) in the intermediate layer, and a first processing unit that executes the processing by the first network, wherein the first processing unit has the first network perform the processing on the input image data in units of n lines and outputs the results of the processing to the second network line by line. (2) The information processing device according to (1), further comprising: (N-1) or more memories used by the first processing unit for the processing, each capable of storing at least the n lines of image data. (3) The information processing device according to (2), wherein the first processing unit uses each of the (N-1) or more memories for each layer of the first network to execute the processing. (4) The information processing apparatus according to (1) or (2), wherein each of the (N-1) or more memories is a line buffer. (5) The information processing apparatus according to (4), wherein the (N-1) or more memories include N line buffers, and the image data stored in the line buffers after the processing of the N layer is performed is read out in units of n lines according to the horizontal blanking period of the input image data and output to the second network. (6) The information processing apparatus according to (3), wherein the (N-1) or more memories include (N-1) line buffers and a frame buffer capable of storing image data for one frame, the first processing unit stores image data read from the line buffers that store the results of the processing of the N layer in the frame buffer, and the image data stored in the frame buffer is read out in units of n lines according to the vertical blanking period of the input image data and output to the second network.(7) An information processing device according to any one of (1) to (6), further comprising: an imaging unit that takes an image and outputs an image, wherein the image is input to the first processing unit as input image data. (8) An information processing device according to any one of (1) to (7), wherein the second network includes an M layer (where M is an integer of 2 or more), a second processing unit that performs processing by the second network, and a plurality of memories used by the second processing unit for the processing, each capable of storing at least the n lines of image data, wherein the image data output from the first network in units of n lines is stored in one of the plurality of memories. (9) An information processing device according to (8), wherein the plurality of memories are two or more frame buffers, and the second processing unit performs the processing by sequentially using the two or more frame buffers for each layer of the M layer. (10) The information processing apparatus according to (9), wherein each of the plurality of memories is a M or more line buffer capable of storing at least the n lines of image data, and the second processing unit performs the processing using the M or more line buffers layer by layer. (11) The information processing apparatus according to any one of (1) to (10), wherein the processing is a convolution process using the filter. (12) An information processing system comprising: a first network including an input layer and a second network including an output layer, wherein the network that processes input image data using an n-pixel × n-line (n is an integer of 2 or more) filter is divided between the Nth layer and the (N+1)th layer (N is an integer of 2 or more) in the intermediate layer, and the first processing unit performs the processing by the first network and the second processing unit performs the processing by the second network, wherein the first processing unit has the first network perform the processing on the input image data in n-line units and outputs the results of the processing line by line to the second network. (13) The information processing system according to (12), further comprising an imaging unit that performs imaging and outputs an image, and inputs the image as input image data to the first processing unit.(14) The information processing system according to (13), wherein the imaging unit and the first processing unit are integrally configured. (15) The information processing system according to any one of (12) to (14), wherein the processing is a convolution process using the filter. (16) An information processing method comprising: a first network including an input layer and a second network including an output layer, wherein the network that processes input image data using an n-pixel × n-line (n is an integer of 2 or more) filter is divided between the Nth layer and the (N+1)th layer (N is an integer of 2 or more) in the intermediate layer, and the first processing step is to perform processing by the first network, wherein the first processing step is to have the first network perform the processing on the input image data in n-line units and output the results of the processing to the second network line by line.

[0281] 1, 1a, 1b, 1c, 1d, 1e, 1f, 1g Information processing device 10, 10a, 10b, 10c, 10d, 10e, 10f Sensor 60 CNN model 60a Sensor CNN 60b Post-processing CNN 100 Imaging device 101 Pixel array section 120, 120a, 120b, 120c, 310, 310a AI engine 122a, 122b, 122c, 301a, 301b Frame buffer 123b, 123c Switch 124a, 124b, 124c, 330a, 330b, 330c Line buffer 200, 210 AI processing section 300a, 300b Post-processing section

Claims

1. An information processing apparatus comprising: a first network including an input layer, wherein the network that processes input image data using an n-pixel × n-line (n is an integer of 2 or more) filter is divided between the Nth layer and the (N+1)th layer (N is an integer of 2 or more) in the intermediate layer; and a second network including an output layer, wherein the first processing unit executes the processing performed by the first network, and the first processing unit performs the processing on the input image data in units of n lines, and outputs the results of the processing line by line to the second network.

2. The information processing apparatus according to claim 1, further comprising (N-1) or more memories used by the first processing unit for the processing, each capable of storing at least the n lines of image data.

3. The information processing apparatus according to claim 2, wherein the first processing unit uses each of the (N-1) or more memories for each layer of the first network to perform the processing.

4. The information processing apparatus according to claim 1, wherein each of the (N-1) or more memories is a line buffer.

5. The information processing apparatus according to claim 4, comprising N line buffers as (N-1) or more memory units, wherein the image data stored in the line buffers after the N-layer processing is performed is read out in units of n lines according to the horizontal blanking period of the input image data and output to the second network.

6. The information processing apparatus according to claim 3, comprising (N-1) or more memory units, including (N-1) line buffers and a frame buffer capable of storing image data for one frame, wherein the first processing unit cumulatively stores image data read from the line buffers that store the results of the processing of the N layers in the frame buffer, and the image data stored in the frame buffer is read out in units of n lines according to the vertical blanking period of the input image data and output to the second network.

7. An imaging unit that performs imaging and outputs an image, the information processing apparatus according to claim 1, further comprising: an imaging unit that performs imaging and outputs an image, and inputs the image as input image data to the first processing unit.

8. The information processing apparatus according to claim 1, wherein the second network includes an M layer (where M is an integer of 2 or more), a second processing unit that performs processing by the second network, and a plurality of memories used by the second processing unit for the processing, each capable of storing at least the n lines of image data, and the image data output from the first network in units of the n lines is stored in one of the plurality of memories.

9. The information processing apparatus according to claim 8, wherein the plurality of memories are two or more frame buffers, and the second processing unit executes the processing by sequentially using the two or more frame buffers for each layer of the M layer.

10. The information processing apparatus according to claim 9, wherein each of the plurality of memories is a line buffer of M or more, each capable of storing at least the n lines of image data, and the second processing unit performs the processing using the M or more line buffers layer by layer.

11. The information processing apparatus according to claim 1, wherein the processing is a convolution process using the filter.

12. An information processing system comprising: a first network including an input layer and a second network including an output layer, wherein the network that processes input image data using an n-pixel × n-line (n is an integer of 2 or more) filter is divided between the Nth layer and the (N+1)th layer (N is an integer of 2 or more) in the intermediate layer, and the first processing unit performs the processing by the first network and the second processing unit performs the processing by the second network, wherein the first processing unit has the first network perform the processing on the input image data in units of n lines and outputs the results of the processing line by line to the second network.

13. An imaging unit that performs imaging and outputs an image, the information processing system according to claim 12, further comprising the imaging unit that performs imaging and outputs an image, and inputs the image as input image data to the first processing unit.

14. The information processing system according to claim 13, wherein the imaging unit and the first processing unit are integrally configured.

15. The information processing system according to claim 12, wherein the processing is a convolution process using the filter.

16. An information processing method comprising: a first network including an input layer and a second network including an output layer, wherein the network that processes input image data using an n-pixel × n-line (n is an integer of 2 or more) filter is divided between the Nth layer and the (N+1)th layer (N is an integer of 2 or more) in the intermediate layer, and the first processing step executes the processing by the first network, wherein the first processing step has the first network perform the processing on the input image data in units of n lines, and outputs the results of the processing line by line to the second network.

Citation Information

Patent Citations

  • Imaging device, imaging method, and imaging program

    JP2022186333A

  • Server device, terminal device, information processing method, and information processing system

    JP2024058463A