Methods, apparatuses, and systems for processing image data representing a scene to extract features
By dividing image data into multiple parts, processing them in parallel, and then combining them for output, the problem of long processing time for large images is solved, while maintaining the ability to extract small features and improving processing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AXIS
- Filing Date
- 2021-11-22
- Publication Date
- 2026-07-31
AI Technical Summary
When processing large images, the computation time and number of calculations of existing convolutional neural networks increase significantly, resulting in low processing efficiency, and reducing the resolution will result in the loss of small feature information.
Image data is divided into multiple parts and processed in parallel through multiple convolutional neural network layers. Each part overlaps, and the combined output is then further processed to extract features.
It reduces the total processing time while maintaining the ability to extract small features, thus reducing the computational load and improving processing efficiency.
Smart Images

Figure CN114627301B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to image processing, and more particularly to capturing image data representing a scene and using a convolutional neural network to process the image data to extract features related to objects in the scene. Background Technology
[0002] Convolutional neural networks can be used when processing images of a scene to extract features related to objects within the scene. In this process, the number of computations required increases significantly with the image size (i.e., the number of pixels). Therefore, processing time increases significantly with the image size. Examples of such large images are high-resolution images and panoramic images representing wide scenes. One way to reduce the number of computations is to reduce the resolution of the image being processed. However, this reduces the likelihood of extracting small features related to objects within the scene. Summary of the Invention
[0003] The purpose of this invention is to facilitate time-efficient processing of image data representing a scene to extract features related to objects in the scene, while maintaining the possibility of extracting small features related to objects in the scene.
[0004] According to a first aspect, a method is provided for processing image data representing a scene using a convolutional neural network to extract features related to objects in the scene. The method includes: processing two or more portions of corresponding parts of two or more parts representing the scene of the image data through a first number of layers of the convolutional neural network using corresponding circuits of two or more circuits to form two or more outputs, wherein the two or more parts of the scene partially overlap. The method further includes: combining the two or more outputs to form a combined output; and processing the combined output through a second number of layers of the convolutional neural network using one of the two or more circuits to extract features related to objects in the scene.
[0005] Therefore, the processing of image data through a first number of layers is divided across two or more different circuits. Specifically, two or more different portions of the image data are processed through the first number of layers by means of one of the two or more different circuits. The two or more different portions are arranged such that corresponding two or more portions of the scene they represent partially overlap. In other words, each portion of the image data representing a part of the scene partially overlaps with a portion of the scene represented by another portion of the image data.
[0006] Two or more portions of image data are processed by a first number of layers of a convolutional neural network on corresponding circuits in two or more circuits, enabling the processing of a portion in parallel, which reduces the total processing time.
[0007] Two or more circuits can be equal in number to two or more portions of the image data. In this case, the processing of the combined output by a second number of layers of a convolutional neural network will be achieved by means of a circuit in one of the two or more portions of the image data that has been processed by a first number of layers of the convolutional neural network.
[0008] The number of two or more circuits can be greater than the number of two or more portions of the image data. In this case, the processing of the combined output by the second number of layers of the convolutional neural network can be accomplished by means of one of the two or more circuits that does not process any one of the two or more portions of the image data by the first number of layers of the convolutional neural network.
[0009] By partially overlapping two or more parts of a scene, the corresponding two or more parts of the image data will include image data representing the partially overlapping portions of the scene. This makes it possible to process the two or more parts of the image data independently on the corresponding circuits of two or more circuits through a first number of layers of a convolutional neural network.
[0010] For a given first number of layers, the size by which two or more parts of the scene partially overlap can be selected such that processing the first number of layers is covered. The size of the partial overlap of the two or more parts of the scene can be further based on the filter (kernel) size and stride of the convolutions in each of the first number of layers of the convolutional neural network. Alternatively, for a given size of partial overlap of two or more parts of the scene, the first number of layers can be selected such that processing the first number of layers is covered. The first number of layers can be further based on the filter size and stride of the convolutions in each of the first number of layers of the convolutional neural network.
[0011] Furthermore, the size of the first number of layers and the partial overlap of two or more parts of the scene can be based on the filter size and stride of each convolution of the first number of layers of the convolutional neural network.
[0012] Two or more circuits may consist of a first circuit and a second circuit. The action of processing two or more portions of the image data may then include: processing a first portion of the image data representing a first part of a scene through a first number of layers of a convolutional neural network using the first circuit to form a first output; and processing a second portion of the image data representing a second part of a scene through a first number of layers of a convolutional neural network using the second circuit to form a second output, wherein the first and second parts of the scene partially overlap. The action of combining the two or more outputs then includes combining the first and second outputs to form a combined output, and the action of processing the combined output includes processing the combined output through a second number of layers of a convolutional neural network using one of the first and second circuits to extract features related to objects in the scene.
[0013] Therefore, the processing of image data is divided into a first circuit and a second circuit by means of a first number of layers. Specifically, the first portion and the second portion of image data are processed by means of the first circuit and the second circuit, respectively, through the first number of layers. The first portion and the second portion of image data are arranged such that the corresponding first portion and the corresponding second portion of the scene they represent partially overlap. In other words, the first portion of image data represents a first portion of the scene that partially overlaps with the second portion of the scene represented by the second portion of image data.
[0014] The first part of the image data may be image data captured by a first image sensor, and the second part of the image data may be image data captured by a second image sensor. In other words, the first image sensor captures a first part of the scene, and the second image sensor captures a second part of the scene, and the first and second parts of the scene partially overlap.
[0015] This is advantageous, for example, when two sensors capture image data representing two parts of a scene in order to form a panoramic image by combining the image data representing the two parts of the scene. In this case, the two parts of the scene typically partially overlap, and the overlap between the two parts of the scene is used (e.g., by means of mixing image data from both sensors) to reduce the risk of seam lines, etc., in the region of the panoramic image corresponding to the boundary between the images captured by the two sensors in the scene. Image data from the first image sensor and the second image sensor can be processed directly through a first number of layers by means of the first circuit and the second circuit, respectively, without first combining them to form an image, for example, by mixing image data from both sensors.
[0016] Two or more circuits can consist of four circuits. Then, the actions of processing two or more portions of the image data can include:
[0017] With the aid of a first circuit, a first portion of the image data representing a first part of the scene is processed through a first number of layers of a convolutional neural network to form a first output;
[0018] With the aid of a second circuit, a second portion of the image data representing a second part of the scene is processed through a first number of layers of a convolutional neural network to form a second output, wherein the first and second portions of the scene partially overlap.
[0019] With the aid of a third circuit, the third part of the image data representing the third part of the scene is processed through a first number of layers of a convolutional neural network to form a third output, wherein the second and third parts of the scene partially overlap; and
[0020] With the aid of a fourth circuit, the fourth part of the image data representing the scene is processed through a first number of layers of a convolutional neural network to form a fourth output, wherein the third and fourth parts of the scene partially overlap.
[0021] Then, the action of combining two or more outputs includes combining the first, second, third and fourth outputs of the processing from the first and second parts of the image data to form a combined output, and the action of processing the combined output includes processing the combined output through a second number of layers of a convolutional neural network by means of one of the first, second, third and fourth circuits to extract features related to objects in the scene.
[0022] Compared to using two circuits to process image data divided into two parts, using four circuits to process image data divided into four parts reduces the number of computations per circuit. However, since the four parts of the image data represent four partially overlapping portions of the scene, the number of computations required to process the four parts of the image data using four circuits is not half the number of computations required to process the two parts of the image data using two circuits.
[0023] In an alternative approach, when two or more circuits consist of four circuits, the actions for processing two or more portions of the image data can include:
[0024] With the aid of a first circuit, a first portion of the image data representing a first part of the scene is processed through a first number of layers of a convolutional neural network to form a first intermediate output;
[0025] With the aid of a second circuit, the second part of the image data representing the second part of the scene is processed by a first number of layers of a convolutional neural network to form a second intermediate output, wherein the first part and the second part of the scene partially overlap.
[0026] The first intermediate output and the second intermediate output from the processing of the first part and the second part of the image data are combined to form a first intermediate combined output;
[0027] With the aid of a third circuit, the third part of the image data representing the third part of the scene is processed by a first number of layers of a convolutional neural network to form a third intermediate output, wherein the second and third parts of the scene partially overlap.
[0028] With the aid of a fourth circuit, the fourth part of the image data representing the fourth part of the scene is processed by a first number of layers of a convolutional neural network to form a fourth intermediate output, wherein the third and fourth parts of the scene partially overlap.
[0029] The third and fourth intermediate outputs from the processing of the third and fourth parts of the image data are combined to form a second intermediate combined output;
[0030] Using one of the first, second, third, and fourth circuits, the first intermediate combined output is processed through a third number of layers of a convolutional neural network to form the first output; and
[0031] The second intermediate combined output is processed by a third number of layers of a convolutional neural network using one of the first, second, third, and fourth circuits to form the second output.
[0032] Then, the action of combining two or more outputs includes combining the first output and the second output to form a combined output, and the action of processing the combined output includes processing the combined output through a second number of layers of a convolutional neural network by means of one of the first circuit, the second circuit, the third circuit and the fourth circuit to extract features related to objects in the scene.
[0033] For a given first number of layers, the size by which the first and second parts of the scene, the second and third parts of the scene, and the third and fourth parts of the scene partially overlap can be selected such that the processing of the first number of layers is covered. Additionally, for a given third number of layers, the size by which the second and third parts of the scene partially overlap can be selected such that the processing of the third number of layers is also covered in addition to the processing of the first number of layers. The size by which the first and second parts of the scene, the second and third parts of the scene, and the third and fourth parts of the scene partially overlap can be further based on the filter size and stride of each convolution in the first number of layers of the convolutional neural network and the filter size and stride of each convolution in the third number of layers of the convolutional neural network.
[0034] Alternatively, for a given size in which the first and second parts of the scene, the second and third parts of the scene, and the third and fourth parts of the scene partially overlap, a first number of layers can be selected such that processing the first number of layers is covered and processing the third number of layers is covered. The first number of layers and the third number of layers can be further based on the filter size and stride of each convolution of the first number of layers of the convolutional neural network and the filter size and stride of each convolution of the third number of layers of the convolutional neural network.
[0035] The first part of the image data may be image data captured by a first image sensor, the second part of the image may be image data captured by a second image sensor, the third part of the image data may be image data captured by a third image sensor, and the fourth part of the image data may be image data captured by a fourth image sensor, wherein the overlap between the second and third parts of the scene is greater than the overlap between the first and second parts and greater than the overlap between the third and fourth parts.
[0036] This is advantageous, for example, when four sensors capture image data representing four parts of a scene in order to form a panoramic image by combining the image data representing the four parts of the scene. In this case, the four parts of the scene typically partially overlap, and the overlap between the four parts of the scene is used (e.g., by mixing image data from both sensors) to reduce the risk of seams, etc., in the regions of the panoramic image corresponding to the boundaries between the images captured by the four sensors in the scene. The image data from the first, second, third, and fourth image sensors can be processed directly through a first number of layers by means of the first, second, third, and fourth circuits, respectively, without first combining them to form an image, for example, by mixing the image data from the four sensors.
[0037] The overlap between the second and third parts of the scene is greater than the overlap between the first and second parts, and also greater than the overlap between the third and fourth parts, which enables the first and second intermediate combined outputs to be processed by a third number of layers of the convolutional neural network.
[0038] Image data can be image data captured by an image sensor.
[0039] In an alternative, the image data can be image data captured by at least two image sensors. Then, each portion of the image data can be captured by a single image sensor of the at least two image sensors, and the at least two image sensors can be arranged such that image data representing corresponding portions of the scene that partially overlap are captured by the respective image sensor.
[0040] Image data can also be combined, such that two or more portions of the image data are captured by the same image sensor, and two or more portions of the image data are captured by separate image sensors.
[0041] The size of the overlap between the first number of layers and two or more parts of the scene can be selected such that the processing of the first number of layers is covered.
[0042] According to a second aspect, a non-transitory computer-readable storage medium is provided, on which instructions are stored, which, when executed by a processing device, are used to implement the method according to the first aspect.
[0043] According to a third aspect, an apparatus is provided for processing image data representing a scene using a convolutional neural network to extract features related to objects in the scene. The apparatus includes two or more circuits configured to perform a first processing function, which is configured to process two or more portions of corresponding portions of two or more parts of the scene representing the image data through a first number of layers of the convolutional neural network by means of corresponding circuits in the two or more circuits, to form two or more outputs, wherein the two or more parts of the scene partially overlap. The two or more circuits are further configured to perform: a combination function, configured to combine the two or more outputs to form a combined output; and a second processing function, configured to process the combined output through a second number of layers of the convolutional neural network by means of one of the two or more circuits, to extract features related to objects in the scene.
[0044] Image data can be image data captured by an image sensor.
[0045] In an alternative, the image data can be image data captured by at least two image sensors.
[0046] The apparatus may include four or more circuits. A first processing function is then further configured to: process four portions of image data captured by corresponding image sensors of four image sensors and representing corresponding portions of four parts of a scene through a first number of layers of a convolutional neural network, using corresponding circuits of the four or more circuits, to form four intermediate outputs, wherein the four parts of the scene partially overlap; and combine two of the four intermediate outputs into a first intermediate combined output, and combine the remaining two of the four intermediate outputs into a second intermediate combined output. The apparatus is further configured to perform a third processing function, which is configured to process corresponding intermediate combined outputs of the first and second intermediate combined outputs through a third number of layers of a convolutional neural network using two of the four or more circuits, to form a first output and a second output, respectively.
[0047] The aforementioned actions according to the method of the first aspect also apply to the apparatus of the third aspect when applicable.
[0048] According to a fourth aspect, a system is provided for capturing image data representing a scene and processing the image data using a convolutional neural network to extract features related to objects in the scene. The system includes means according to a third aspect, wherein the means includes four or more circuits. The system further includes a camera for capturing image data representing the scene. The camera includes: a first image sensor arranged to capture image data representing a first portion of the scene; a second image sensor arranged to capture image data representing a second portion of the scene; a third image sensor arranged to capture image data representing a third portion of the scene; and a fourth image sensor arranged to capture image data representing a fourth portion of the scene. The first, second, third, and fourth image sensors are arranged such that the overlap between the second and third portions of the scene is greater than the overlap between the first and second portions and greater than the overlap between the third and fourth portions.
[0049] The actions described above according to the method of the first aspect also apply to the system of the fourth aspect when applicable.
[0050] Further applicability of the invention will become apparent from the detailed description given below. However, it should be understood that while the detailed description and specific examples indicate preferred embodiments of the invention, they are given by way of example only, as various variations and modifications within the scope of the invention based on this detailed description will become apparent to those skilled in the art.
[0051] Therefore, it will be understood that the present invention is not limited to the specific components of the described apparatus or the operation of the described method, as such apparatus and method can vary. It will also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It must be noted that, as used in the specification and appended claims, the articles “a,” “the,” and “the” are intended to indicate the presence of one or more elements unless the context explicitly indicates otherwise. Thus, for example, references to “a unit” or “the unit” can include several devices, etc. Furthermore, the words “comprising,” “including,” “containing,” and similar wording do not exclude other elements or steps. Attached Figure Description
[0052] The above and other aspects of the invention will now be described in more detail with reference to the accompanying drawings. The drawings should not be considered limiting, but rather for explanation and understanding. The same reference numerals always denote the same elements.
[0053] Figure 1 This is a flowchart of an embodiment of a method for using convolutional neural networks to process image data representing a scene in order to extract features related to objects in the scene.
[0054] Figure 2 This is a flowchart of another embodiment of a method for using convolutional neural networks to process image data representing a scene in order to extract features related to objects in the scene.
[0055] Figure 3 This is a flowchart of yet another embodiment of a method for using convolutional neural networks to process image data representing a scene in order to extract features related to objects in the scene.
[0056] Figure 4a and Figure 4b This is a flowchart of yet another embodiment of a method for using convolutional neural networks to process image data representing a scene in order to extract features related to objects in the scene.
[0057] Figure 5 This is a schematic diagram of multiple pixels of an image related to the processing of layers in a convolutional neural network.
[0058] Figure 6 This is a schematic block diagram of an embodiment of a device for processing image data representing a scene using a convolutional neural network to extract features related to objects in the scene.
[0059] Figure 7 This is a schematic block diagram of an embodiment of a device for processing image data representing a scene using a convolutional neural network to extract features related to objects in the scene.
[0060] Figure 8This is a schematic block diagram of an embodiment of a system for capturing image data representing a scene and using a convolutional neural network to process the image data to extract features related to objects in the scene.
[0061] Figure 9 This is a schematic block diagram of an embodiment of a system for capturing and processing image data representing a scene. Detailed Implementation
[0062] The invention will now be described more fully below with reference to the accompanying drawings, in which presently preferred embodiments of the invention are illustrated. However, the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided for thoroughness and completeness and to convey the scope of the invention to those skilled in the art.
[0063] Figure 1 This is a flowchart of an embodiment of a method 100 for processing image data representing a scene using a convolutional neural network to extract features related to objects in the scene. For example, if the image data representing the scene involves high-resolution images or panoramic images, a large amount of data needs to be processed through layers of the convolutional neural network, and therefore the computational load will increase. Therefore, the image data can be divided into corresponding parts representing two or more parts of the scene, and processed by a first number of layers of the convolutional neural network S110 by means of corresponding circuits in two or more circuits to form two or more parts of two or more outputs. Therefore, these two or more parts of the image data can be processed in parallel by the first number of layers of the convolutional neural network, and thus the total processing time will be reduced. A circuit is an arrangement of circuits or circuit systems. It can be arranged, for example, on a chip, and can further include software for performing the processing or otherwise arranged together with the software for performing the processing.
[0064] However, when image data is divided into two or more parts representing a scene, some objects in the scene will appear at the boundary between the two parts of the scene, and thus will be represented in the image data at the boundary between the two parts of the scene. Therefore, due to the inherent characteristics of convolutional neural networks, each of the two or more parts of the image data needs to include image data representing the overlapping portion of each of the adjacent parts of the scene, so as to also be able to identify features related to objects in the scene appearing at the boundary between the two or more parts of the scene. Therefore, the two or more parts of the scene are arranged such that they partially overlap. The size of the partial overlap between the two or more parts of the scene and the first number of layers in the convolutional neural network are interdependent. The larger the overlap, the larger the first number of layers can be, and the larger the first number of layers to be processed, the larger the overlap needs to be. Therefore, if processing through a given first number of layers is possible, the size of the partial overlap between the two or more parts of the scene needs to be chosen large enough to cover such processing. The size by which two or more parts of a scene partially overlap is typically determined as the number of pixels required for the two or more parts of image data to partially overlap in order to cover the processing of a first number of layers in a convolutional neural network. The characteristics of the first layer of the convolutional neural network that influence the size by which two or more parts of a scene need to partially overlap are the filter size and stride of the convolutions in each of the first number of layers in the convolutional neural network. Conversely, for a given size of partial overlap between two or more parts of a scene, the filter size and stride of the convolutions in each layer of the convolutional neural network will dominate the first number of layers that can be processed.
[0065] The partial overlap between parts of a scene corresponds to the pixel overlap between parts of image data. If a single image sensor has been used for both parts of the image data, the overlapping pixels between the two parts of the image data can be the same pixels. Alternatively, overlapping pixels can be captured from two different image sensors, such that the overlapping pixels are different for the two parts of the image data, but represent the same part of the scene. To illustrate how the filter size of a convolution in a layer affects the amount of pixel overlap required for processing layers in a convolutional neural network, refer to... Figure 5 , Figure 5This is a schematic diagram of layers 510, 520, and 530, each involving a 5×5 pixel subset of the image related to the processing of the layers in the convolutional neural network. The convolutions in each of layers 520 and 530 are 3×3. This means that the features in these layers have a 3×3 pixel data locality. To create a "pixel" in a layer, 3×3 pixels are needed in the previous layer. Figure 5 In this context, this is illustrated by pixels in the second layer 520 that are illustrated with a grid pattern and require 3×3 pixels in the first layer 510 that are illustrated with a grid pattern. Similarly, pixels in the third layer 530 that are illustrated with a stripe pattern require 3×3 pixels in the second layer 520 that are illustrated with both stripe and grid patterns. Furthermore, as... Figure 5 As illustrated in the diagram, the pixels in the third layer 530, represented by a striped pattern, require 5×5 pixels in the first layer 510, represented by both striped and grid patterns. Therefore, if the image is to be divided into two or more parts, and the boundary between two adjacent parts is perpendicular to the vertical line 540 between the second and third pixel columns in the first, second, and third layers 510, then an additional pixel column must be processed by the second layer 520 for each of the two adjacent parts to create pixels for the third layer 530 at the boundary corresponding to the vertical line 540. Furthermore, two additional pixel columns must be processed by the second layer 520 for each of the two adjacent parts to create pixels for the third layer 530 at the boundary corresponding to the vertical line 540. The number of additional pixel columns required in the first layer 510 is based on the filter size of the convolution in the second and third layers 520. In this case, the number of additional columns is (3-1) / 2 + (3-1) / 2. For an N×N convolution in a layer, the number of additional columns required from the previous layer to create pixels at the boundaries of the layer will be (N-1) / 2. Furthermore, the layer stride (i.e., the number of pixels the N×N filter moves within the layer when creating pixels for the next layer) also affects the number of additional columns required in that layer to create pixels at the boundaries of the next layer. Figure 5 In this context, the stride is 1 for both the first layer 510 and the second layer 520. If the stride is 2 in a layer with N×N convolutions, the number of additional columns required relative to that layer in order to create pixels at the boundaries in the next layer will be ((N-1) / 2)×2.
[0066] It should be noted that in order to process a portion of image data through the first number of layers of the convolutional neural network, the overlapping of adjacent portions of the image data also needs to be processed through that portion of the image data by the first number of layers of the convolutional neural network. Therefore, the additional pixels between two adjacent portions of two or more portions of the image data are twice the number of additional pixels of a single portion. Furthermore, the more portions of the image data there are, the more boundaries exist between the portions, and thus the more additional pixels are required. However, the required number of additional pixels depends on the length of the boundaries. Therefore, the division should preferably make the boundaries as short as possible. For example, images with a width greater than their height (e.g., panoramic images) are preferably divided by width, with vertical boundaries between the portions.
[0067] The processing of image data by a first number of layers of a convolutional neural network is divided into separate processing of two or more portions of the image data by the first number of layers of the convolutional neural network on corresponding circuits in two or more circuits. This allows for the parallel processing of a portion, reducing the overall processing time. Generally, the more portions of the image data are divided and processed by the first number of layers of the convolutional neural network on the corresponding separate circuits, the greater the reduction in overall processing time. However, since the size of the overlap required between the portions depends on the characteristics of the first number of layers of the convolutional neural network, the size will be the same regardless of the number of portions into which the image data is divided. Therefore, the number of portions of the image data to be processed on two different circuits will increase with the number of portions into which the image data is divided, and thus the larger the number of portions into which the image data is divided, the less significant the reduction in processing time becomes.
[0068] Since the amount of image data that needs to be processed by two different circuits through a first number of layers of a convolutional neural network is determined by the size (width) and length of the overlap, and the required size (width) of the overlap is given by the first number of layers, it is preferable to divide the image data into two or more parts such that the overlap is as short as possible. For images where the width is greater than the height (e.g., panoramic images), it is preferable to divide vertically such that the two or more parts of the image data represent corresponding portions of two or more vertical sections of the scene, wherein the vertical sections overlap at the vertical boundaries between adjacent vertical sections.
[0069] When an image representing a scene is processed through a first number of layers, the partial overlap between two or more parts of the scene corresponds to the partial overlap between two or more parts of the image data. The required magnitude of the partial overlap between the two or more parts of the image data representing the scene depends on the number of layers, as well as the filter size and stride of each layer. Therefore, given the filter size and stride of each layer, the required magnitude of the partial overlap will depend on the number of layers. Thus, if the overlap increases, the number of layers that can be processed in parallel (i.e., the first number of layers) can increase. On the other hand, a large overlap between two or more parts of the image data will result in a larger amount of the same image data being processed by two circuits. Instead of processing the two or more parts of the image data through all of the first number of layers on two or more circuits respectively, an alternative method can be used. In the alternative method, the two or more parts of the image data with reduced partial overlap are processed through a subset of the first number of layers on the corresponding circuits in the two or more circuits. Subsequently, data obtained from processing on each of two or more circuits (data required for processing subsequent subsets of the first number of layers on other circuits in the two or more circuits) is provided to the other circuits in the two or more circuits, i.e., the necessary overlapping image data between the copying circuits. This step is repeated until processing has been completed through all of the first number of layers. In this method, redundant processing is reduced because the initial overlap can be less than the overlap required to process two or more portions of the image data through all of the first number of layers on two or more circuits respectively. However, the smaller the initial overlap, the larger the number of subsets of the first number of layers that need to be processed using intermediate data exchange between the two or more circuits.
[0070] The image data may comprise two or more portions of image data relating to an image of a scene. The image may have been captured by one image sensor or generated by combining images captured by more than one image sensor. In this case, the image data can be divided into two or more portions, each of which partially overlaps with an adjacent portion of the image data; that is, the adjacent portion corresponding to the overlap comprises a portion of the same image data. This overlap then corresponds to the overlap between two or more portions of the scene represented by the two or more portions of the image data.
[0071] In an alternative, the two or more portions of the image data may involve image data from corresponding images in two or more images captured by two or more image sensors, wherein the two or more image sensors capture corresponding portions of two or more parts of a scene, and wherein the two or more parts of the scene partially overlap. In this case, the two or more portions of the image data will not include a portion of the same image data as the adjacent portion corresponding to the overlapping image data, because the overlapping image data was captured by different image sensors for different portions of the image data. However, if the image sensors are calibrated, the overlapping corresponding portions of the image data will be similar, except for noise. By using image data directly from two or more image sensors to process through a first number of layers, it is not necessary to exchange the image data captured by the two or more image sensors between two or more circuits before processing through the first number of layers.
[0072] Furthermore, combination is also feasible. In this case, at least two of the two or more portions of the image data relate to image data captured by corresponding image sensors of at least two image sensors, and at least two of the two or more portions of the image data relate to image data captured by one image sensor. For example, for three portions of the image data, one of the three portions may relate to one image sensor, and two of the three portions may relate to different image sensors.
[0073] When two or more portions of image data have been processed (S130) through a first number of layers of a convolutional neural network via corresponding circuits in two or more circuits to form two or more outputs, the two or more outputs are combined (S120) to form a combined output. The combination of two or more outputs can be achieved by concatenating the two outputs together and cutting data related to overlap. The combined output is then processed (S130) through a second number of layers of the convolutional neural network by one of the two or more circuits to extract features related to objects in the scene. Using one of the two or more circuits used to process one portion of the image data through the first number of layers of the convolutional neural network and another circuit used to process the combined output through the second number of layers of the convolutional neural network means that the total number of circuits required is the same as the number of portions of the image data. Then, during the time it takes to process the combined output through the second number of layers of the convolutional neural network, one of the two or more circuits will be busy, while the other circuits will not be busy. During this time, new processing of corresponding portions of the two or more portions of image data related to one or more other images (such as one or more next image frames of a video stream) can be initiated by a circuit in one of the two or more circuits that is not currently processing the combined output. Two or more circuits used to process the combined output through a second number of layers of a convolutional neural network can alternate between different image frames, such as between different image frames in a video stream, to achieve load balancing between the two or more circuits. Conversely, if the total number of circuits is greater than the number of portions of image data, circuits other than those used to process corresponding portions of the two or more portions of image data through the first number of layers of the convolutional neural network can be used to process the combined output through the second number of layers of the convolutional neural network. In this case, circuits other than those used to process the combined output through the second number of layers of the convolutional neural network can be directly used to begin processing two or more portions of image data associated with another image (such as the next image frame in a video stream). This reduces the need for load balancing between the two or more circuits, but on the other hand, requires at least one additional circuit.
[0074] Figure 2 This is a flowchart of another embodiment of a method 200 that uses a convolutional neural network to process image data representing a scene to extract features related to objects in the scene. Specifically, Figure 2This embodiment involves two parts of image data, with processing performed on corresponding circuits in two separate circuits. A first part of the image data representing a first part of a scene is processed (S210) by a first circuit through a first number of layers of a convolutional neural network to form a first output. A second part of the image data representing a second part of a scene is processed (S215) by a second circuit through a first number of layers of a convolutional neural network to form a second output. The first and second parts of the scene partially overlap. The first and second outputs are then combined (S220) to form a combined output. Combining two or more outputs can be achieved by concatenating the two outputs and cutting data related to the overlap. The combined output is processed (S230) by a circuit in one of the first and second circuits through a second number of layers of a convolutional neural network to extract features related to objects in the scene. Alternatively, the combined output can be processed (S230) by a circuit other than the first and second circuits through a second number of layers of a convolutional neural network.
[0075] For example, if image frames of a video stream are to be processed by a first circuit and a second circuit through a first number of layers and a second number of layers of a convolutional neural network, the first circuit can process a first portion of the image data associated with the first image frame through the first number of layers, and the second circuit can process a second portion of the image data associated with the first image frame through the first number of layers. Then, the first circuit can be used to process the combined output through the second number of layers. Once the first circuit has completed processing the combined output associated with the first frame through the second number of layers, the first circuit can begin processing the first portion of the image data associated with the second image frame through the first number of layers. In parallel with the first circuit processing the combined output through the second number of layers and processing the first portion of the image data associated with the second image frame through the first portion of the image data, the second circuit can be used to process the second portion of the image data associated with the second image frame through the first number of layers, and depending on the processing time of the first circuit processing the combined output through the second number of layers and processing the first portion of the image data associated with the second image frame, the second circuit can continue to be used in parallel to process the first portion of the image data associated with a third image frame, and so on. Once the first circuit has completed processing the combined output associated with the first image frame through a second number of layers and has completed processing the first portion of the image data associated with the second image frame through a second number of layers, the second circuit can begin processing the combined output associated with the second image frame through a second number of layers. Once the second circuit has completed processing the combined output associated with the second frame through a second number of layers, the second circuit can begin processing the second portion of the image data associated with the third image frame through a first number of layers. Then, the circuits of the first and second circuits can alternate between image frames to achieve load balancing.
[0076] The sizes of the first and second portions of the image data associated with each image frame can also differ. For example, when a first circuit that processes the first portion through a first number of layers also processes the combined output associated with the first image frame through a second number of layers, the first portion of the image data associated with the first frame can be smaller than the second portion, for example, half the size. When a second circuit that processes the second portion through a first number of layers also processes the combined output associated with the second image frame through a second number of layers, the first portion of the image data associated with the second frame can be larger than the second portion, for example, twice the size. This can reduce the number of image frames that need to be processed in parallel.
[0077] Conversely, if there are one or more additional circuits besides the first and second circuits, one of these additional circuits can be used to process the combined output through a second number of layers of the convolutional neural network. In this case, the first and second circuits can be directly used to process a first portion of the image data associated with the second image frame and a second portion of the image data associated with the second image frame, respectively. This reduces the need for load balancing across the image frames for two or more circuits, but on the other hand, it requires at least one or more additional circuits. The number of the one or more additional circuits can be selected based on the processing time for processing a portion of the image data through the first number of layers and the processing time for processing the combined output through the second number of layers. For example, if the processing time of the first number of layers is half the processing time of the second number of layers, then using two additional circuits to alternately process the combined output through the second number of layers will balance the load relative to the first and second circuits.
[0078] The first and second portions of the image data may relate to image data of an image of the scene. The image may have been captured by a single image sensor or may be generated by combining images captured by more than one image sensor. In this case, some of the image data in the first portion of the image data will be identical to some of the image data in the second portion of the image data; that is, image data representing partial overlap between the first and second portions of the scene.
[0079] Alternatively, the first portion of the image data can be image data captured by a first image sensor, and the second portion of the image data can be image data captured by a second image sensor. In other words, the first image sensor captures a first portion of the scene, and the second image sensor captures a second portion of the scene, and the first and second portions of the scene partially overlap. In this case, the first and second portions of the image data will not include some of the same data as the image data representing the partial overlap between the first and second portions of the scene, because the image data representing this overlap is captured by different sensors for the first and second portions of the image data. However, if the first and second image sensors are calibrated, the image data corresponding to the overlap will be similar for the first and second portions of the image data, except for noise. The image data from the first and second image sensors can be used as the first and second portions of the image data, respectively, and processed directly through a first number of layers by means of the first and second circuits, respectively, without first combining them to form an image. By using image data directly from the first image sensor and the second image sensor as the first part and the second part of the time image data, respectively, it is not necessary to exchange the image data captured by the first image sensor and the second image sensor between the first circuit and the second circuit before processing through a first number of layers.
[0080] For example, if the original image representing the scene has a size (resolution) of 8192 × 2048 pixels, and an image representing the scene with a size (resolution) of 2048 × 512 pixels is to be processed through layers of a convolutional neural network (such as MobileNet-SSD), then the original image must first be downsized by a factor of 4. Furthermore, if two parts of the image data to be processed are to be processed through a first number of layers, and that first number of layers is 48, then each of the two parts of the image data needs to overlap the other part by 148 pixels (relative to the original image's resolution of 592 pixels). Therefore, for two parts of the image data divided into two equal-sized parts according to the image width, the total overlap, including each part of the image data, must then be 1172 × 512 pixels. Due to this overlap, processing two parts of the image data through 48 layers on corresponding circuits in two circuits requires more computation than processing the entire image through 48 layers on a single circuit. However, since the processing can be performed in parallel on the two circuits, the total time will still be reduced.
[0081] Figure 3This is a flowchart of yet another embodiment of a method 300 that uses a convolutional neural network to process image data representing a scene to extract features related to objects in the scene. Specifically, Figure 3 An embodiment involving four parts of image data, wherein processing is performed on corresponding circuits in four circuits. A first part of the image data representing a scene is processed by a first circuit through a first number of layers of a convolutional neural network (S310) to form a first output. A second part of the image data representing a scene is processed by a second circuit through a first number of layers of a convolutional neural network (S312) to form a second output, wherein the first and second parts of the scene partially overlap. A third part of the image data representing a scene is processed by a third circuit through a first number of layers of a convolutional neural network (S314) to form a third output, wherein the second and third parts of the scene partially overlap. A fourth part of the image data representing a scene is processed by a fourth circuit through a first number of layers of a convolutional neural network (S316) to form a fourth output, wherein the third and fourth parts of the scene partially overlap. The first, second, third, and fourth outputs from the processing of the first and second parts of the image data are combined (S320) to form a combined output. The combined output is processed by a second number of layers of a convolutional neural network through one of the first, second, third, and fourth circuits (S330) to extract features related to objects in the scene.
[0082] Compared to using two circuits to process image data divided into two parts, using four circuits to process image data divided into four parts reduces the number of computations per circuit. However, since the four parts of the image data represent four partially overlapping portions of the scene, the number of computations required to process the four parts of the image data using four circuits is not half the number of computations required to process the two parts of the image data using two circuits.
[0083] The first, second, third, and fourth parts of the image data may relate to image data of an image of the scene. The image may have already been captured by one image sensor, or it may be generated by combining images captured by more than one image sensor into a single image, as similar to [the image data described above]. Figure 2 The image data is described in two parts as follows.
[0084] Alternatively, the first, second, third, and fourth portions of the image data can be image data captured by the respective image sensors among the first, second, third, and fourth image sensors, as similar to those described above. Figure 2The image data is described as being in two parts. In this alternative, the image data captured by the respective image sensors among the first, second, third, and fourth image sensors does not need to be first combined into an image and then divided into parts of the image data; instead, it can be directly used as the respective parts of the first, second, third, and fourth parts of the image data. Therefore, it is not necessary to exchange the image data captured by the first, second, third, and fourth image sensors between the first, second, third, and fourth circuits before processing through a first number of layers.
[0085] Figure 4a and Figure 4b This is a flowchart of yet another embodiment of a method 400 that uses a convolutional neural network to process image data representing a scene to extract features related to objects in the scene. Specifically, Figure 4a and Figure 4b This embodiment involves four parts of image data, where processing is first performed by corresponding circuits in four circuits and then by corresponding circuits in two of the four circuits. The first part of the first part of the image data representing a scene is processed by a first circuit through a first number of layers of a convolutional neural network (S410) to form a first intermediate output. The second part of the second part of the image data representing a scene is processed by a second circuit through a first number of layers of a convolutional neural network (S411) to form a second intermediate output, wherein the first and second parts of the scene partially overlap. The first and second intermediate outputs from the processing of the first and second parts of the image data are combined (S412) to form a first intermediate combined output. The third part of the third part of the image data representing a scene is processed by a third circuit through a first number of layers of a convolutional neural network (S413) to form a third intermediate output, wherein the second and third parts of the scene partially overlap. The fourth part of the fourth part of the image data representing a scene is processed by a fourth circuit through a first number of layers of a convolutional neural network (S413) to form a fourth intermediate output, wherein the third and fourth parts of the scene partially overlap. The third and fourth intermediate outputs from the processing of the third and fourth parts of the image data are combined in step S415 to form a second intermediate combined output. The first intermediate combined output is passed through a third number of layers (relative to) of the convolutional neural network by means of one of the first, second, third, and fourth circuits. Figure 4a and Figure 4bThe intermediate number of layers (referred to as "intermediate number layers") is processed in S416 to form the first output. The second intermediate combined output is processed in S417 by passing through the intermediate number of layers of the convolutional neural network via one of the first, second, third, and fourth circuits to form the second output. The first and second outputs are combined in S420 to form a combined output. The combined output is passed through the second number of layers of the convolutional neural network (relative to the intermediate number of layers) via one of the first, second, third, and fourth circuits. Figure 4a and Figure 4b The "last layer" (referred to as the "last number layer") is processed by S430 to extract features related to objects in the scene.
[0086] For a given first number of layers and an intermediate number of layers in a convolutional neural network, the first and second parts of the scene, as well as the third and fourth parts of the scene, need to partially overlap by a certain size so that the processing of the first number of layers is covered, and the second and third parts of the scene also need to partially overlap by a certain size so that the processing of the first number of layers and the processing of the intermediate number of layers are covered. Specifically, the size by which the first and second parts of the scene, as well as the third and fourth parts of the scene, need to partially overlap is based on the filter size and stride of each convolution in the first number of layers of the convolutional neural network. The size by which the second and third parts of the scene need to partially overlap is based on the filter size and stride of each convolution in the first number of layers of the convolutional neural network and the filter size and stride of each convolution in the intermediate number of layers of the convolutional neural network.
[0087] Alternatively, for a given size where the first and second parts of the scene, the second and third parts of the scene, and the third and fourth parts of the scene partially overlap, processing of a first number of layers is covered, and processing of an intermediate number of layers is covered. Specifically, the covered first number of layers and intermediate number of layers are based on the filter size and stride of each convolution of the first number of layers of the convolutional neural network and the filter size and stride of each convolution of the intermediate number of layers of the convolutional neural network.
[0088] The first, second, third, and fourth parts of the image data may relate to image data of an image of the scene. The image may have already been captured by one image sensor, or it may be generated by combining images captured by more than one image sensor into a single image, as similar to [the image data described above]. Figure 2 The image data is described in two parts as follows.
[0089] Alternatively, the first, second, third, and fourth portions of the image data can be image data captured by the respective image sensors among the first, second, third, and fourth image sensors, as similar to those described above. Figure 2 The image data is described as two parts. In this alternative, the image data captured by the respective image sensors among the first, second, third, and fourth image sensors does not need to be first combined into one image and then divided into parts of the image data, but can be directly used as the respective parts of the first, second, third, and fourth parts of the image data. Therefore, it is not necessary to exchange the image data captured by the first, second, third, and fourth image sensors between the first, second, third, and fourth circuits before processing through a first number of layers. Additionally, since the two combined intermediate outputs are processed through an intermediate number of layers of the convolutional neural network by means of the respective circuits in the four circuits, the overlap between the second and third parts of the scene is greater than the overlap between the first and second parts and greater than the overlap between the third and fourth parts.
[0090] Figure 6This is a schematic block diagram of an embodiment of an apparatus 600 for processing image data representing a scene using a convolutional neural network to extract features related to objects in the scene. The apparatus 600 includes two or more circuits 610 configured to perform a first processing function 661, which is configured to process two or more portions of corresponding portions of two or more parts of a scene representing the image data through a first number of layers of the convolutional neural network by means of corresponding circuits in the two or more circuits 610, to form two or more outputs, wherein the two or more parts of the scene partially overlap. The two or more circuits 610 are further configured to perform a combination function 663 and a second processing function 665, whereby the combination function 663 is configured to combine the two or more outputs to form a combined output, and the second processing function 665 is configured to process the combined output through a second number of layers of the convolutional neural network by means of one of the two or more circuits 610, to extract features related to objects in the scene. One of the two or more circuits 610 that processes the combined output by means of a second number of layers in the convolutional neural network may or may not be a circuit among two or more circuits that process one of two or more portions of the image data representation by means of a first number of layers in the convolutional neural network. In the former case, the number of circuits needed is only as many as the number of portions of the image data that should be processed by the first number of layers in the convolutional neural network, because the second number of layers in the neural network is processed by means of a circuit used to process the first number of layers. In the latter case, one more circuit is needed than the number of portions of the image data that should be processed by the first number of layers, because the second number of layers should be processed by a circuit that does not process a portion of the image data that is processed by the first number of layers.
[0091] Image data can be image data captured by a single image sensor. Alternatively, image data can be image data captured by at least two image sensors.
[0092] Two or more circuits 610 are configured to perform the functions of device 600. Each of the two or more circuits 610 may include a processor (not shown), such as a central processing unit (CPU), microcontroller, or microprocessor. The processor is configured to execute program code, such as program code configured to perform the functions of device 600.
[0093] Device 600 may further include memory 650. Memory 650 may be one or more of a buffer, flash memory, hard disk drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or another suitable device. In a typical arrangement, memory 650 may include non-volatile memory for long-term data storage and volatile memory serving as system memory for two or more circuits 610. Memory 650 may exchange data with two or more circuits 610 via a data bus. Accompanying control lines and address buses may also be present between memory 650 and the two or more circuits 610.
[0094] The functionality of device 600 can be implemented in the form of executable logic routines (e.g., lines of code, software programs, etc.) stored on a non-transitory computer-readable medium (e.g., memory 650) of device 600 and executed by two or more circuits 610 (e.g., using a processor). Furthermore, the functionality of device 600 can be a standalone software application or part of a software application that performs additional tasks associated with device 600. The described functionality can be considered as a method configured to execute by a processing unit (e.g., a processor of two or more circuits 610). Moreover, while the described functionality can be implemented in software, such functionality can also be implemented by dedicated hardware or firmware, or some combination of hardware, firmware, and / or software.
[0095] The functions performed by device 600 and circuit 610 can be further adapted to relate to Figure 1 Method 100 described, about Figure 2 Method 200 described, about Figure 3 Method 300 described and about Figure 4a and Figure 4b The corresponding steps of method 400 are described.
[0096] Figure 7 This is a schematic block diagram of a further embodiment of an apparatus 700 for processing image data representing a scene using a convolutional neural network to extract features related to objects in the scene. Specifically, Figure 6An embodiment involving four circuits 710, 720, 730, and 740 is described. The device 700 is configured to perform a first processing function 761. The first processing function 761 is configured to process four portions of image data captured by the respective image sensors of the four image sensors, representing four parts of a scene, through a first number of layers in a convolutional neural network using the respective circuits of the four circuits 710, 720, 730, and 740, to form four intermediate outputs, wherein the four parts of the scene partially overlap. The first processing function 761 is further configured to combine two of the four intermediate outputs into a first intermediate combined output, and combine the remaining two of the four intermediate outputs into a second intermediate combined output. The device 700 is further configured to perform a third processing function 764, which is configured to process the image data through a third number of layers in a convolutional neural network using two of the four circuits 710, 720, 730, and 740 (relative to...). Figure 7 The apparatus 700 is further configured to perform a combination function 763, which is configured to combine the first output and the second output into a combined output, respectively. The apparatus 700 is further configured to perform a second processing function 665, which is configured to combine the first output and the second output into a combined output via one of four circuits 710, 720, 730, 740. The apparatus 700 is further configured to perform a second processing function 665, which is configured to pass through a second number of layers (relative to the first and second intermediate combined outputs) of the convolutional neural network using one of four circuits 710, 720, 730, 740. Figure 7 The combined output is processed by a second number of layers (referred to as the "final number of layers") to extract features related to objects in the scene. The apparatus 700 may include additional circuitry (not shown) besides the four circuits 710, 720, 730, and 740. A second processing function 765 can then be configured to process the combined output by means of additional circuitry (not shown) through a second number of layers of the convolutional neural network to extract features related to objects in the scene.
[0097] The four parts of image data can refer to image data of an image of a scene. The image can have been captured by a single image sensor, or it can be generated by combining images captured by more than one image sensor into a single image, as similar to the image data about... Figure 2 The image data is described in two parts as follows.
[0098] Alternatively, the four parts of the image data can be image data captured by the respective image sensors in the four image sensors, such as similar to the image data about... Figure 2The image data is described as two parts. In this alternative, the image data captured by the respective image sensors among the first, second, third, and fourth image sensors does not need to be first combined into an image and then divided into parts of the image data, but can be directly used as the respective parts of the first, second, third, and fourth parts of the image data. Therefore, it is not necessary to exchange the image data captured by the first, second, third, and fourth image sensors between the first, second, third, and fourth circuits before processing through a first number of layers. Additionally, since the two combined intermediate outputs are processed through an intermediate number of layers of the convolutional neural network by means of the respective circuits in the four circuits, the overlap between two parts of the four parts of the scene is greater than the overlap between parts of the four parts of the scene.
[0099] Each of the four parts of the scene can involve a portion of the width of the wide scene. In this case, there will be two outer parts and two central parts of the scene. To enable image data processing via the first layer of a convolutional neural network on four circuits 710, 720, 730, and 740, and then via the third layer of a convolutional neural network on two of the four circuits 710, 720, 730, and 740, as per [the relevant documentation]... Figure 4a and Figure 4b It is disclosed that the overlap of the central part of the scene is greater than the overlap of each of the peripheral parts of the scene with the corresponding part in the central part of the scene.
[0100] Four circuits 710, 720, 730, and 740 are configured to perform the functions of device 700. Each of the four circuits 710, 720, 730, and 740 may include a processor 715, 725, 735, or 745, such as a central processing unit (CPU), microcontroller, or microprocessor. The processor is configured to execute program code, such as program code configured to perform the functions of device 700.
[0101] Device 700 may further include memory 750. Memory 750 may be one or more of a buffer, flash memory, hard disk drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or another suitable device. In a typical arrangement, memory 750 may include non-volatile memory for long-term data storage and volatile memory serving as system memory for the four circuits 710, 720, 730, and 740. Memory 750 may exchange data with the four circuits 710, 720, 730, and 740 via a data bus. Accompanying control lines and address buses may also be present between memory 750 and the four circuits 710, 720, 730, and 740.
[0102] The functionality of device 700 can be implemented in the form of executable logic routines (e.g., lines of code, software programs, etc.) stored on a non-transitory computer-readable medium (e.g., memory 750) of device 600 and executed by four circuits 710, 720, 730, 740 (e.g., using processors 715, 725, 735, 745). Furthermore, the functionality of device 700 can be a standalone software application or part of a software application that performs additional tasks associated with device 700. The described functionality can be considered as a method configured to be executed by processing units (e.g., processors 715, 725, 735, 745 of the four circuits 710, 720, 730, 740). Moreover, while the described functionality can be implemented in software, it can also be implemented by dedicated hardware or firmware, or some combination of hardware, firmware, and / or software.
[0103] The functions performed by device 700 and four circuits 710, 720, 730, and 740 can be further adapted to relate to Figure 3 Method 300 described and about Figure 4a and Figure 4b The corresponding steps of method 400 are described.
[0104] Figure 8 This is a schematic block diagram of an embodiment of a system 800 for capturing image data representing a scene and processing that image data using a convolutional neural network to extract features related to objects in the scene. System 800 includes, as described above... Figure 7 The described apparatus 700. The system further includes a camera 810 for capturing image data representing a scene. The camera 810 includes: a first image sensor 821 arranged to capture image data representing a first portion of the scene; a second image sensor 822 arranged to capture image data representing a second portion of the scene; a third image sensor 823 arranged to capture image data representing a third portion of the scene; and a fourth image sensor 824 arranged to capture image data representing a fourth portion of the scene. The four image sensors 821, 822, 823, and 824 are arranged such that the overlap between the second and third portions of the scene is greater than the overlap between the first and second portions and greater than the overlap between the third and fourth portions. For example, the four image sensors 821, 822, 823, and 824 can be arranged to capture a wide scene by arranging them to each capture a portion of the width of a wide scene, and then stitching together the image data captured by the four image sensors 821, 822, 823, and 824 to form a panoramic image. This arrangement is in Figure 9The image is disclosed in the document. First image sensors 821 and 827 capture image data representing two peripheral portions of the scene, while second image sensors 823 and 825 capture image data representing two central portions of the scene. In this arrangement, the image sensors are positioned such that portions of the scene captured by the image sensors partially overlap, enabling a more seamless combination of image data from each sensor into a panoramic image. To enable image data processing via a first layer of a convolutional neural network on four circuits and then via a third layer of a convolutional neural network on two circuits, as per [the document's description]... Figure 4a and Figure 4b Disclosed, the four image sensors 821, 822, 823, and 824 are arranged such that the overlap of the central portion of the scene is greater than the overlap of each of the peripheral portions of the scene with the corresponding portion in the central portion of the scene. In other words, the partial overlap 960 between the second portion of the scene captured by the second image sensor 823 and the third portion of the scene captured by the third image sensor 825 is greater than the partial overlap 950 between the first portion of the scene captured by the first image sensor 821 and the second portion of the scene captured by the second image sensor 823, and the partial overlap 970 between the third portion of the scene captured by the third image sensor 825 and the fourth portion of the scene captured by the fourth image sensor 827. The image data captured by the corresponding image sensors 821, 823, 825, and 827 do not need to be first combined into one image and then divided into portions of image data, but can be directly used as the corresponding portions of the first, second, third, and fourth portions of image data. Therefore, it is not necessary to exchange image data captured by the first image sensor, the second image sensor, the third image sensor, and the fourth image sensor between the first circuit, the second circuit, the third circuit, and the fourth circuit before processing through the first number of layers.
[0105] Those skilled in the art will recognize that the present invention is not limited to the embodiments described above. Rather, many modifications and variations are possible within the scope of the appended claims. These modifications and variations will be understood and implemented by those skilled in the art in practicing the claimed invention through a study of the drawings, this disclosure, and the appended claims.
Claims
1. A method for using a convolutional neural network to process image data representing a scene to extract features related to objects in the scene, the method comprising: By means of corresponding circuits in two or more circuits, the image data representing two or more parts of the scene is processed through a first number of layers of the convolutional neural network to form two or more outputs, wherein the two or more parts of the scene partially overlap. By stitching together the two or more outputs and cutting the image data associated with each overlapping portion, the two or more outputs are combined to form a combined output; and The combined output is processed by a second number of layers of the convolutional neural network using one of the two or more circuits to extract features related to objects in the scene.
2. The method according to claim 1, wherein, The actions of processing the two or more portions of the image data include: With the aid of a first circuit, a first portion representing a first part of the scene is processed through the first number of layers of the convolutional neural network to form a first output; and With the aid of a second circuit, the image data representing a second portion of the scene is processed through the first number of layers of the convolutional neural network to form a second output, wherein the first and second portions of the scene partially overlap. The actions that combine the two or more outputs include: By stitching the first output and the second output together and cutting the image data related to the overlapping portion, the first output and the second output are combined to form the combined output. The actions for processing the combined output include: The combined output is processed by a second number of layers of the convolutional neural network using one of the first and second circuits to extract features related to objects in the scene.
3. The method according to claim 2, wherein, The first portion of the image data is image data captured by a first image sensor, and the second portion of the image data is image data captured by a second image sensor.
4. The method of claim 1, wherein, The actions of processing the two or more portions of the image data include: With the aid of a first circuit, the first portion of the image data representing a first part of the scene is processed through the first number of layers of the convolutional neural network to form a first output; With the aid of a second circuit, the image data representing a second portion of the scene is processed through the first number of layers of the convolutional neural network to form a second output, wherein the first and second portions of the scene partially overlap. With the aid of a third circuit, the third portion of the image data representing a third part of the scene is processed through the first number of layers of the convolutional neural network to form a third output, wherein the second and third portions of the scene partially overlap; and With the aid of a fourth circuit, the fourth portion of the image data representing a fourth part of the scene is processed through the first number of layers of the convolutional neural network to form a fourth output, wherein the third and fourth portions of the scene partially overlap. The actions that combine the two or more outputs include: By stitching together the first, second, third, and fourth outputs of the processing from the first and second portions of the image data and cutting out image data related to each overlapping portion, the first, second, third, and fourth outputs are combined to form the combined output. The actions for processing the combined output include: The combined output is processed by the second number of layers of the convolutional neural network using one of the first, second, third, and fourth circuits to extract features related to objects in the scene.
5. The method of claim 1, wherein, The actions of processing the two or more portions of the image data include: With the aid of a first circuit, the first portion of the image data representing a first part of the scene is processed through the first number of layers of the convolutional neural network to form a first intermediate output; With the aid of a second circuit, the image data representing a second part of the scene is processed through the first number of layers of the convolutional neural network to form a second intermediate output, wherein the first part and the second part of the scene partially overlap. By stitching together the first intermediate output and the second intermediate output of the processing from the first part and the second part of the image data and cutting out the image data related to the overlapping part of the first part of the image data, the first intermediate output and the second intermediate output are combined to form a first intermediate combined output; With the aid of a third circuit, the third part of the image data representing the third part of the scene is processed through the first number of layers of the convolutional neural network to form a third intermediate output, wherein the second part and the third part of the scene partially overlap. With the aid of a fourth circuit, the fourth part of the image data representing the fourth part of the scene is processed through the first number of layers of the convolutional neural network to form a fourth intermediate output, wherein the third part and the fourth part of the scene partially overlap. By stitching together the third intermediate output and the fourth intermediate output of the processing from the third and fourth portions of the image data and cutting out the image data related to the overlapping portion of the third portion of the image data, the third intermediate output and the fourth intermediate output are combined to form a second intermediate combined output; The first intermediate combined output is processed by a third number of layers of the convolutional neural network using one of the first, second, third, and fourth circuits to form the first output; and The second intermediate combined output is processed by the third number of layers of the convolutional neural network using one of the first, second, third, and fourth circuits to form the second output. The actions that combine the two or more outputs include: The combined output is formed by stitching the first intermediate combined output and the second intermediate combined output together and cutting out image data related to the overlapping portion of the first intermediate combined output. The actions for processing the combined output include: The combined output is processed by the second number of layers of the convolutional neural network using one of the first, second, third, and fourth circuits to extract features related to objects in the scene.
6. The method of claim 5, wherein, The first part of the image data is image data captured by a first image sensor, the second part of the image data is image data captured by a second image sensor, the third part of the image data is image data captured by a third image sensor, and the fourth part of the image data is image data captured by a fourth image sensor, wherein the overlap between the second part and the third part of the scene is greater than the overlap between the first part and the second part and greater than the overlap between the third part and the fourth part.
7. The method of claim 1, wherein, The image data is image data captured by an image sensor.
8. The method according to claim 1, wherein, The image data is image data captured by at least two image sensors.
9. The method of claim 8, wherein, Each portion of the image data is captured by a single image sensor among the at least two image sensors, wherein the at least two image sensors are arranged such that image data representing corresponding portions of the scene that partially overlap are captured.
10. The method according to claim 1, wherein, The first number of layers and the size of each overlapping filter are selected such that processing the first number of layers is covered.
11. A non-transitory computer-readable storage medium storing instructions that, when executed by a processing-capable device, are used to perform the method according to any one of claims 1 to 9.
12. An apparatus for processing image data representing a scene using a convolutional neural network to extract features associated with objects in the scene, the apparatus comprising two or more circuits configured to perform: A first processing function is configured to process two or more portions of corresponding parts of two or more parts representing the scene in the image data through a first number of layers of the convolutional neural network, using corresponding circuits in the two or more circuits, to form two or more outputs, wherein, The two or more parts of the scene partially overlap; The combination function is configured to combine two or more outputs to form a combined output by stitching the two or more outputs together and cutting the image data associated with each overlapping part; and The second processing function is configured to process the combined output through a second number of layers of the convolutional neural network by means of one of the two or more circuits to extract features related to objects in the scene.
13. The apparatus of claim 12, wherein, The image data is image data captured by one image sensor or by at least two image sensors.
14. The apparatus according to claim 12, wherein, The device includes four or more circuits. The first processing function is further configured as follows: By means of corresponding circuits in the four or more circuits, the image data, captured by the corresponding image sensors in the four image sensors and representing corresponding parts of the scene, is processed through the first number of layers of the convolutional neural network to form four intermediate outputs, wherein the four parts of the scene partially overlap; and By stitching the four or more outputs together and cutting the image data associated with each overlapping part, two of the four intermediate outputs are combined into a first intermediate combined output, and the remaining two of the four intermediate outputs are combined into a second intermediate combined output. Furthermore, the device is configured to perform: The third processing function is configured to process the corresponding intermediate combined outputs of the first intermediate combined output and the second intermediate combined output through a third number of layers of the convolutional neural network by means of two of the four or more circuits, so as to form the first output and the second output, respectively.
15. A system for capturing image data representing a scene and processing the image data using a convolutional neural network to extract features related to objects in the scene, the system comprising: The apparatus according to claim 14; and A camera for capturing image data representing the scene, the camera comprising: A first image sensor is arranged to capture a first portion of the image data representing a first part of the scene; A second image sensor is arranged to capture a second portion of the image data representing a second part of the scene; A third image sensor is arranged to capture the third portion of the image data representing a third part of the scene; A fourth image sensor is arranged to capture the fourth portion of the image data representing the fourth part of the scene. The first image sensor, the second image sensor, the third image sensor, and the fourth image sensor are arranged such that the overlap between the second part and the third part of the scene is greater than the overlap between the first part and the second part, and greater than the overlap between the third part and the fourth part.