Image processor and method for processing an image

Through the collaborative computing of the processing unit array and the data processor array, the demand for high-throughput small-area image processors is solved, efficient image processing capabilities are achieved, and image processing for driver assistance systems and self-driving cars is supported.

CN115100017BActive Publication Date: 2025-10-21MOBILEYE VISION TECH LTD

Patent Information

Application Number
CN202210482943.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-02-11
Filing Date
2016-06-09
Publication Date
2025-10-21
Estimated Expiration
2036-06-09

AI Technical Summary

Technical Problem

Existing technologies have difficulty providing high-throughput, small-area image processors to support the image processing needs of driver assistance systems and autonomous vehicles.

Method used

A processing unit array is used to perform distortion calculations. By receiving weights associated with target pixels and values ​​of adjacent source pixels, the distortion results are calculated and provided in a memory module. The absolute difference and disparity are calculated using a data processor array. The image processing flow is optimized in combination with a collection unit and a memory module.

Benefits of technology

It achieves efficient image processing, improves the throughput and processing power of image processors, is suitable for small-area image processing needs, and supports image processing for driver assistance systems and self-driving cars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100017B_ABST
    Figure CN115100017B_ABST
Patent Text Reader

Abstract

The present application relates to image processors and methods for processing images. A method of computing a warp result is disclosed that can include performing, for each target pixel in a set of target pixels, a warp computation process that includes receiving, by a first set of processing units in an array of processing units, a first weight and a second weight associated with the target pixel; receiving, by a second set of processing units in the array, values of neighboring source pixels associated with the target pixel; computing, by the second set, a warp result responsive to the values of the neighboring source pixels and the pair of weights; and providing the warp result to a memory module.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of an application filed on June 9, 2016, with application number 201680045334.8 and invention name “Image processor and method for processing images”.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to U.S. Provisional Patent Serial No. 62 / 173,389 filed on June 10, 2015; U.S. Provisional Patent Serial No. 62 / 173,392 filed on June 10, 2015; U.S. Provisional Patent Serial No. 62 / 290,383 filed on February 2, 2016; U.S. Provisional Patent Serial No. 62 / 290,389 filed on February 2, 2016; U.S. Provisional Patent Serial No. 62 / 290,392 filed on February 2, 2016; U.S. Provisional Patent Serial No. 62 / 290,395, filed February 2, 2016; U.S. Provisional Patent Serial No. 62 / 290,400, filed February 2, 2016; U.S. Provisional Patent Serial No. 62 / 293,145, filed February 9, 2016; U.S. Provisional Patent Serial No. 62 / 293,147, filed February 9, 2016; and U.S. Provisional Patent 62 / 293,908, filed February 11, 2016, the entire contents of which are incorporated herein by reference. background

[0004] Over the past few years, camera-based driver assistance systems (DAS) have entered the market, with significant progress in the development of autonomous vehicles. DAS include lane departure warning (LDW), automatic high-beam control (AHC), pedestrian recognition, and forward collision warning (FCW). These driver assistance systems utilize real-time image processing of multiple patches detected from multiple image frames captured by cameras installed in the vehicle.

[0005] There is an increasing need to provide high throughput low footprint image processors for supporting DAS and / or autonomous vehicles.

[0006] Overview

[0007] A system and method are provided as shown in the claims and description.

[0008] Any combination of any subject matter in any claims may be provided.

[0009] Any combination of any methods and / or method steps disclosed in any figures and / or in the specification may be provided.

[0010] Any combination of any units, devices and / or components disclosed in any figures and / or in the specification may be provided. Non-limiting examples of such units include a collection unit, an image processor, etc.

[0011] Any combination of methods and / or method steps according to the present application, any combination of any image processors and / or image processor components, any combination of any image processor with any collection unit and / or any processing module may be provided.

[0012] According to an embodiment of the present invention, a method for calculating a warp result may be provided. The method may include performing a warp calculation process for each target pixel in a set of target pixels. The warp calculation process may include receiving, by a first group of processing units in a processing unit array, a first weight and a second weight associated with the target pixel; receiving, by a second group of processing units in the array, a value of a neighboring source pixel associated with the target pixel; calculating, by the second group, a warp result based on the values ​​of the neighboring source pixels and a pair of weights; and providing the warp result to a memory module.

[0013] The calculation of the warping result may include relaying the values ​​of some adjacent source pixels between the processing units in the second group.

[0014] The calculation of the warped result may include relaying intermediate results calculated by the second group and values ​​of some neighboring source pixels among the processing elements of the second group.

[0015] Calculation of the distortion result may include: calculating a first difference between a first pair of adjacent source pixels and a second difference between a second pair of adjacent source pixels by a first processing unit of the second group; providing the first difference to a second processing unit of the second group; and providing the second difference to a third processing unit of the second group.

[0016] The calculation of the distortion result may also include: calculating a first modified weight by a fourth processing unit of the second group in response to the first weight; providing the first modified weight from the fourth processing unit to a second processing unit of the second group; and calculating a first intermediate result based on the first difference, the first adjacent source pixel and the first modified weight by the second processing unit of the second group.

[0017] Calculation of the distortion result may also include: providing a second difference value from the third processing unit of the second group to the sixth processing unit of the second group; providing a second adjacent source pixel from the fifth processing unit of the second group to the sixth processing unit of the second group; and calculating a second intermediate result based on the second difference value, the second adjacent source pixel and the first modification weight by the sixth processing unit of the second group.

[0018] The calculation of the distorted result may also include: providing a second intermediate result from the sixth processing unit of the second group to the seventh processing unit of the second group; providing a first intermediate result from the second processing unit of the second group to the seventh processing unit of the second group; and calculating a third intermediate result based on the first intermediate result and the second intermediate result by the seventh processing unit of the second group.

[0019] The calculation of the distortion result may also include: providing a third intermediate result from the seventh processing unit of the second group to the eighth processing unit of the second group; providing a second intermediate result from the sixth processing unit of the second group to the ninth processing unit of the second group; providing a second intermediate result from the ninth processing unit of the second group to the eighth processing unit of the second group; providing a second modification weight from the third processing unit of the second group to the eighth processing unit of the second group; and calculating the distortion result based on the second intermediate result, the third intermediate result and the second modification weight by the eighth processing unit of the second group.

[0020] The method may include performing, in parallel through the array, a plurality of warp calculation processes associated with the subset of target pixels.

[0021] The method may comprise fetching, in parallel, adjacent source pixels associated with each destination pixel in a subset of pixels from a collection unit; wherein the collection unit may comprise a set associate cache and may be arranged to access a memory module which may comprise a plurality of independently accessible memory banks.

[0022] The method may include receiving first and second warping parameters for each target pixel in a subset of pixels; wherein the first and second warping parameters may include first and second weights and position information indicating a position of a neighboring source pixel associated with the target pixel.

[0023] The method may include providing position information of each target pixel in the subset of pixels to a collection unit.

[0024] The method may include converting the position information into addresses of adjacent source pixels by a collection unit.

[0025] The method may include calculating, by a third set of processing units of the array and for each destination pixel of the subset of pixels, first and second warp parameters; wherein the first and second warp parameters may include first and second weights and position information indicating a position of a neighboring source pixel associated with the destination pixel.

[0026] The method senses the first weight and the second weight from the third group to the first group.

[0027] The method may include providing position information of each target pixel of the subset of pixels to a collection unit.

[0028] The method may include converting the position information into addresses of adjacent source pixels by a collection unit.

[0029] According to one embodiment of the present invention, a method for calculating a distortion result may be provided, which may include simultaneously receiving a first weight and a second weight for each target pixel in a subgroup of pixels through a first group of processing units in a processing unit array; simultaneously providing position information indicating the position of an adjacent source pixel associated with each target pixel in the subgroup of pixels to a collection unit; simultaneously receiving adjacent source pixels associated with each target pixel in the subgroup of pixels through the array and from the collection unit; wherein different groups of the array receive adjacent source pixels associated with different target pixels in the subgroup of pixels; and simultaneously calculating distortion results related to different target pixels through different groups of the array.

[0030] The method may include receiving or calculating first and second warp parameters for each target pixel in a subset of pixels; wherein the first and second warp parameters may include first and second weights and position information indicating a position of a neighboring source pixel associated with the target pixel.

[0031] According to one embodiment of the present invention, a method for calculating a warping result may be provided, the method comprising: repeating the following steps for each subgroup of target pixels in a group of target pixels: receiving, through an array of processing units, adjacent source pixels associated with each target pixel in the subgroup of target pixels; and calculating, through the array, the warping result for the target pixels from the subgroup of target pixels; wherein the calculation may comprise calculating intermediate results and relaying at least some of the intermediate results between the processing units in the array.

[0032] Each processing unit of the array may be directly coupled to one set of processing units in the array and may be indirectly coupled to another set of processing units in the array.The terms "processing unit" and "data processor" may be used interchangeably.

[0033] According to an embodiment of the present invention, an image processor that can be configured to calculate a distortion result can be provided. The image processor can be configured to perform a distortion calculation process for each target pixel in a group of target pixels. The distortion calculation process can include receiving a first weight and a second weight associated with the target pixel through a first group of processing units in a processing unit array of the image processor; receiving a value of an adjacent source pixel associated with the target pixel through a second group of processing units in the array; calculating a distortion result based on a response to the value of the adjacent source pixel and a pair of weights through the second group; and providing the distortion result to a memory module.

[0034] According to an embodiment of the present invention, an image processor can be provided, which can be configured to calculate a distortion result. The image processor may include an array of processing units, which can be configured to simultaneously receive a first weight and a second weight for each target pixel in a subgroup of pixels through a first group of processing units in the array; simultaneously provide position information indicating the position of an adjacent source pixel associated with each target pixel in the subgroup of pixels to a collection unit of the image processor; simultaneously receive adjacent source pixels associated with each target pixel in the subgroup of pixels through the array and from the collection unit; wherein different groups in the array receive adjacent source pixels associated with different target pixels in the subgroup of pixels; and simultaneously calculate distortion results related to different target pixels through different groups in the array.

[0035] According to an embodiment of the present invention, an image processor may be provided, which may be configured to calculate a distortion result. The image processor may be configured to repeat the following steps for each subgroup of target pixels in a group of target pixels: receiving adjacent source pixels associated with each target pixel in the subgroup of target pixels through an array of processing units of the image processor; and calculating the distortion result for the target pixels from the subgroup of target pixels through the array; wherein the calculation may include calculating intermediate results and relaying at least some of the intermediate results between the processing units in the array.

[0036] According to an embodiment of the present invention, a method for calculating disparity may be provided, which may include calculating a set of sums of absolute differences (SADs) by a first group of data processors of a data processor array; wherein the set of SADs may be associated with a subgroup of source pixels and target pixels; wherein each SAD may be calculated based on a previously calculated SAD and based on the absolute difference between another source pixel currently calculated and a target pixel belonging to a subgroup of target pixels; and determining the best matching target pixel in the subgroup of target pixels in response to the value of the set of SADs by a second group of data processors of the array.

[0037] A given SAD in the set of SADs reflects the absolute difference between a given rectangular source pixel array and a given rectangular target pixel array; wherein the previously calculated SAD may include (a) a first previously calculated SAD reflecting the absolute difference between (i) a rectangular source pixel array that differs from the given rectangular source pixel array by a first source pixel column and a second source pixel column and (ii) a rectangular target pixel array that differs from the given rectangular target pixel array by a first target pixel column and a second target pixel column; and (b) a second previously calculated SAD reflecting the absolute difference between the first source column and the first source column.

[0038] For a given SAD - the other source pixel may be the lowest source pixel of the second source pixel column and the target pixel belonging to the subset of target pixels may be the lowest target pixel of the second target pixel column.

[0039] The method may include: calculating an intermediate result by subtracting (a) a second previously calculated SAD and (b) the absolute difference between (i) a target pixel that may be located at the top of a second target pixel column and (ii) a source pixel that may be located at the top of a second source pixel column from a first previously calculated SAD; and calculating a given SAD by adding the absolute difference between the lowest target pixel in the second target pixel column and the lowest source pixel in the second source pixel column to the intermediate result.

[0040] The method may include storing, for a given SAD, in a data processor array a first previously calculated SAD, a second previously calculated SAD, a destination pixel that may be positioned at the top of a second destination pixel column, and a source pixel that may be positioned at the top of a second source pixel column.

[0041] The calculation of a given SAD may be performed before extracting the lowest destination pixel of the second destination pixel column and the lowest source pixel of the second source pixel column.

[0042] The subset of target pixels may include target pixels that may be sequentially stored in a memory module; wherein the calculation of a set of SADs may be performed before retrieving the subset of target pixels from the memory module.

[0043] The fetching of the subset of target pixels from the memory module may be performed by a collection unit that may include a content addressable memory buffer.

[0044] The subgroup of target pixels belongs to a group of target pixels that may include multiple subgroups of target pixels; wherein the method may include repeating the following steps for each subgroup of target pixels: calculating a group of SADs for each subgroup of target pixels by a first group of processing units; and finding the best matching target pixel in the group of target pixels in response to the value of the group of SADs for each subgroup of target pixels by a second group of data processors of the array.

[0045] The method may include: calculating, by a first group of data processors in a data processor array, a plurality of sets of SADs that may be associated with a plurality of source pixels and a plurality of target pixel subsets; wherein each SAD in the plurality of sets of SADs may be calculated based on an absolute difference between a previously calculated SAD and a currently calculated SAD; and finding, by a second group of data processors in the array and for a source pixel, a best matching target pixel in response to the value of the SAD that may be associated with the source pixel.

[0046] The multiple sets of SADs may include subsets of SADs, and each subset of SADs may be associated with a plurality of source pixels and a plurality of subsets of destination pixels in the plurality of subsets of destination pixels.

[0047] The plurality of source pixels may belong to columns of a rectangular pixel array and may be adjacent to each other.

[0048] The calculation of multiple groups of SADs may include calculating the SADs in different subgroups of the SADs in parallel.

[0049] The method may include calculating the SADs belonging to the same subgroup of SADs in a sequential manner.

[0050] The plurality of source pixels may be a pair of source pixels.

[0051] The plurality of source pixels may be four source pixels.

[0052] Different subsets of SADs may be calculated by different first subsets of data processors in the array of data processors.

[0053] The method may include calculating SADs belonging to the same subgroup of SADs in a sequential manner; and sequentially fetching target pixels associated with different SADs in the same subgroup of SADs to a data processor array.

[0054] According to an embodiment of the present invention, an image processor may be provided, which may include an array of data processors and may be configured to calculate disparity by calculating a set of sums of absolute differences (SADs) by a first group of data processors of the data processor array; wherein the set of SADs may be associated with a subgroup of source pixels and target pixels; wherein each SAD may be calculated based on a previously calculated SAD and based on a currently calculated absolute difference between another source pixel and a target pixel belonging to a subgroup of target pixels; and determining the best matching target pixel in the subgroup of target pixels in response to the value of the set of SADs by a second group of data processors of the array.

[0055] According to an embodiment of the present invention, a collection unit may be provided, which may include an input interface that may be arranged to receive multiple requests for retrieving multiple requested data units; a cache memory that may include multiple entries and may be configured to store multiple tags and multiple cached data units; wherein each tag may be associated with a cached data unit and may indicate a group of memory cells in a memory module that are different from the cache memory and store the cached data units; a comparator array that may be arranged to simultaneously compare between the multiple tags and multiple requested memory group addresses to provide a comparison result; wherein each requested memory group address may indicate a group of memory cells in the memory module that store the requested data unit among the multiple requested data units; a contention evaluation unit; a controller that may be arranged to: (a) classify the multiple requested data units into cached data units and uncached data units that can be stored in the cache memory based on the comparison result; and (b) send information about the cached and uncached data units to the contention evaluation unit; wherein the contention evaluation unit may be arranged to check for the occurrence of at least one contention; and an output interface that may be arranged to request any uncached data units from the memory module in a contention-free manner.

[0056] The comparator array may be arranged to perform comparisons between a plurality of tags and a plurality of requested memory bank addresses simultaneously during a single collection unit clock cycle; and wherein the contention evaluation unit may be arranged to check for an occurrence of the at least one contention during a single collection unit clock cycle.

[0057] The contention evaluation unit may be arranged to recheck the occurrence of at least one contention in response to a new tag of the cache memory.

[0058] The collection unit may be arranged to operate in a pipelined manner; wherein the duration of each stage of the pipeline may be one collection unit clock cycle.

[0059] Each group of memory cells in a row of a memory bank is from a plurality of independently accessible memory banks; wherein the contention evaluation unit may be arranged to determine that potential contention occurs when two uncached data cells belong to different rows in the same memory bank.

[0060] The cache memory may be a fully associative memory cache.

[0061] The collection unit may comprise an address translator which may be arranged to translate the location information contained in the plurality of requests into a plurality of requested memory bank addresses.

[0062] The plurality of requested data units may belong to an array of data units; wherein the location information comprises coordinates of the plurality of requested data units within the array of data units.

[0063] The contention evaluation unit may include a plurality of groups of nodes; wherein each group of nodes may be arranged to evaluate contention between a plurality of requested memory bank addresses and tags of the plurality of tags.

[0064] According to an embodiment of the present invention, a method for responding to multiple requests to retrieve multiple requested data units can be provided, the method may include: receiving multiple requests for retrieving multiple requested data units through an input interface of a collection unit; storing multiple tags and multiple cached data units through a cache memory that may include multiple entries; wherein each tag can be associated with a cached data unit and can indicate a group of memory cells in a memory module that are different from the cache memory and store the cached data units; simultaneously comparing between the multiple tags and multiple requested memory group addresses through a comparator array to provide a comparison result; wherein each requested memory group address can indicate a group of memory cells in the memory module that store requested data units among the multiple requested data units; classifying the multiple requested data units into cached data units and uncached data units that can be stored in the cache memory based on the comparison result through a controller; and sending information about the cached and uncached data units to a contention evaluation unit; checking the occurrence of at least one contention through the contention evaluation unit; and requesting any uncached data units from the memory module in a contention-free manner through an output interface.

[0065] According to an embodiment of the present invention, a processing module may be provided, which may include a data processor array; wherein each data processor unit in a plurality of data processors in the data processor array may be directly coupled to some data processors in the data processor array, may be indirectly coupled to some other data processors in the data processor array, and may include a relay channel for relaying data between relay ports of the data processors.

[0066] The retransmission channel of each of the plurality of data processors may exhibit substantially zero latency.

[0067] Each of the plurality of data processors may include a core; wherein the core may include an arithmetic logic unit and memory resources; wherein the cores of the plurality of data processors may be coupled to each other via a configurable network.

[0068] Each data processor of the plurality of data processors may include a plurality of data flow components of a configurable network.

[0069] Each data processor of the plurality of data processors may include a first non-rebroadcast input port directly coupleable to a first set of neighbors.

[0070] The first group of neighbors may be formed by data processors that may be located within a distance of less than four data processors from the data processor.

[0071] A first non-rebroadcast input port of a data processor may be directly coupled to the rebroadcast ports of data processors of a first set of neighbors.

[0072] The data processor may further include a second non-rebroadcast input port that may be directly coupled to the non-rebroadcast ports of the data processors of the first set of neighbors.

[0073] A first non-rebroadcast input port of a data processor may be directly coupled to non-rebroadcast ports of data processors of a first set of neighbors.

[0074] The first set of neighbors may consist of eight data processors.

[0075] The first repeating port of each data processor of the plurality of data processors may be directly coupled to a second set of neighbors.

[0076] For each data processor in the plurality of data processors, the second set of neighbors is different from the first set of neighbors.

[0077] For each data processor of the plurality of data processors, the second set of neighbors may include data processing units that are further away from the data processor than any data processor belonging to the first set of neighbors.

[0078] In addition to the plurality of data processors, the processor array may also include at least one further data processor.

[0079] The data processors in a data processor array may be arranged in rows and columns.

[0080] Some data processors in each row may be coupled to each other in a cyclic manner.

[0081] The data processors in each row can be controlled by a shared microcontroller.

[0082] Each of the plurality of data processors may include a configuration instruction register; wherein the instruction register may be arranged to receive a configuration instruction during a configuration process and store the configuration instruction in the configuration instruction register; wherein the data processors of a given row may be controlled by a given shared microcontroller; wherein each data processor in a given row may be arranged to receive selection information for selecting a selected configuration instruction from the given shared microcontroller, and configure the data processor to operate according to the selected configuration instruction under specific conditions.

[0083] The particular condition may be satisfied when the data processor may be arranged to respond to the selection information; wherein the particular condition may not be satisfied when the data processor may be arranged to ignore the selection information.

[0084] Each of the plurality of data processors may include a controller, an arithmetic logic unit, a register file, and a configuration instruction register; wherein the instruction register may be arranged to receive a configuration instruction during a configuration process and store the configuration instruction in the configuration instruction register; wherein the controller may be arranged to receive selection information for selecting a selected configuration instruction and configure the data processor to operate according to the selected configuration instruction.

[0085] Each data processor of the plurality of data processors may include up to three configuration instruction registers.

[0086] According to an embodiment of the present invention, a method for operating a processing module that may include a data processor array may be provided; wherein the operation may include processing data by a data processor in the array; wherein each data processor unit in a plurality of data processors in the data processor array may be directly coupled to some data processors in the data processor array, may be indirectly coupled to some other data processors in the data processor array, and wherein data is relayed between relay ports of the processing data processors using one or more relay channels of the one or more data processors.

[0087] According to an embodiment of the present invention, an image processor may be provided, which may include a data processor array, a first microcontroller, a buffer unit, and a second microcontroller; wherein the data processors in the array may be arranged to receive data processor configuration instructions during a data processor configuration process; wherein the buffer unit may be arranged to receive buffer unit configuration instructions during a buffer unit configuration process; wherein the first microcontroller may be arranged to control the operation of the data processor by providing data processor selection information to the data processor; wherein the data processor may be arranged to select the selected data processor configuration instruction in response to the data processor selection information and perform one or more data processing operations according to the selected data processor configuration instruction; wherein the second microcontroller may be arranged to control the operation of the buffer unit by providing buffer unit selection information to the buffer unit; wherein the buffer unit may be arranged to select the selected buffer unit configuration instruction in response to at least a portion of the buffer unit selection information and perform one or more buffer unit operations according to the selected buffer unit configuration instruction; and wherein the size of the data processor selection information may be a portion of the size of the data processor configuration instruction.

[0088] The data processors in the array may be arranged in groups of data processors, wherein different groups of data processors may be controlled by different first microprocessors.

[0089] A data processor group may be a row of data processors.

[0090] The data processors in the same group of data processors receive the same data processor selection information in parallel.

[0091] The buffer unit may include a plurality of memory resource groups; wherein different memory resource groups may be coupled to different data processor groups.

[0092] The image processor may comprise a second microcontroller; wherein different second microcontrollers may be arranged to control different groups of memory resources.

[0093] Different memory resource groups may be different shift register groups.

[0094] Different shift register groups may be coupled to multiple buffer groups that may be arranged to receive data from the memory module.

[0095] The plurality of buffer groups may not be controlled by the second microcontroller.

[0096] The buffer unit selection information selects connectivity between the plurality of memory resource groups and the plurality of data processor groups.

[0097] Each data processor may include an arithmetic logic unit and a data flow component; wherein the data processor configuration instructions define operation codes for the arithmetic logic unit and define data flow to the arithmetic logic unit via the data flow component.

[0098] The image processor may further comprise a memory module which may comprise a plurality of memory banks; wherein the buffer unit may be arranged to extract data from the memory module and send the data to the data processor array.

[0099] The first microcontroller shares the program memory.

[0100] Each of the first microcontrollers may include a control register storing a first instruction address, a plurality of header instructions, and a plurality of loop instructions.

[0101] The image processor may include a memory module that may be coupled to the buffer unit; wherein the memory module may include a memory buffer, a load store unit, and a plurality of memory banks; wherein the memory buffer may be controlled by the third microcontroller.

[0102] The memory buffer may be arranged to receive a memory buffer configuration instruction during a memory buffer configuration procedure; wherein the third microcontroller may be arranged to control the operation of the memory buffer by providing memory buffer selection information to the memory buffer.

[0103] The image processor may include a memory buffer that may be controlled by a third microprocessor.

[0104] The third microcontroller, the first microcontroller, and the second microcontroller may have the same structure.

[0105] According to an embodiment of the present invention, an image processor may be provided, which may include multiple configurable circuits and multiple microcontrollers; wherein the multiple configurable circuits may include memory circuits and multiple data processors; wherein each configurable circuit may be arranged to store up to a limited number of configuration instructions; wherein the multiple microcontrollers may be arranged to control the multiple configurable circuits by repeatedly providing selection information to the multiple configurable circuits, the selection information being used to select a selected configuration instruction from the limited number of configuration instructions by each configurable circuit.

[0106] The plurality of configurable circuits may include a memory module which may include a plurality of memory banks; and a buffer unit for exchanging data between the memory module and the data processor.

[0107] The size of the information is selected to be no more than two bits.

[0108] According to an embodiment of the present invention, a method for configuring an image processor that may include multiple configurable circuits and multiple microcontrollers may be provided; wherein the multiple configurable circuits may include a memory circuit and multiple data processors; wherein the method may include storing up to a limited number of configuration instructions in each configurable circuit; and controlling the multiple configurable circuits by repeatedly providing selection information to the multiple configurable circuits through multiple microcontrollers, the selection information being used to select a selected configuration instruction from the limited number of configuration instructions through each configurable circuit.

[0109] According to an embodiment of the present invention, a method for operating an image processor may be provided, which may include: providing an image processor, which may include a data processor array; a memory module that may include multiple storage bodies; a buffer unit; a collection unit; and multiple microcontrollers; controlling the data processor array through a portion of the multiple microprocessors, the memory module, and the buffer unit; retrieving data from the memory module through the buffer unit; sending the data to the data processor array through the buffer unit; receiving multiple requests for retrieving multiple requested data units from the memory module through the collection unit; and sending the multiple requested data units to the data processor array through the collection unit.

[0110] According to an embodiment of the present invention, a method for configuring an image processor may be provided, which image processor may include an array of a data processor array, a first microcontroller, a buffer unit, and a second microcontroller; wherein the method may include providing data processor configuration instructions to data processors in the array during a data processor configuration process; providing buffer unit configuration instructions to the buffer unit during a buffer unit configuration process; providing data processor selection information to the data processor by the first microcontroller to control the operation of the data processor; selecting the selected data processor configuration instruction by the data processor in response to the data processor selection information, and performing one or more data processing operations according to the selected data processor configuration instruction; providing buffer unit selection information to the buffer unit by the second microcontroller to control the operation of the buffer unit; selecting the selected buffer unit configuration instruction by the buffer unit in response to at least a portion of the buffer unit selection information, and performing one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information may be a portion of the size of the data processor configuration instruction.

[0111] According to an embodiment of the present invention, a non-transitory computer-readable medium may be provided, wherein the non-transitory computer-readable medium stores instructions for calculating a distortion result, wherein once the instructions are executed by a processing unit array, the following steps are executed: performing a distortion calculation process for each target pixel in a group of target pixels, the distortion calculation process comprising: receiving, by a first group of processing units in the processing unit array, a pair of weights comprising a first weight and a second weight associated with the target pixel; receiving, by a second group of processing units in the array, a value of an adjacent source pixel associated with the target pixel; calculating, by the second group, a distortion result based on a response to the values ​​of the adjacent source pixels and the pair of weights; and providing the distortion result to a memory module.

[0112] According to an embodiment of the present invention, a non-transitory computer-readable medium may be provided, which stores instructions for calculating a distortion result, and which, once executed by a processing unit array, results in the execution of the following steps: simultaneously receiving a first weight and a second weight for each target pixel in a pixel subgroup through a first group of processing units in the processing unit array; simultaneously providing position information indicating the position of an adjacent source pixel associated with each target pixel in the pixel subgroup to a collection unit; simultaneously receiving adjacent source pixels associated with each target pixel in the pixel subgroup through the array and from the collection unit; wherein different groups in the array receive adjacent source pixels associated with different target pixels in the pixel subgroup; and simultaneously calculating distortion results related to the different target pixels through different groups of the array.

[0113] According to an embodiment of the present invention, a non-transitory computer-readable medium may be provided, the non-transitory computer-readable medium storing instructions for calculating a warping result, the instructions, once executed by a processing unit array, resulting in the execution of the following steps: repeating the following steps for each target pixel subset in a set of target pixels: receiving, through the processing unit array, adjacent source pixels associated with each target pixel in the target pixel subset; and calculating, through the array, the warping result for the target pixels from the target pixel subset; wherein the calculating includes calculating intermediate results and relaying at least some of the intermediate results between the processing units in the array.

[0114] According to an embodiment of the present invention, a non-transitory computer-readable medium may be provided, which stores instructions for calculating disparity, which, once executed by a data processor array, result in the execution of the following steps: calculating a set of sums of absolute differences (SADs) by a first group of data processors in the data processor array; wherein the set of SADs is associated with source pixels and target pixel subgroups; wherein each SAD is calculated based on a previously calculated SAD and based on a currently calculated absolute difference between another source pixel and a target pixel belonging to the target pixel subgroup; and determining the best matching target pixel in the target pixel subgroup in response to the value of the set of SADs by a second group of data processors in the array.

[0115] According to an embodiment of the present invention, a non-transitory computer-readable medium can be provided, which stores instructions for responding to multiple requests to retrieve multiple requested data units, which instructions, once executed by a collection unit, result in the following steps: receiving multiple requests for retrieving multiple requested data units through an input interface of the collection unit; storing multiple tags and multiple cached data units through a cache memory including multiple entries; wherein each tag is associated with a cached data unit and indicates a group of memory cells in a memory module that are different from the cache memory and store the cached data units; performing a comparison between the multiple tags and multiple requested memory bank addresses simultaneously through a comparator array to provide a comparison result; wherein each requested memory bank address indicates a group of memory cells in the memory module that store requested data units among the multiple requested data units; classifying the multiple requested data units into cached data units and uncached data units stored in the cache memory based on the comparison result through a controller; and sending information about the cached and uncached data units to a contention evaluation unit; checking for the occurrence of at least one contention through the contention evaluation unit; and requesting any uncached data units from the memory module in a contention-free manner through an output interface.

[0116] According to an embodiment of the present invention, a non-transitory computer-readable medium storing instructions for operating a processing module may be provided, wherein once the instructions are executed by the processing module, the following steps are performed: processing data by a data processor in a data processor array in the processing module; wherein each data processor unit in a plurality of data processors in the data processor array is directly coupled to some data processors in the data processor array, is indirectly coupled to some other data processors in the data processor array, and data is relayed between relay ports of the data processors using one or more relay channels of the one or more data processors.

[0117] According to an embodiment of the present invention, a non-transitory computer-readable medium may be provided, storing instructions for configuring an image processor comprising a plurality of configurable circuits and a plurality of microcontrollers; wherein the plurality of configurable circuits include a memory circuit and a plurality of data processors, and wherein once the instructions are executed by the image processor, the following steps are performed: storing up to a limited number of configuration instructions in each configurable circuit; and controlling the plurality of configurable circuits by repeatedly providing selection information to the plurality of configurable circuits through the plurality of microcontrollers, the selection information being used to select a selected configuration instruction from the limited number of configuration instructions through each configurable circuit.

[0118] A non-transitory computer-readable medium storing instructions for operating an image processor, the image processor comprising a data processor array; a memory module comprising a plurality of memory banks; a buffer unit; a collection unit; and a plurality of microcontrollers; wherein execution by the image processor results in the following steps: sending data to the data processor array through the buffer unit; receiving, through the collection unit, a plurality of requests for retrieving a plurality of requested data units from the memory module; and sending, through the collection unit, the plurality of requested data units to the data processor array.

[0119] A non-transitory computer-readable medium storing instructions for configuring an image processor, the image processor comprising a data processor array, a first microcontroller, a buffer unit, and a second microcontroller; wherein the plurality of configurable circuits comprise memory circuits and a plurality of data processors, wherein the instructions, once executed by the image processor, result in the following steps being performed: providing data processor configuration instructions to the data processors in the array during a data processor configuration process; providing buffer unit configuration instructions to the buffer unit during a buffer unit configuration process; providing data processor selection information to the data processor via the first microcontroller to control the operation of the data processor; selecting the selected data processor configuration instruction by the data processor in response to the data processor selection information and performing one or more data processing operations according to the selected data processor configuration instruction; providing buffer unit selection information to the buffer unit via the second microcontroller to control the operation of the buffer unit; selecting the selected buffer unit configuration instruction by the buffer unit in response to at least a portion of the buffer unit selection information and performing one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information is a portion of the size of the data processor configuration instruction.

[0120] According to an embodiment of the present invention, a method may be provided, including:

[0121] performing a warp calculation process for each target pixel in a set of target pixels, the warp calculation process comprising: receiving, by a first group of processing units in an array of processing units, a pair of weights including a first weight and a second weight associated with the target pixel; receiving, by a second group of processing units in the array, a value of a neighboring source pixel associated with the target pixel; calculating, by the second group, a warp result based on the values ​​of the neighboring source pixels and the pair of weights; and providing the warp result to a memory module;

[0122] A set of sums of absolute differences (SADs) is calculated by a first set of data processors in the array of data processors; wherein the set of SADs is associated with a subset of source pixels and a subset of destination pixels; wherein each SAD is calculated based on a previously calculated SAD and based on a currently calculated absolute difference between another source pixel and a destination pixel belonging to the subset of destination pixels; and a best matching destination pixel in the subset of destination pixels is determined by a second set of data processors in the array in response to values ​​of the set of SADs.

[0123] According to an embodiment of the present invention, a method can be provided, comprising: (a) performing a warping calculation process for each target pixel in a set of target pixels, the warping calculation process comprising: receiving a pair of weights including a first weight and a second weight associated with the target pixel through a first group of processing units in a processing unit array; receiving a value of an adjacent source pixel associated with the target pixel through a second group of processing units in the array; calculating a warping result based on the value of the adjacent source pixel and the pair of weights in response to the second group; and providing the warping result to a memory module; and (b) receiving a plurality of requests for retrieving a plurality of requested data units through an input interface of a collection unit; storing a plurality of tags and a plurality of cached data units through a cache memory comprising a plurality of entries; wherein each tag is associated with the cached data unit and indicates the storage a group of memory cells in a memory module that are different from the cache memory and store the cached data cells; performing, by a comparator array, a comparison between the plurality of tags and a plurality of requested memory bank addresses simultaneously to provide a comparison result; wherein each requested memory bank address indicates a group of memory cells in the memory module that store a requested data cell among the plurality of requested data cells; classifying, by a controller, the plurality of requested data cells into cached data cells and uncached data cells stored in the cache memory based on the comparison result; and sending information about the cached data cells and the uncached data cells to a contention evaluation unit; detecting, by the contention evaluation unit, occurrence of at least one contention; and requesting any uncached data cells from the memory module in a contention-free manner through an output interface.

[0124] According to an embodiment of the present invention, a method can be provided, comprising: (a) performing a distortion calculation process for each target pixel in a set of target pixels, the distortion calculation process comprising: receiving a pair of weights including a first weight and a second weight associated with the target pixel through a first group of processing units in a processing unit array; receiving a value of an adjacent source pixel associated with the target pixel through a second group of processing units in the array; calculating a distortion result based on the value of the adjacent source pixel and the pair of weights in response to the second group; and providing the distortion result to a memory module; and (b) operating a processing module including a data processor array; wherein the operation comprises processing data through the data processors in the array; wherein each data processor unit of a plurality of data processors in the data processor array is directly coupled to some data processors in the data processor array, is indirectly coupled to some other data processors in the data processor array, and uses one or more relay channels in one or more data processors to relay data between relay ports of the data processors.

[0125] According to an embodiment of the present invention, a method can be provided, which includes: (a) performing a distortion calculation process for each target pixel in a group of target pixels, the distortion calculation process including: receiving a pair of weights through a first group of processing units in a processing unit array, the pair of weights including a first weight and a second weight associated with the target pixel; receiving the value of an adjacent source pixel associated with the target pixel through a second group of processing units in the array; calculating a distortion result based on the value of the adjacent source pixel and the pair of weights in response to the second group; and providing the distortion result to a memory module; and (b) configuring an image processor including a plurality of configurable circuits and a plurality of microcontrollers; wherein the plurality of configurable circuits include a memory circuit and a plurality of data processors; wherein configuring the image processor includes storing up to a limited amount of configuration instructions in each configurable circuit; controlling the plurality of configurable circuits through the plurality of microcontrollers by repeatedly providing selection information to the plurality of configurable circuits, the selection information being used to select a selected configuration instruction from the limited amount of configuration instructions through each configurable circuit.

[0126] According to an embodiment of the present invention, a method can be provided, comprising: (a) performing a distortion calculation process for each target pixel in a set of target pixels, the distortion calculation process comprising: receiving a pair of weights by a first group of processing units in a processing unit array, the pair of weights comprising a first weight and a second weight associated with the target pixel; receiving a value of an adjacent source pixel associated with the target pixel by a second group of processing units in the array; calculating a distortion result by the second group based on a response to the value of the adjacent source pixel and the pair of weights; and providing the distortion result to a memory module; and (b) operating an image processor, wherein operating the image processor comprises providing an image processor comprising a data processor array; a memory module comprising a plurality of memory banks; a buffer unit; a collection unit; and a plurality of microcontrollers; controlling the data processor array by the plurality of microcontrollers, a portion of the memory module and the buffer unit; retrieving data from the memory module by the buffer unit; sending the data to the data processor array by the buffer unit; receiving a plurality of requests to retrieve a plurality of requested data units from the memory module by the collection unit; and sending the plurality of requested data units to the data processor array by the collection unit.

[0127] According to an embodiment of the present invention, a method may be provided, comprising: (a) performing a distortion calculation process for each target pixel in a set of target pixels, the distortion calculation process comprising: receiving a pair of weights by a first group of processing units in a processing unit array, the pair of weights comprising a first weight and a second weight associated with the target pixel; receiving a value of an adjacent source pixel associated with the target pixel by a second group of processing units in the array; calculating a distortion result by the second group based on a response to the value of the adjacent source pixel and the pair of weights; and providing the distortion result to a memory module; and (b) configuring an image processor comprising an array of data processors, a first microcontroller, a buffer unit, and a second microcontroller; wherein configuring the image processor comprises: providing a data processor configuration to the data processors in the array during the data processor configuration process. configuration instruction; providing a buffer unit configuration instruction to the buffer unit during a buffer unit configuration process; controlling the operation of the data processor by the first microcontroller by providing data processor selection information to the data processor; selecting the selected data processor configuration instruction by the data processor in response to the data processor selection information, and performing one or more data processing operations according to the selected data processor configuration instruction; controlling the operation of the buffer unit by the second microcontroller by providing buffer unit selection information to the buffer unit; selecting the selected buffer unit configuration instruction by the buffer unit in response to at least a portion of the buffer unit selection information, and performing one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information is a portion of the size of the data processor configuration instruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0128] The subject matter which is regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. Figure 1 The invention, both as to its method of operation and organization, together with objects, features, and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings, in which:

[0129] Figure 1 A method according to an embodiment of the present invention is shown;

[0130] Figure 2 An image processor according to an embodiment of the present invention is shown;

[0131] Figure 3 An image processor according to an embodiment of the present invention is shown;

[0132] Figure 4 A portion of an image processor according to an embodiment of the present invention is shown;

[0133] Figure 5shows a clock tree according to an embodiment of the present invention;

[0134] Figure 6 illustratively showing a memory module according to an embodiment of the present invention;

[0135] Figure 7 shows a mapping between an LSU of a memory module and a memory bank of the memory module according to an embodiment of the present invention;

[0136] Figure 8 shows a storage buffer according to an embodiment of the present invention;

[0137] Figure 9 and Figure 10 shows the instruction Row, Sel according to an embodiment of the present invention;

[0138] Figure 11 illustratively showing a buffer unit according to an embodiment of the present invention;

[0139] Figure 12 shows a collecting unit according to an embodiment of the present invention;

[0140] Figure 13 is a timing diagram illustrating a process including address translation, cache hit / miss, contention, and information output;

[0141] Figure 14 Showing a contention evaluation unit according to an embodiment of the present invention;

[0142] Figure 15 and Figure 16 shows a data processing unit according to an embodiment of the present invention;

[0143] Figure 17 A distortion calculation method according to an embodiment of the present invention is shown;

[0144] Figure 18 and Figure 19 shows an array of data processors performing warp calculations according to an embodiment of the present invention;

[0145] Figure 20 shows distortion parameters output from various data processors according to an embodiment of the present invention;

[0146] Figure 21 shows a data processor group performing warp calculations according to an embodiment of the present invention;

[0147] Figure 22 A distortion calculation method according to an embodiment of the present invention is shown;

[0148] Figure 23 shows a group of processing units according to an embodiment of the present invention;

[0149] Figure 24 There is shown a first subset of source pixels, a first subset of destination pixels, a second subset of source pixels, and a second subset of destination pixels;

[0150] Figure 25 A subgroup SG(B) of source pixels is shown with a central pixel SB;

[0151] Figure 26 A corresponding subgroup TG(B) of target pixels (not shown) is shown with a central pixel TB;

[0152] Figure 27 A method according to an embodiment of the present invention is shown;

[0153] Figure 28 shows eight source pixels and thirty-two destination pixels processed by DPA according to an embodiment of the present invention;

[0154] Figure 29 shows an array of source pixels according to an embodiment of the present invention;

[0155] Figure 30 shows an array of target pixels according to an embodiment of the present invention;

[0156] Figure 31 shows a plurality of groups of data processors DPU according to an embodiment of the present invention;

[0157] Figure 32 Eight groups of data processors DPUs are shown according to an embodiment of the present invention - each group comprising four DPUs; and

[0158] Figure 33 A distortion calculation method according to an embodiment of the present invention is shown.

[0159] Detailed description of the drawings

[0160] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details. In other cases, well-known methods, processes, and components are not described in detail in order to avoid obscuring the present invention.

[0161] The subject matter which is regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. Figure 1 The invention, both as to its method of operation and organization, together with objects, features, and advantages thereof, will be best understood by reference to the following detailed description when read together.

[0162] It should be understood that for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. In addition, reference numerals may be repeated in the figures to indicate corresponding or similar elements where deemed appropriate.

[0163] Because the illustrated embodiments of the present invention can be implemented to a large extent using electronic components and circuits known to those skilled in the art, the details will not be explained to a greater extent than those that must be considered as described above in order to understand and appreciate the basic concepts of the invention and so as not to obscure or distract from the teachings of the invention.

[0164] Any reference to a method in this specification should apply, mutatis mutandis, to a system capable of performing the method, and should apply, mutatis mutandis, to a non-transitory computer-readable medium storing instructions that, once executed by a computer, result in the performance of the method. For example, any method step provided herein can be performed by a system. In this sense, a system can be an image processor, a collection unit, or any component of an image processor. A non-transitory computer-readable medium storing instructions that, once executed by a computer, result in the performance of each method provided herein can be provided.

[0165] Any reference in the specification to a system and any other components shall apply, mutatis mutandis, to methods that may be performed by a memory device and to non-transitory computer-readable media storing instructions that may be executed by a memory device. For example, the methods and / or method steps performed by an image processor as described herein may be provided.

[0166] Any reference in the specification to a non-transitory computer-readable medium should apply, mutatis mutandis, to systems capable of executing instructions stored in the non-transitory computer-readable medium and should apply, mutatis mutandis, to methods executable by a computer that reads the instructions stored in the non-transitory computer-readable medium.

[0167] Any combination of any modules or units listed in any figure, any part of the description and / or any claims may be provided. In particular, any combination of any claimed features may be provided.

[0168] A pixel may be an image element obtained by a camera, or may be a processed image element.

[0169] The terms "row" and "line" are used interchangeably.

[0170] The term automobile is used as a non-limiting example of a vehicle.

[0171] To simplify the description, some of the figures and some of the following text include numerical examples (e.g., bus width, number of memory rows, number of registers, register length, data unit size, instruction size, number of components per unit or module, number of microprocessors, number of data processors per row and / or column in an array). Each numerical example is merely a non-limiting example.

[0172] Figure 1 A method 90 according to an embodiment of the present invention is shown.

[0173] System 90 may be part of a DAS, an autonomous vehicle control module, or the like.

[0174] System 90 may be installed in automobile 10. At least some components of system 90 are within the vehicle.

[0175] The system 90 may include a first camera 81, a first processor 83, a storage unit 85, a human-machine interface 86, and an image processor 100. These components may be coupled to each other via a bus or network 82 or by any other arrangement.

[0176] System 90 may include additional cameras and / or additional processors and / or additional image processors.

[0177] The first processor 83 may determine which task the image processor 100 should perform and instruct the image processor 100 to operate accordingly.

[0178] It is to be noted that the image processor 100 may be part of the first processor 83, and it may be part of any other system.

[0179] The human-machine interface 86 may include a display, a speaker, one or more light-emitting diodes, a microphone, or any other type of human-machine interface. The human-machine interface may communicate with the vehicle driver's mobile device, the vehicle's multimedia system, and the like.

[0180] Figure 2 An image processor 100 according to an embodiment of the present invention is shown.

[0181] Master port 101 and slave port 103 provide an interface between image processor 100 and any other components of system 90 .

[0182] The image processor 100 includes:

[0183] 1) Direct Memory Access (DMA) for accessing external memory resources such as the memory unit 85 .

[0184] 2) A controller, such as but not limited to the scalar unit 104 .

[0185] 3) Scalar Unit (SU) program memory 106.

[0186] 4) Scalar unit (SU) data memory 108.

[0187] 5) Memory module (MM) 200.

[0188] 6) MM control unit 290.

[0189] 7) Collection unit (GU) 300.

[0190] 8) Buffer unit (BU) 400.

[0191] 9) BU control unit 490.

[0192] 10) Data Processing Array (DPA) 500.

[0193] 11) DPA control unit 590.

[0194] 12) Configure bus 130.

[0195] 13) Multiplexers and buffers 110 , 112 , 114 , 116 , 118 , 120 .

[0196] 14) Buses 132, 133, 134, 135, 136 and 137.

[0197] 15)PMA status and configuration buffer 109.

[0198] The image processor 100 also includes a plurality of microcontrollers. Figure 4 Shown in.

[0199] DMA 102 is coupled to multiplexers 112, 114, and 120. Scalar unit 104 is coupled to buffer 118 and multiplexer 112. Buffer 118 is coupled to multiplexer 116. Buffer 110 is coupled to multiplexers 112, 116, and 114. Multiplexer 112 is coupled to SU program memory 106. Multiplexer 114 is coupled to SU data memory 108.

[0200] The memory unit 200 is coupled to the collection unit 300 via a (unidirectional) bus 132 , to the buffer unit 400 via a (unidirectional) bus 134 , and to the DPA 500 via a (unidirectional) bus 133 .

[0201] The collection unit 300 is coupled to the buffer unit 400 via a (unidirectional) bus 135 and to the DPA 500 via a (unidirectional) bus 137. The buffer unit 400 is coupled to the DPA 500 via a (unidirectional) bus 136.

[0202] The units of the image processor may be coupled to each other via other buses, via additional fewer buses, via interconnects and / or networks, via buses of other widths and directionality, and the like.

[0203] It should be noted that Figure 2 The collection unit 300, buffer unit 400, DPA 500, memory module 200, scalar unit 104, SU program memory 106, SU data memory 108 and any other multiplexers and / or buffers may be coupled to each other in other ways, via additional and / or other common buses, networks, meshes, etc.

[0204] The scalar unit 104 can control other components of the image processor 100 to perform tasks. The scalar unit 104 can receive (for example, from Figure 1 The first processor 83 of the SU may determine which tasks to perform and may extract relevant instructions from the SU program memory 106.

[0205] The scalar unit 104 may determine which programs the microcontrollers within the SB control unit 290 , the BU control unit 490 , and the DPA control unit 590 will execute.

[0206] The program executed by the microcontroller of the SB control unit 290 controls the storage buffer (not shown) of the memory module 200. The program executed by the microcontroller of the BU control unit 490 controls the buffer unit 400. The program executed by the microcontroller of the DPA control unit 590 controls the data processing unit of the DPA 500.

[0207] Any of the microcontrollers can control any module or unit by providing a short selection message (e.g., 2-3 bits, less than one byte, or any number of bits less than the number of bits of the selected configuration instruction) for selecting a configuration instruction already stored in the controlled module or unit. This allows for reducing traffic and performing fast configuration changes (since a configuration change may require selecting between different configuration registers already stored in the relevant unit or module).

[0208] It should be noted that the number of control units and their distribution among the components of the image processor may differ from Figure 2 Those shown in .

[0209] The memory module 200 is the highest level memory resource of the image processor 100. The buffer unit 400 and the collection unit 300 are lower level memory resources of the image processor 100 and may be configured to extract data from the memory module 200 and provide the data to the DPA 500. The DPA 500 may send data directly to the memory module 200.

[0210] The DPA 500 includes a plurality of data processors and is arranged to perform computational tasks such as, but not limited to, image processing algorithms. Non-limiting examples of image processing algorithms include warp algorithms, parallax algorithms, and the like.

[0211] The collection unit 300 includes a cache memory. The collection unit 300 is configured to receive a request from the DPA 500 to extract a plurality of data units (such as pixels) from the cache memory or from the memory unit and extract the requested pixels. The collection unit 300 can operate in a pipelined manner and have a limited number (e.g., three) pipeline stages with very low latency (e.g., one (or less than five or ten) clock cycles). As shown below, the collection unit can also extract data units in an append mode, while using the address generator of the memory module to extract information.

[0212] The buffer unit 400 is configured to act as a buffer for data between the memory module 200 and the DPA 500. The buffer unit 400 may be arranged to provide data to a plurality of data processors of the DPA 500 in parallel.

[0213] Configuration bus 130 is coupled to DMA 102 , memory module 200 , collection unit 300 , buffer unit 400 , and DPA 500 .

[0214] The DPA 500 demonstrates an architecture that can support both parallel and pipelined implementations. It also demonstrates flexible connectivity, enabling it to connect nearly every data processing unit (DPU) to every other DPU.

[0215] The units of the image processor 100 are controlled by a compact microprocessor that can execute zero-delay loops and implement nested loops.

[0216] Figure 3 An image processor 100 according to an embodiment of the present invention is shown.

[0217] Figure 3 Non-limiting examples of the widths of various buses and the contents of the memory module 200 , the collection unit 300 , the buffer unit 400 , and the DPA 500 are provided.

[0218] The DPU 500 may include 6 rows and 6 columns of data processing units (DPUs) 510 ( 0 , 0 ) - 510 ( 5 , 15 ).

[0219] Configuration bus 130 is 32 bytes wide.

[0220] Bus 132 is 8 x 64 bytes wide.

[0221] Bus 134 is 6 x 128 bytes wide.

[0222] Bus 135 is 2 x 128 bytes wide.

[0223] Bus 137 is 2 x 16 x 16 bytes wide.

[0224] Bus 133 is 6 x 2 x 16 x 16 bytes wide.

[0225] Bus 136 is 2 x 16 x 16 bytes wide.

[0226] The memory module 200 is shown to include an address generator, 6 load store units, 16 multi-port memory interfaces, and 16 independently accessible memory banks of 8 byte lines.

[0227] The collection unit 300 includes a cache memory comprising 18 registers, each register being 8 bytes.

[0228] The buffer unit 400 includes 6 rows and 4 columns of 16-byte registers and 6 rows and 16 columns of 2:1 multiplexers.

[0229] Figure 4 A portion of the image processor 100 according to an embodiment of the present invention is shown.

[0230] The two memory buffers of memory module 200 may be controlled by an SB control unit 290. SB control unit 290 may include an SB program memory 293 and SB microcontrollers 291 and 292. SB program memory 293 stores instructions to be executed by SB microcontrollers 291 and 292. SB microcontrollers 291 and 292 may be fed (via configuration bus 130 and / or via scalar unit 104) with information (stored in configuration registers 298) indicating which instructions (among the instructions stored in SB program memory 293) to execute.

[0231] The different register rows of the buffer unit 400 may be controlled by a BU control unit 490. The BU control unit 490 may include a BU program memory 497, configuration registers 498, and BU microcontrollers 491-496.

[0232] The BU program memory 297 stores instructions to be executed by the BU microcontrollers 491-496. The BU microcontrollers 491-496 may be fed (via the configuration bus 130 and / or via the scalar unit 104) with information (stored in the configuration register 498) indicating which instructions (of the instructions stored in the BU program memory 497) to execute.

[0233] The DPUs of different rows of DPA 500 may be controlled by a DPA control unit 590. DPA control unit 590 may include a DPA program memory 597, configuration registers 598, and DPA microcontrollers 591-596.

[0234] The DPA program memory 297 stores instructions to be executed by the DPA microcontrollers 591-596. The DPA microcontrollers 591-596 may be fed with information (via the configuration bus 130 and / or via the scalar unit 104) indicating which instructions (of the instructions stored in the DPA program memory 597) to execute.

[0235] It should be noted that the microcontrollers can be grouped in other ways, for example, there can be one microprocessor group, two, three, or more microprocessor groups.

[0236] Figure 5 A clock tree according to an embodiment of the present invention is shown.

[0237] The input clock signal 2131 is fed to the scalar unit 104. The scalar unit sends clk_mem 2132 to the banks 610-625 of the memory module and sends clk 2133 to the buffer unit 400, the collection unit 300, and the load store units (LSUs) 630-635 of the memory module 200. clk 2133 is converted to dpa_clk 2134, which is sent to the DPA 500.

[0238] Memory modules

[0239] Figure 6 A memory module 200 is shown according to an embodiment of the present invention. Figure 7 1 and 2. A mapping between an LSU of a memory module and a memory bank of the memory module according to an embodiment of the present invention is shown.

[0240] The memory module 200 includes 16 independently accessible memory banks M0-M15 610-625, 6 load-store units LSU0-LSU5 630-625, large and small address generators AG0-AG5 640-645, and two store buffers 650 and 660.

[0241] Memory banks M0-M15 610-625 are eight bytes wide (64 bits per row) and contain 1K rows to provide a total memory capacity of 96KB. Each memory bank may include (or may be coupled to) a multi-port memory interface for arbitrating between requests sent to the memory bank.

[0242] exist Figure 6In , there are four clients coupled to each memory bank (four arrows), and the multi-port memory interface must arbitrate between access requests appearing on these four inputs.

[0243] The multi-port memory interface can apply any arbitration scheme. For example, it can apply priority-based arbitration.

[0244] Each LSU can select one of 6 addresses from the address generator and is connected to 4 memory banks, and can access 16 bytes (from 2 memory banks) per access, so that 6 LSUs can access 12 of the 16 memory banks at a time.

[0245] Figure 7 The mapping between different values ​​of the control signal SysMemMap and the mapping between the LSUs 630 - 635 and the memory banks M0 - M15 610 - 625 is shown.

[0246] Figure 6 Memory module 200 is shown outputting data cells to a collection unit through eight eight-byte wide buses (part of bus 132 ) and to a buffer unit via six sixteen-byte wide buses (part of bus 134 ).

[0247] Each address generator in AGO-AG5 640-645 can implement a four-dimensional (4D) iterator by using the following variables and registers:

[0248] Baddr defines the base address in the memory bank.

[0249] 'W' direction - variable wDepth defines the distance in bytes of one step in the W direction. Variable wCount defines the maximum value of the W counter - when this value is reached, zCounter is incremented and wCounter is cleared.

[0250] 'Z' direction - zArea defines the byte distance of one step in the Z direction, and the variable zCount defines the maximum value of the Z counter. When this value is reached, the X counter increments and the Z counter is cleared.

[0251] 'X' direction - variable xStep defines the step size (can be 1, 2, 4, 8 or 16 bytes). variable xCount defines the maximum value of the X counter before the next 'Y'

[0252] 'Y' direction - the variable stride defines the distance in bytes between the start points of consecutive 'rows'. The variable yCount defines the maximum value of the Y counter.

[0253] When all counters reach their maximum value, a stop condition is generated.

[0254] The generated address is: Addr = BAddr + wCount * wDepth + xcounter * xstep + ycounter * Stride + zcounter * Area.

[0255] The variables are stored in registers that can be configured via the configuration bus 130 .

[0256] The following is an example of an address generator configuration map:

[0257]

[0258] Storage can also be written in sizes other than 2, 4, 8, and 16 bytes.

[0259] Each LSU can perform load / store operations to or from 2 of the 4 connected memory banks (16 bytes).

[0260] The memory bank accessed may depend on the selected mapping (see for example Figure 7 ) and address.

[0261] Data to be stored is prepared in one of storage buffers 650 and 660 described below.

[0262] Each LSU can select an address generated from one of the six address generators AG0-AG5.

[0263] Load Operation

[0264] Data read from the memory banks is stored in a buffer (not shown) of the load store unit and then transferred (via bus 134) to the buffer unit 400. The buffer helps avoid stalls due to contention on the memory banks.

[0265] Storage Operations

[0266] The data to be stored in the memory bank is prepared in one of the memory buffers 650 and 660. There are two memory buffers (memory buffer 0 650 and memory buffer 1 660). Each memory buffer can request between one and four 16-byte words to be written to one of the LSUs.

[0267] Each LSU can therefore get up to 8 simultaneous requests, which are granted one after another in a predetermined order: 1) Store Buffer 0 - Word 0, 2) Store Buffer 0 - Word 1, ... 4) Store Buffer 1 - Word 0, ... , 8) Store Buffer 1 - Word 3.

[0268] When the store buffer is configured to operate in store-conditional mode, the store buffer can either ignore (not send to the memory bank) or process (send to the memory bank) data.

[0269] When a memory buffer is configured to operate in a scattered mode, a portion of a data unit received thereby may be considered an address associated with the storage of the remaining data unit.

[0270] LSU operation priority. Store operations take precedence over load operations so that stores will not stall due to contention with loads. Because load operations use the buffer, contention is usually swallowed and does not cause a stall.

[0271] Storage Buffer

[0272] Memory buffers 650 and 660 are controlled by memory buffer microcontrollers 291 and 292 .

[0273] During configuration, each of the storage buffers 650 and 660 receives (and stores) three configuration instructions (sb_instr[1]-sb_instr[3]). The configuration instructions of different storage buffers (also referred to as storage buffer configuration instructions) can be different from each other or can be the same.

[0274] During configuration, each memory buffer microcontroller receives the address of the instruction to be executed by each memory buffer microcontroller. The first and last PCs indicate the first and last instructions to be read from the memory buffer program memory 293. In the following configuration example, the location of the program memory for each memory buffer microcontroller is also defined:

[0275] Memory buffer configuration

[0276]

[0277] Memory Buffer Microcontroller Configuration Map:

[0278]

[0279] The microcontroller instructions for storing buffers can be either execute instructions or do loop instructions. They have the following format:

[0280] Execute the command:

[0281]

[0282] Do loop:

[0283]

[0284] Figure 8A memory buffer 660 is shown in accordance with an embodiment of the present invention.

[0285] The storage buffer 660 has four multiplexers 661-664, four buffers Word0-Word3 671-674, and four demultiplexers 681-684.

[0286] Buffer word 0 671 is coupled between multiplexer 661 and demultiplexer 681. Buffer word 1 672 is coupled between multiplexer 662 and demultiplexer 682. Buffer word 2 673 is coupled between multiplexer 663 and demultiplexer 683. Buffer word 3 674 is coupled between multiplexer 664 and demultiplexer 684.

[0287] Each of the multiplexers 661 - 664 has four inputs for receiving different lines of the bus 133 and is controlled by a control signal Row, Sel.

[0288] Each of the demultiplexers 681-684 has six output terminals for providing data to any one of the LSU0-LSU5 and is controlled by the control signal En,LSU.

[0289] The store buffer configuration instructions control the operation of the store buffer and even generate the commands Row, Sel and En, LSU.

[0290] An example of the configuration directive format is provided below:

[0291] Store buffer instruction encoding

[0292]

[0293] The five bits called "Data Select" are actually the instructions Row, Sel, and are used in Figure 9 and Figure 10 The values ​​between 0 and 28 are mapped to different ports of the DPU of the DPA 500. Figure 9 In , D# and E# represent the outputs D and E of DPU[row,#] respectively, where 'row' is the row select of the currently selected instruction, and D'# is the output D of DPU[row+1,#].

[0294] Buffer unit

[0295] Figure 11 A buffer unit 400 according to an embodiment of the present invention is shown.

[0296] The buffer unit 400 includes a read buffer (RB) collectively labeled 402, a register file (RF) 404, a buffer unit internal network 408, multiplexer control circuits 471-476, an output multiplexer 406, BU configuration registers 401(1)-401(5) (each storing two configuration instructions), a history configuration buffer 405 and a BU read buffer configuration register 403.

[0297] The BU microcontroller can select which configuration instruction to read for each row (of the two configuration instructions stored in each BU configuration register in 401 (1)-401 (5)).

[0298] There are six rows of multiplexers, and they include multiplexers 491(0)-491(15) and 491'(0)-491'(15), multiplexers 492(0)-492(15) and 492'(0)-492'(15), multiplexers 493(0)-493(15) and 493'(0)-493'(15), multiplexers 494(0)-494(15) and 494'(0)-494'(15), and multiplexers 495(0)-495(15) and 495'(0)-495'(15).

[0299] To simplify the explanation, Figure 11 Only multiplexer control circuits 471 and 476 are shown.

[0300] A buffer unit internal network 408 couples the read buffer 402 to the register file 404 .

[0301] The first row of four read buffers 415, 416, 417, and 417 are coupled (via the buffer unit internal network 408) to the first row of four registers R3 413, R2 412, R1 411, and R0 410.

[0302] The second row of four read buffers 425, 426, 427, and 427 are coupled (via the buffer unit internal network 408) to the second row of four registers R3 423, R2 422, R1 421, and R0 420.

[0303] The third row of four read buffers 435, 436, 437, and 437 are coupled (via the buffer unit internal network 408) to the third row of four registers R3 433, R2 432, R1 431, and R0 430.

[0304] The fourth row of four read buffers 445, 446, 447, and 447 are coupled (via the buffer unit internal network 408) to the fourth row of four registers R3 443, R2 442, R1 441, and R0 440.

[0305] Different rows of the register file and corresponding rows of multiplexers are controlled by different BU microcontrollers in 491-495.

[0306] The multiplexers of different rows are also controlled by the DPU microcontrollers. Specifically, each DPU microcontroller 591-596 controls the corresponding DPU row and sends control instructions (MuxCtl) to the corresponding multiplexer row (via multiplexer control circuits 471-476). Each multiplexer control circuit stores the last (e.g., sixteen) MuxCtl instructions (instruction history), and the history configuration buffer 405 stores selection information for determining which MuxCtl instruction to send to the multiplexer row.

[0307] The multiplexer control circuit 471 controls the multiplexers of the first row and includes a FIFO 481(1) for storing MuxCtl instructions sent from the DPU microcontroller 491, and includes controlling the multiplexer 481(2) to select which stored MuxCtl instruction is taken out from the FIFO 481(1) and sent to the multiplexers of the first row including multiplexers 491(0)-491(15) and 491'(0)-491'(15).

[0308] The multiplexer control circuit 476 controls the sixth row of multiplexers and includes a FIFO 486(1) for storing MuxCtl instructions sent from the DPU microcontroller 496 and includes controlling multiplexer 486(2) to select which stored MuxCtl instruction to take out from FIFO 482(1) and send to the sixth row of multiplexers including multiplexers 495(0)-495(15) and 495'(0)-495'(15).

[0309] The register file can be controlled by the BU microcontroller. Operations performed by the register file may include (a) shifting from the most significant to the least significant bit, where the shift jump is a power of 2 bytes (or any other value), and (b) loading one or two registers from the read buffer. The contents of the file registers can be manipulated. For example, the contents can be interleaved and / or interlaced. Some examples are provided in the instruction set provided below.

[0310] The buffer configuration map includes the addresses of the configuration buffers used to store configuration instructions for the buffer units and to store instructions for which commands are to be fetched by the buffer unit microcontrollers (BUuC0-BuuC5). The latter are referred to as the first and last PCs of a BU instruction pair (each thirty-two-bit MuxConfig instruction includes two separate buffer unit configuration instructions):

[0311] Buffer unit configuration map

[0312]

[0313]

[0314] Instructions executed by the BU microcontroller include bits [0:8] for loop control or instruction repetition, and bits [9:13] that contain a value for controlling the execution of the instruction. This is true for both register file instructions and do loop instructions.

[0315] Instruction encoding:

[0316] RFL stands for Register File Line and RBL stands for Read Buffer Line.

[0317] RRR[r / r1:r0] represents register number "r" (16 bytes) / registers from register r1 to t0 (RFL or RBL). RRR[r][b / b1:b0] represents register number "r" byte b / registers from bytes b1 to b0 (RFL or RBL).

[0318] RF instruction: (Each cycle, the BU microcontroller can send bits 9-14 to the BU)

[0319]

[0320]

[0321] Do loop:

[0322]

[0323] Read buffer loaded.

[0324] The RB loading operation from an LSU is configuration controlled and self-triggered by its status and the status of the associated LSU.The BU Read Buffer Configuration Register (also known as RBSrcCnf) 403 specifies from which LSU each RB line is loaded.

[0325] The configuration instructions stored in the BU read buffer configuration register 403 have the following format:

[0326]

[0327] "Read buffer" refers to the rows of the read buffer. The four bits of each read buffer row can have the following meanings: 0-7: no load, 8: LSU0, 9: LSU1, 10: LSU2, 11: LSU3, 12: LSU4, 13: LSU5, 14: GU0, 15: GU1 (GUI exchanges 8-byte loads as d8..d15, d0..d7 instead of d0..dl5, or the last 8 short integers of the GU in short integer mode).

[0328] Multiplexer Configuration

[0329] The configuration of the multiplexer is shown below (this example refers to the first line, and my instruction MuxCtl is represented as MuxCt10):

[0330]

[0331] MuxCtl operates Muxes (selecting registers of the register file and ports of the DPU) in the following way:

[0332]

[0333]

[0334] Selection of the history may be accomplished (by FIFOs 471-476) by reading the contents of the four-bit history configuration buffer 405 which stores the selection information for the multiplexers of each row.

[0335] Collection Unit

[0336] Figure 12 A collection unit 300 according to an embodiment of the present invention is shown.

[0337] The collection unit 300 includes an input buffer 301 , an address translator 302 , a cache memory 303 , an address tag comparator 304 , a contention evaluation unit 306 , a controller 307 , a memory interface 308 , an iterator 310 , and a configuration register 311 .

[0338] The collection unit 300 is configured to collect up to 16 byte or short pixels from eight memory banks MB0..MB7 or MB8..MB15 via a fully associative cache memory (CAM) 303 comprising 16 8-byte registers, depending on the bit gu_Ctrl

[15] of the configuration register 311. The pixel addresses or coordinates are generated or received from the array according to the mode gu_Ctrl[3:0] of the configuration register 311.

[0339] The collection unit 300 may access the banks of the memory module 200 using addresses (duplets) indicating the bank number and the row within the bank.

[0340] The collection unit may receive or generate X and Y coordinates representing the location of the requested pixel in the image instead of an address.The collection unit comprises an address converter 302 for converting the X, Y coordinates into an address.

[0341] Iterator 310 can operate in one of two modes - (a) internal iterator only and (b) address generator using the memory module.

[0342] When operating in mode (a), iterator 310 may generate sixteen addresses using the following control parameters (which are stored in configuration register 311):

[0343] 1)AddBase - 16 base addresses.

[0344] 2)AddStep-iterator address step.

[0345] 3) xCount: Counter of the maximum number of steps before stride. stride is taken from the previous step or base coordinate.

[0346] 4)AddStride: stride (or y step size).

[0347] When operating in mode (b), the iterator provides AddBase to the address generator, and the address generator uses that address to perform the iteration.

[0348] For example, during disparity calculations, the iterator pattern may be useful, where a collection unit may retrieve data units from both the source and destination images (especially pixels that are close to each other).

[0349] Another mode of operation includes receiving an address of requested data (the address may be an X, Y coordinate or a memory address such as a doublet to be converted by the address translator 302), checking whether the requested data unit is stored in the cache memory, and if not, retrieving the data unit from the memory module 200.

[0350] Another mode of operation includes receiving the address of a requested data unit and converting the address into further requested addresses, and then fetching the contents of the further requested addresses from cache memory 303 or from a memory unit. This mode of operation may be useful, for example, during warp calculations, where a collection unit may receive the address of a pixel and obtain the pixel and several other adjacent data units.

[0351] Cache memory 303 stores tags as doublets, and these tags are used to determine whether a requested data unit is in cache memory 303. These doublets are also used to detect contention - when multiple requested data units reside in different rows of the same memory bank during the same cycle.

[0352] Figure 13 is a timing diagram illustrating a process including address translation, cache hit / miss, contention, and information output.

[0353] The coordinate (X, Y) is accepted or generated in cycle 0 and converted to a doublet in cycle 1. The memory bank access is calculated in the same cycle for the access performed in cycle 2. If several coordinates address the same memory bank at different addresses (as is the case in this example), there is contention, and the corresponding pixel fetch is delayed until the next cycle, and stall is asserted in cycle 1. The coordinate causing the contention is retired in cycle 1, and the memory bank is accessed in cycle 2. The delay between the coordinate and the pixel is 5 cycles + the number of stall cycles. In the extreme case, accessing 16 pixels will result in 15 stall cycles. For warp operations, due to the proximity of the accessed pixels, 16 pixels are fetched in an average of 1.0-1.4 cycles, depending on the type of warp.

[0354] Configuration Mapping

[0355]

[0356]

[0357] Return Reference Figure 12 Input buffer 301 is coupled to address translator 302. Address tag comparator 304 receives input from cache memory 303 and from address translator 302 (if address translation is required) or from input buffer 301. Address tag comparator 304 sends output signals indicating the comparison (such as cache miss, cache hit, and if hit, where the hit occurred) to controller 307, as well as contention evaluation unit 306 and memory interface 308. Iterator 310 is coupled to input buffer 301 and memory interface 308.

[0358] An input interface, such as input buffer 301 , is arranged to receive a plurality of requests for retrieving a plurality of requested data units.

[0359] Cache memory 303 includes entries (such as sixteen entries or lines) that store a plurality of tags (each tag may be a doublet) and a plurality of cached data units.

[0360] Each tag is associated with a cached data unit and represents a group of memory cells (such as rows) in a memory module (such as a memory bank) that is different from the cache memory and stores the cached data unit.

[0361] The address tag comparator 304 includes a comparator array arranged to perform comparisons simultaneously between a plurality of tags and a plurality of requested memory bank addresses to provide comparison results.

[0362] The address tag comparator includes K×J nodes—covering each pair of tag and requested bank address. The (k, j)th node 304 (k, j) compares the kth requested address with the jth tag.

[0363] If, for example, a requested bank / row address does not match any tag, then address tag comparator 304 will signal a miss.

[0364] The controller 307 may be arranged to (a) classify the plurality of requested data units into data units stored in the cache memory 303 and uncached data units (not stored in the cache memory 303) based on the comparison result; and (b) when there is at least an uncached data unit, send information about it to the contention evaluation unit 306.

[0365] The contention evaluation unit 306 is arranged to check for the occurrence of at least one contention.

[0366] The memory interface 308 is arranged to request any uncached data units from the memory module in a contention-free manner.

[0367] When one or more uncached data units are retrieved by the collection unit, they are stored in the cache memory. In a cycle following the given cycle in which contention was detected, the address tag comparator 304 may receive the same requested data units from the given cycle from the input buffer or from the address translator. The previously uncached data units now stored in the cache memory 303 will modify the comparison result made by the address tag comparator 304, and only the uncached memory data units that were not requested in the previous iteration will be retrieved from the memory module.

[0368] The contention evaluation unit may include multiple groups of nodes. Figure 14 An example is provided in .

[0369] The number of node groups is the maximum number of memory banks (eg, eight) that the collection unit can access simultaneously.

[0370] Each group of nodes is arranged to evaluate contention associated with a single memory bank. Figure 14There are eight sets of nodes in the _ for comparison between the sixteen requested addresses and the eight memory bank rows.

[0371] The nodes of the first group of nodes (305(0,0)-306(0,15)) are connected in series. The leftmost node 306(0,0) of the first group of nodes receives the input signal (used, bank 0, row) and the signal (valid 0, address 0).

[0372] Address 0 is the first of the 16 requested addresses (of data units), and Valid 0 indicates whether the first requested address is valid - the first requested address refers to a cached data unit (invalid) or to an uncached data unit (valid).

[0373] The input signal (used, bank 0, row) indicates whether bank 0 is used and the row requested by the first node of the group of nodes. The input signal (used, bank 0, row) fed to the leftmost node 306 (0, 0) indicates that bank 0 is not used.

[0374] For example, if the leftmost node 306(0,0) (or any other node) determines that a previously unused memory bank (not currently associated with any uncached data cells) should be used (for retrieving a valid data cell address associated with node node), the node changes the signal (USED, MEMORY BANK 0, ROW) to indicate that the memory bank is being used - and also updates "ROW" to the row requested by the node.

[0375] If any node in the first group of nodes (among nodes 306(0,1)-306(0,15)) receives a valid address indicating bank 0, the node compares the row of the requested address with the row indicated by (used, bank 0, row). If the row values ​​do not match, the node outputs a contention signal.

[0376] The same process is executed simultaneously by any group of nodes.

[0377] If there are J labels, there are J groups of nodes connected in series, and each group may contain K nodes.

[0378] Thus, each node of the group is arranged to (a) receive an access request indication (e.g., a signal (used, bank, row)) indicating whether any previous node of the group has requested access to the memory bank identified by the doublet (valid, address) and (b) update the access request indication to indicate whether the group is requesting access to its corresponding memory bank.

[0379] Data processing array (DPA) 500.

[0380] DPA 500 includes 96 DPUs arranged in six rows (lines), with each row including sixteen DPUs.

[0381] Each row of the DPU can be controlled by a separate DPA microcontroller.

[0382] Figure 15 and Figure 16 A DPU 510 is shown in accordance with an embodiment of the present invention.

[0383] The DPU 510 includes:

[0384] 1) Arithmetic Logic Unit (ALU) 540

[0385] 2) Register file 550 comprising sixteen registers 550(0)-550(15).

[0386] 3) Two output multiplexers MuxD 534 and MuxE 535.

[0387] 4) Multiple input multiplexers MuxIn0 570, MuxIn1 571, MuxIn2 572, MuxIn3 573, MuxA 561, MuxB 562, MuxCl 563, MuxCh 564, MuxF 526 and MuxG 527.

[0388] 5) Internal multiplexers MuxH 529 and MuxG' 528.

[0389] 6) Triggers 565 and 566.

[0390] 7) Registers RegA 531 , RegB 532 , RegCl 533 and RegCh 534 .

[0391] The multiplexers mentioned above are non-limiting examples of data flow components.

[0392] Each input multiplexer is coupled to an input port of the DPU 510 and may be coupled to other DPUs, to the collection unit 300, to the memory module 200, to the buffer unit 400, or to an output port of the DPU.

[0393] Input multiplexers MuxA 561, MuxB 562, MuxC1 563, MuxCh 564, and MuxH 529 also include inputs coupled to bus 581. Bus 581 is also coupled to outputs of MuxIn0 570, MuxIn1 571, MuxIn2 572, and MuxIn3 573.

[0394] Registers RegA 531, RegB 532, RegCl 533, and RegCh 534 are connected between input multiplexers MuxA 561, MuxB 562, MuxCl 563, MuxCh 564 (one register per multiplexer, respectively) and the ALU 540 and feed data to the ALU.

[0395] Using two sets of multiplexers, where one set of multiplexers can receive the output of the other set, increases the number of sources that provide data that feeds ALU 540 .

[0396] The outputs of the ALU 540 are coupled to the inputs of the register file 550. A first register file RegO 550(0) is also connected to the ALU 540 as an input.

[0397] The output of register file 550 is coupled to output multiplexers MuxD 534 and MuxE 535. The outputs of output multiplexers MuxD 534, MuxE 535 are coupled to output ports D (521) and E (522), respectively, and to output port G 523 and flip-flop 566 (via MuxG' 528).

[0398] Register RegH 539 is connected between MuxH 529 and MuxG 527. MuxF 526 is directly connected to port F 522 and flip-flop 565, thereby providing a low-latency retransmission channel. MuxG 527 is coupled to MuxG' 528.

[0399] MuxIn0, MuxIn1, MuxIn2, and MuxIn3 can implement short routing to other DPUs: (a) MuxIn0 and MuxIn1 get their input from the D outputs of the eight DPUs, and (b) MuxIn2 and MuxIn3 get their input from the E outputs of the same eight DPUs.

[0400] The other five input multiplexers, MuxA, MuxB, MuxCl, MuxCh, and MuxG, can implement the following routing:

[0401] MuxA can get its input from the buffer unit and also get input from MuxIn0..MuxIn3.

[0402] Each of MuxB, MuxCl and MuxCh may obtain its input from buffer units Muxln0..MuxIn3 and from an internal register of the register file (eg R14 or R15 of the register file).

[0403] Most ALU operations produce a short integer (Out0) result. Some operations generate a word result or two short integer results ({Out1, Out0}). The outputs are stored in constant locations in register file 550: R(0) <= Out0 and R(1) <= Out1 (for operations that produce two short integers).

[0404] The DPU 510 and other DPUs in the same row are controlled by a shared row PDA microcontroller that generates a selection information stream for selecting between configuration instructions stored in the DPUs (see configuration registers 511).

[0405] The configuration register (also called dpu_Ctrl) 511 can store the following:

[0406]

[0407]

[0408] The four configuration registers 511 ( 1 ) - 511 ( 3 ) are referred to as dpu_Inst[ 0 ] - dpu_Inst[ 1 ].

[0409] As described above, each DPU of the DPA 500 is directly coupled to some DPUs in the PMA and directly coupled (via one or more intermediate DPUs) to some other data processors in the data processor array. Each DPU has a relay channel (between ports F and G) for relaying data between the DPU's relay ports (ports F and G). This simplifies and reduces connections while providing sufficient connectivity and flexibility to efficiently perform image processing tasks.

[0410] The retransmission channel of each of the plurality of data processors (particularly the path between port F and output G of port G) exhibits substantially zero latency. This allows the use of PDUs as zero-latency retransmission channels to indirectly couple between DPUs, and also allows data to be broadcast to multiple DPUs by using retransmission channels between different DPUs.

[0411] Refer again Figure 16 Each DPU includes a core. The core includes an ALU 540 and memory resources such as a register file 550. The cores of multiple DPUs are coupled to each other through a configurable network. The configurable network includes data flow components such as multiplexers MuxA-MuxCh, MuxD-MuxE, MuxIn0-MuxIn3, MuxF-MuxH and MuxG'. These data flow components can be included in the DPU (such as Figure 16shown), but may be located at least partially outside the DPU.

[0412] The DPU may include non-relay input ports directly coupled to a first set of neighbors. For example, the non-relay input ports may include input ports A, B, C1, Ch, In0, In1, In2, and In3. Their connectivity to the first set of neighbors is listed in the example below. The first set of neighbors may include, for example, eight neighbors.

[0413] The first set of neighbors is formed by DPUs that are within a distance of less than four DPUs from the DPU (cyclic distance). The distance and direction are cyclic. For example, MuxIn0 is coupled to the D ports of DPUs that are: (a) in the same row but one column to the left (D(0, -1)), (b) in the same row but one column to the right (D(0, +1)), (c) in the same column but one row above (D(-1, 0)), (d) one row above and one column to the left (D(-1, -1)), (e) one row above and one column to the right (D(-1, +1)), (f) two rows above and in the same column (D(-2, 0)), (g) two rows above and one column to the left (D(-2, -1)), and (h) two rows above but one column to the right (D(-2, +1)).

[0414] A first non-relay input port of a data processor can be directly coupled to a relay port of a first set of neighboring data processors. For example, see port A, which is directly coupled to an F port of the same DPU and is coupled to an F port of a DPU in the same row but one column to the left (F / Fd(0,-1)).

[0415] The first retransmission ports (e.g., ports G and F) can be directly coupled to the second set of neighbors. For example, input multiplexer F (coupled to port F) can be coupled to the G output and delayed G output (Gd) of the G port of the DPU, which are (a) one row down and in the same column (G / Gd(+1,0)), (b) two rows down and in the same column (G / Gd(+2,0), (c) three rows down and in the same column (G / Gd(+3,0)), (d) the same row but one column to the left (G / Gd(0,+1)), (e) the same row but two columns to the left (G / Gd(0,+2), (f) the same row but four columns to the left (G / Gd(0,+4)), and (g) the same row but eight columns to the left (G / Gd(0,+8)).

[0416] In the following configuration example, the locations of the different configuration buffers for the PMA and DPU microcontroller configuration registers 598 (including registers p_FLIP0-p_FLIPS5 - one for each DPU microcontroller) are provided:

[0417]

[0418] Configuration registers 511(0)-511(3) can store up to four configuration instructions. Configuration instructions can be 64 bits long and can be read with one or two read operations.

[0419] Configuration instructions control the multiplexer selects (A, B, C, D, E, F, and G) and register file 550 shifts:

[0420] 1) Shift as in step 1: for n in [0,15]: for (i=15; i>0; i--) Ri<=R(i-1).

[0421] 2) Shift by step 2: for n in [0,2,4...14]: for (i=7; i>0; i--) {R(2i+1), R(2i)}<={R(2i-1), R(2i-2)}

[0422] These fields can be applied at different times:

[0423] 1) Input multiplexers A, B, C, F and output multiplexer G controls are not delayed.

[0424] 2) ALU control is delayed by one clock cycle.

[0425] 3) Output multiplexers D and E control the delay of two clock cycles.

[0426] The following table describes the different fields of the DPU configuration command:

[0427]

[0428]

[0429]

[0430]

[0431] The input multiplexers MuxA, MuxB, MuxC, MuxF, the output multiplexers MuxD, MuxE, output F, delayed output Fd, output G, and delayed output Gd provide the connectivity listed in this paragraph. The symbol X(N,M) represents the output X(row+N%6, col+M%16) of the DPU. The DPUs of each row are connected to each other in a cyclic manner, and the DPUs of each column are connected to each other in a cyclic manner. It should be noted that the various multiplexers listed below have multiple (e.g., sixteen) inputs, and the following list provides connections to each of these inputs. For example, A[0]-A

[15] are the sixteen inputs of MuxA. In the following, 6% represents modulo 6 operation, and 16% represents modulo 16 operation. R14 and 51R are the last two registers of the register file.

[0432]

[0433]

[0434]

[0435]

[0436]

[0437] Input (A, B, Cl, Ch, G) register configuration specifications (configuration bits are stored in the DPU's configuration registers and are used to control various components of the DPU).

[0438] The following example lists the values ​​of the various bits contained in the DPU configuration instruction. The symbol X(N, M) represents the output of the DPU X(row+N%6,col+M%16).

[0439]

[0440]

[0441]

[0442]

[0443]

[0444]

[0445]

[0446]

[0447] The inputs (A, B, Cl, Ch, G) have two additional inputs multiplexed to (MuxIn0, ..., MuxIn3). In order to save configuration bits, dynamic resource allocation of (MuxIn0, ..., MuxIn3) to the D(N, M) and E(N, M) configuration bits of (A, B, Cl, Ch, G) can be used. The allocation can be done as follows: if one or two of the input (A, B, Cl, Ch, G) configurations have the form D(N, M), the first input is allocated to MuxIn0, and the second (if present) is allocated to MuxIn1. For convenience, the inputs are represented by inputD1 and inputD2, and the configurations are represented by D1(N, M) and D2(N, M), respectively. The control of the MuxIn0 and MuxIn1 multiplexers can be based on D1(N, M) and D2(N, M), respectively. In addition, inputD1 and inputD2 multiplex control MuxIn0 and MuxIn1. If one or two of the input (A, B, Cl, Ch, G) configurations are of the form E(N, M), then the same dynamic allocation applies to MuxIn2 and MuxIn3.

[0448] For example:

[0449] Assume the following configuration:

[0450]

[0451]

[0452] The control of the multiplexer will then be as follows:

[0453]

[0454] ALU opcodes

[0455] Integer operations

[0456]

[0457]

[0458]

[0459] Floating-point operations

[0460] The C exponent bias parameter in the conversion can be viewed as a 7-bit signed integer (the other 25 bits of C are ignored and the sign bit is extended)

[0461]

[0462] Addition

[0463]

[0464] DPU microcontroller instruction encoding

[0465] Execute the command:

[0466]

[0467]

[0468] The DPUs may be configured one at a time (each DPU has a unique unicast address) or may be configured in broadcast mode - there are row and / or column addresses that may reflect the DPUs that share rows and / or columns and this allows configuration information to be broadcasted.

[0469] Data processing array: linear mapping (programming a single DPU configuration register)

[0470]

[0471]

[0472] It should be noted that any image processing algorithm can be executed in an iterative manner by the image processor. Results regarding some pixels are processed by the DPA 500. Some results may be stored in the DPA for a period of time and then sent to the memory module. A certain time period is usually set based on the size of the PMA's memory resources and the number of source pixels or target pixels processed by the DPA during a certain task. Once these results are needed again, they may be retrieved from the memory module. For example, when the DPA 500 performs calculations on certain source pixels of the source image, these results can be stored for a period of time (for example, when performing calculations related to adjacent source pixels) and then sent to the memory. When the results are further needed, they can be retrieved from the memory module.

[0473] Distortion calculation

[0474] Warp calculations may be applied for various reasons, for example, to compensate for imbalances in image acquisition.

[0475] The warp calculations may be performed by the DPA 500 .

[0476] According to an embodiment of the present invention, a distortion calculation is applied to each target pixel (pixel of a target image) in a group of target pixels. The target pixel group can include the entire target image or a portion of the target image. Typically, the target image is virtually segmented into multiple windows, and each window is a group of target pixels.

[0477] The warp calculation may receive or may calculate a corresponding set of source pixels. Source pixels in the corresponding set of source pixels are processed during the warp calculation. The selection of the corresponding set of source pixels is typically fed to the PMA and may depend on, for example, a desired warp function.

[0478] The warp value of the target pixel is calculated by applying weights (Wx, Wy) to the neighboring source pixels associated with the target pixel. The weight and coordinates (x, y) of at least one neighboring source pixel are defined with warp parameters (X', Y').

[0479] Figure 17 A method 1700 according to an embodiment of the present invention is shown.

[0480] The method 1700 may begin with selecting a target pixel from a group of target pixels at step 1710. The selected target pixel will be referred to as a "target pixel."

[0481] Step 1710 may be followed by step 1720 of performing a warp calculation process for each target pixel in a set of target pixels, which includes:

[0482] 1) Calculate (1721) or receive warp parameters for a selected destination pixel. The warp parameters may include first and second weights (Wx, Wy) and coordinates (x, y) for a given source pixel that should be processed during the warp calculation. The first and second weights are received by a first set of processing units (DPUs) in a processing unit array (DPA).

[0483] 2) Request 1722 the neighboring source pixels (which include the given source pixel) from a memory unit such as a collection unit. The collection unit can receive 4 coordinates and convert them into 16 source pixels - four groups of neighboring source pixels - in various modes of operation.

[0484] 3) Receive (1723) the adjacent source pixels associated with the destination pixel via the second set of processing units.

[0485] 4) Calculating (1724) a warping result responsive to the values ​​of adjacent source pixels and a pair of weights by a second set of processing units; and providing the warping result to a memory module.

[0486] Steps 1721 , 1722 , 1723 , and 1724 may be performed in a pipelined manner.

[0487] refer to Figure 18 , a first group of processing units is represented as 505 and may include the leftmost four DPUs in the top first row of DPA 500. A second group of processing units is represented as 501 and may include the rightmost two columns of DPA 500.

[0488] Following step 1720 is step 1730, which checks whether the warp has been calculated for all target pixels of the group. If not, the warp calculation is terminated.

[0489] Step 1726 may include relaying the values ​​of some adjacent source pixels among the processing units in the second group.

[0490] Figure 18 and Figure 19 The output signal of DPU (0, 4) (X' of group 504) is shown to be sent to DPU (0, 15) and then relayed to DPU (1, 15). It should be noted that Figure 18 and Figure 19 In , PMA computes the warping function for four pixels in parallel:

[0491] 1) DPU (0, 3), DPU (1, 3) and the DPUs in group 501 are involved in calculating the distortion of the first pixel.

[0492] 2) DPU (0, 2), DPU (1, 2), and the DPU of group 502 are involved in calculating the distortion of the second pixel.

[0493] 3) DPU (0, 1), DPU (1, 1), and the DPU of group 503 are involved in calculating the distortion of the third pixel.

[0494] 4) DPU (0, 0), DPU (1, 0), and the DPU of group 504 are involved in calculating the distortion of the third pixel.

[0495] Step 1726 may include relaying intermediate results computed by the second group and values ​​of some adjacent source pixels between processing units of the second group.

[0496] Figure 20 Warp parameters (X' for groups 501-504 and Y' for groups 501-504)) are shown being sent from the DPUs (510(0,0)-510(0,3) and 510(1,0)-510(1,3)) of groups 505 and 506 to groups 501, 502, 503 and 504.

[0497] Figure 21 and Figure 22 Warp computations performed by DPUs 510 ( 0 , 15 ) - 510 ( 3 , 15 ) and DPUs 510 ( 0 , 14 ) - 510 ( 5 , 14 ) of group 501 are shown in accordance with an embodiment of the present invention.

[0498] Figure 21 The distortion calculation of includes the following steps (some of which are performed in parallel with each other). Steps 1751-1762 are also Figure 22 Shown in.

[0499] A first difference value (P0-P2) between a first pair of adjacent source pixels and a second difference value (P1-P3) between a second pair of adjacent source pixels are calculated (1751) by a first processing unit (DPU 510(5,14)) of the second group.

[0500] The first difference value is provided (1752) to a second processing unit of the second group (DPU 510(1, 14)), and the second difference value is provided to a third processing unit of the second group.

[0501] A first modified weight Wy' is calculated (1753) by the fourth processing unit of the second group (DPU 510(1, 15)) in response to the first weight.

[0502] The first modified weights are provided (1754) from the fourth processing unit to the second processing unit of the second group (DPU 510(1, 14)).

[0503] A first intermediate result (Var0) is calculated (1755) by a second processing unit of the second group based on the first difference (P0-P2), the first adjacent source pixel (P0) and the first modified weight (Wy'): Var0=(P0-P0)*Wy'-P0.

[0504] The second difference value (P1-P3) is provided (1756) from the third processing unit of the second group to the sixth processing unit of the second group (DPU 510(0,15)).

[0505] The second adjacent source pixel (P1) is provided (1757) from the fifth processing unit of the second group (DPU 510(0,14)) to the sixth processing unit of the second group (DPU 510(2,14)).

[0506] A second intermediate result Var1 is calculated (1758) by the sixth processing unit of the second group based on the second difference value, the second adjacent source pixel and the first modified weight: Var1 = (P1-P3)*Wy'-P1.

[0507] The second intermediate result Var1 from the sixth processing unit of the second group is provided (1759) to the seventh processing unit of the second group (DPU 510(2,15)), and the first intermediate result Var0 is provided from the second processing unit of the second group to the seventh processing unit of the second group.

[0508] The seventh processing unit of the second group calculates (1760) a third intermediate result Var2 relative to the first intermediate result and the second intermediate result: Var2 = Var0 - Var1.

[0509] The third intermediate result is provided (1761) from the seventh processing unit of the second group to the eighth processing unit of the second group (DPU 510 (3, 15)). The second intermediate result is provided from the sixth processing unit of the second group to the ninth processing unit of the second group (DPU 510 (3, 14)).

[0510] The second intermediate result is provided (1762) from the ninth processing unit of the second group to the eighth processing unit of the second group. The second modified weight (Wx') is provided from the third processing unit of the second group to the eighth processing unit of the second group.

[0511] The warping result is calculated (1763) by the eighth processing unit of the second group based on the second intermediate result and the third intermediate result and the second modification weight: Warp_result=Var2*Wx'+Var1.

[0512] like Figure 20 As shown, DPU 510 (5, 14) may receive pixels P0, P1, P2, and P3 from the collection unit. When DPA 500 processes four pixels at a time, groups 501, 502, 503, and 504 receive 16 pixels from the collection unit (in parallel).

[0513] It should be noted that the DPA 500 also receives (eg, from a collection unit) distortion parameters X', Y' associated with each pixel.

[0514] According to one embodiment of the present invention, the warping parameter of each pixel may be calculated by the DPU of the DPA—for example, when the warping parameter can be expressed by a mathematical formula such as a polynomial.

[0515] Figure 23 A set of DPUs 507 is shown calculating X′ and Y′, and these calculated X′ and Y′ may be fed to sets 505 and 506 .

[0516] It should be noted that Figures 18 to 22 Only a non-limiting grouping scheme is shown. The warp calculations can be performed by groups of DPUs of other shapes and sizes.

[0517] Parallax

[0518] The disparity calculation aims to find the best matching destination pixel for a source pixel. The search can be performed for all source pixels in the source image and for all destination pixels in the destination image - but this is not necessarily the case, and the disparity may be applied to only some source pixels of the source image and / or some destination pixels of the destination image.

[0519] Disparity calculations not only compare the difference between a single source pixel and a single destination pixel, but also compare a subset of source pixels to a subset of destination pixels. The comparison may include calculating a function such as the sum of absolute differences (SAD) between a source pixel and the corresponding destination pixel.

[0520] The source pixel may be located at the center of a subset of source pixels, and the destination pixel may be located at the center of a subset of destination pixels.Other positions of the source pixel and / or destination pixel may be used.

[0521] The subset of source pixels and the subset of destination pixels may be rectangular in shape (or may have any other shape) and may include N rows and N columns, where N may be an odd positive integer that may exceed three.

[0522] Most of the disparity calculations may benefit from previous computer disparity calculations. These are shown in Figure 24 and 25 is provided.

[0523] Figure 24 Shown are a first subset 1001 of 5×5 source pixels S(1,1)-S(5,5), a first subset 1002 of 5×5 target pixels T(1,1)-T(5,5), a second subset 1003 of 5×5 source pixels S(1,2)-S(5,6), and a second subset 1004 of 5×5 target pixels T(1,2)-T(5,6).

[0524] Source pixels S(3,3) and S(3,4) are located at the center of the first and second subsets 1001 and 1003 of source pixels. Destination pixels T(3,3) and T(3,4) are located at the center of the first and second subsets 1002 and 1004 of destination pixels.

[0525] The SAD associated with S(3,3) and T(3,3) is equal to:

[0526] SAD(S(3,3),T(3,3))=SUM(|S(i,j)−T(i,j)|)—for indices i and j between 1 and 5.

[0527] The SAD associated with S(3,4) and T(3,4) is equal to:

[0528] SAD(S(3,4), T(3,4))=SUM(|S(i,j)−T(i,j)|)—for index i between 2 and 6 and for index j between 1 and 5.

[0529] Assume that SAD is calculated from left to right. Under this assumption, the calculation of SAD(S(3,4), T(3,4)) may benefit from the calculation of SAD(S(3,3), T(3,3)).

[0530] In particular: SAD(S(3,4), T(3,4)) = SAD(S(3,3), T(3,3)) - SAD(rightmost column of the first subset of source and destination pixels) + SAD(leftmost column of the second subset of source and destination pixels).

[0531] Since the source and destination images are two-dimensional, and assuming that the source pixels are scanned from left to right (per slice) and from top to bottom - then the calculation of SAD is more efficient.

[0532] Figure 25 A subgroup SG(B) of source pixels is shown with a central pixel SB. Figure 26 A corresponding sub-group TG(B) of target pixels (not shown) is shown with a central pixel TB.

[0533] The SUD is calculated for source pixels in the row above the row of the SB and for pixels located to the left of the SB and on the same row.

[0534] The pixel SA is the center of the sub-group SG(A) and is the left neighbor of the pixel SB. The target pixel TA is the left neighbor of the pixel SB and is the center of the sub-group TG(A).

[0535] The pixel SC is the center of the sub-group SG(C) and is the upper neighbor of the pixel SB. The target pixel TC is the upper neighbor of the pixel SB and is the center of the sub-group TG(C).

[0536] The leftmost column of SG(A) is represented as 1110. The rightmost column of SG(C) is represented as 1114. The current rightmost column of SG(B) is represented as 1115. The rightmost lowest pixel of SG(B) (also referred to as the new source pixel NSP) is represented as 1116. The old pixel (belonging to SG(C)) at the top of the current rightmost column of SG(B) (also referred to as the old source pixel NSP) is represented as 1112.

[0537] The leftmost column of TG(A) is represented as 1110′. The rightmost column of TG(C) is represented as 1114′. The current rightmost column of TG(B) is represented as 1115′. The rightmost lowest pixel of TG(B) (also referred to as the new target pixel NTP) is represented as 1116′. The old pixel (belonging to TG(C)) at the top of the current rightmost column of TG(B) (also referred to as the old target pixel NTP) is represented as 1112′.

[0538] Calculating the SAD of (SB, TB) can be equal to:

[0539] SAD (SA, TA).

[0540] -SAD(leftmost column of SG(A), leftmost column of SG(B)).

[0541] +SAD(rightmost column of SG(C)), rightmost column of TG(C)).

[0542] +The absolute difference between the rightmost lowest source and destination pixels of SG(B) and TG(B).

[0543] - The absolute difference between the top source and destination pixels in the rightmost columns of SG(C) and TG(C).

[0544] Figure 27 Method 2600 according to an embodiment of the present invention is shown.

[0545] Method 2600 may begin at step 2610 by selecting source pixels and selecting a subset of target pixels. The subset of target pixels may be an entire target image that is a portion of a target image.

[0546] Step 2610 may be followed by step 2620 of computing a set of sums of absolute differences (SADs) by a first set of data processors of the data processor array.

[0547] The set of SADs is associated with the subset of target pixels and source pixels including the target pixel selected in step 2610. Different SADs for the set are calculated from different target pixels in the subset of (the same) source pixel and target pixel.

[0548] Calculating the set of SADs for the same source pixels reduces the amount of data fetched to the DPA.

[0549] The subset of target pixels may include target pixels sequentially stored in a memory module. Prior to calculating the group SAD, the subset of target pixels is extracted from the memory module. Extracting the subset of target pixels from the memory module is performed by a collection unit including a content addressable memory buffer.

[0550] Each SAD is calculated based on the previously calculated SAD and based on the absolute differences between the currently calculated other source pixels and the other target pixels belonging to the subgroup of target pixels. Figure 25 An example of such a calculation is provided.

[0551] Step 2620 may be followed by step 2630 of finding, by a second set of data processors in the array, a best matching target pixel in the subset of target pixels in response to the set of SAD values.

[0552] Steps 2620 and 2630 may include storing the results of the calculations in the data processor array - the SAD for the entire rectangular pixel array, the SAD for a column, etc. It should be noted that the depth of each DPU's register file can be long enough to store the SAD for the rightmost column of the previous rectangular array. For example - if there are 15 columns in SG(A), then the DPU's register file 550 should be at least 15.

[0553] After storing the previous SADs, then for a given SAD, the first previously calculated SAD, the second previously calculated SAD, the destination pixel at the top of the second destination pixel column, and the source pixel at the top of the second source pixel column are stored.

[0554] Referring to step 2620, the first previously calculated SAD may reflect the absolute difference between: (i) a rectangular source pixel array that differs from a given rectangular source pixel array by a first source pixel column and a second source pixel column, and (ii) a rectangular destination pixel array that differs from a given rectangular destination pixel array by a first destination pixel column and a second destination pixel column. For example, SAD(SA, TA).

[0555] The second previously calculated SAD may reflect the absolute difference between the first source column and the second source column, for example - SAD (leftmost column of SG(A)), leftmost column of SG(B)).

[0556] Step 2620 may include:

[0557] 1) by subtracting from a first previously calculated SAD (e.g., SAD(SA, TA)) (a) a second previously calculated SAD (e.g., -SAD(leftmost column of SG(A)), leftmost column of SG(B)), and (b) the absolute difference between: (i) the destination pixel located at the top of the second destination pixel column, and (ii) the source pixel located at the top of the second source pixel column (e.g., -absolute difference between OSP 1112 and OTP 1112').

[0558] 2) Add the absolute difference between the lowest destination pixel of the second destination pixel column and the lowest source pixel of the second source pixel column (eg, the absolute difference between -NSP 1116 and NTP 1116') to the intermediate result.

[0559] It should be noted that finding the best matching target pixel may involve an iterative process and multiple repetitions of steps 2610, 2620 and 2630 may be performed - for different subsets of pixels and by comparing the results of these multiple iterations - the best matching target pixel in the group of target pixels may be found.

[0560] Note also that the processing unit array may perform multiple disparity calculations (for different source pixels and / or for different destination pixels) in parallel.

[0561] Figure 28 Eight source pixels and thirty-two destination pixels are shown being processed by DPA according to an embodiment of the present invention. Figure 29 A source pixel array according to an embodiment of the present invention is shown. Figure 30 A target pixel array according to an embodiment of the present invention is shown.

[0562] The SAD associated with the following pixels is calculated: source pixels (SP0, SP1, SP2 and SP3) and (SP'0, SP'1, SP'2 and SP'3), 4×8 destination pixels (including the leftmost column of TP0, TP1, TP2 and TP3), and another 4×8 destination pixels (including the leftmost column of TP'0, TP'1, TP'2 and TP'3).

[0563] Source pixels SP0, SP1, SP2 and SP3 belong to the same column, and their SADs are calculated in a pipelined manner:

[0564] 1) Calculate the SAD for SP0 and a target pixel.

[0565] 2) Use the previous calculation when calculating the SAD of SP1 and a certain target pixel.

[0566] 3) Use the previous calculation when calculating the SAD of SP2 and a certain target pixel.

[0567] 4) Use the previous calculation when calculating SP3 and the SAD of a certain target pixel.

[0568] While calculating the SAD of source pixels SP0, SP1, SP2 and SP3, PMA also calculates the SAD of SP'0, SP'1, SP'2 and SP'3. SP'0, SP'1, SP'2 and SP'3 belong to the same column and their SADs are calculated in a pipelined manner:

[0569] 1) Calculate the SAD for SP'0 and a target pixel.

[0570] 2) Use the previous calculation when calculating the SAD between SP'1 and a certain target pixel.

[0571] 3) Use the previous calculation when calculating the SAD of SP'2 and a certain target pixel.

[0572] 4) Use the previous calculation when calculating the SAD of SP'3 and a certain target pixel.

[0573] DPA 500 may calculate the SAD for each source pixel and multiple other destination pixels in parallel.

[0574] For example, assuming that the first row of 4×8 target pixels includes TP0 and seven shifted target pixels (TP0, Ts1P0, Ts2P0, Ts2P0, Ts3P0, Ts4P0, Ts5P0, Ts6P0, Ts7P0), the calculation of the SAD of SP0 may include calculating the SAD for SP0 and each of TP0, Ts1P0, Ts2P0, Ts2P0, Ts3P0, Ts4P0, Ts5P0, Ts6P0, Ts7P0.

[0575] When calculating any SAD, you need to calculate the absolute difference of the new pixel. Figure 29 Four new source pixels NSO, NS1, NS2 and NS3 are shown (for computing the SAD associated with SP0, SP1, SP2 and SP3 and only one destination pixel column).

[0576] Figure 30 The 32 new target pixels are shown:

[0577] 1) A new target pixel for calculating the SAD between SP0 and eight different target pixels - NT0, Ns1T0, Ns2T0, Ns3T0, Ns4T0, Ns5T0, Ns6T0, Ss7T0.

[0578] 2) A new target pixel for calculating the SAD of SP1 and eight different target pixels - NT1, Ns1T1, Ns2T1, Ns3T1, Ns4T1, Ns5T1, Ns6T1, Ss7T1.

[0579] 3) A new target pixel for calculating the SAD of SP2 and eight different target pixels - NT2, Ns1T2, Ns2T2, Ns3T2, Ns4T2, Ns5T2, Ns6T2, Ss7T2.

[0580] 4) A new target pixel for calculating the SAD of SP3 and eight different target pixels - NT3, Ns1T3, Ns2T3, Ns3T3, Ns4T3, Ns5T3, Ns6T3, Ss7T3.

[0581] Figure 31 Groups of eight DPUs 1131 , 1132 , 1133 , 1134 , 1135 , 1136 , 1137 , and 1138 are shown—each group including 4 DPUs.

[0582] Each group 1131 , 1132 , 1133 and 1134 calculates the SAD for pixels SP0 , SP1 , SP2 and SP3 - but for different target pixels ( TP0 , TP2 , TP3 and TP4 ).

[0583] Each group 1135, 1136, 1137 and 1138 calculates the SAD of pixels SP'0, SP'1, SP'2 and SP'3 - but for different target pixels (TP0, TP2, TP3 and TP4)

[0584] Pixel group 1140 performs a minimization operation on the SAD calculated by groups 1131-1138.

[0585] Accordingly, method 2600 may include calculating, by a first group of data processors in a data processor array, multiple sets of SADs associated with multiple subsets of multiple source pixels and target pixels; wherein each SAD in the multiple sets of SADs is calculated based on the absolute difference between a previously calculated SAD and a currently calculated SAD; and finding, by a second group of data processors in the array and for the source pixel, the best matching target pixel in response to the value of the SAD associated with the source pixel.

[0586] The multiple sets of SADs may include sub-sets of SADs, each sub-set of SADs being associated with a plurality of source pixels and a plurality of sub-sets of destination pixels in the plurality of sub-sets of destination pixels. For example, the sets 1131-1138 calculate different sub-sets of SADs.

[0587] A plurality of source pixels may belong to columns of a rectangular pixel array and be adjacent to each other.

[0588] Calculating the multiple groups of SADs may include calculating the SADs of different SAD subgroups in parallel.

[0589] The calculating may include calculating SADs belonging to the same SAD subgroup in a sequential manner.

[0590] The following text describes some PMA state and configuration buffers according to an embodiment of the present invention.

[0591] These status and configuration buffers 109 include a PMA control status register, a PMA suspend enable control register, and a PMA suspend event status register.

[0592] The control registers may allow, for example, the scalar unit to determine a predetermined period of operation for the image processor. Additionally or alternatively, the scalar unit may suspend the image processor (without changing the state of the PMA) and program the program processor, send a control signal to the program processor, and resume operation of the image processor from the same point (except for the changes introduced by the scalar unit).

[0593] PMA Control Status Register (p_PmaCsr)

[0594]

[0595]

[0596] PMA Halt Enable Control Register (HaltOnEvent)

[0597]

[0598] PMA Halt Event Status Register (HoeStatus)

[0599]

[0600] This function enables suspend operation, changes certain configurations without flushing the computation pipeline, and resumes operation. It is implemented through the following registers: (a) suspend counter enable control bit (suspCntEn in p_PmaCsr), (b) suspend counter (p_SuspCnt), and (c) suspend reset control (p_RstCtl).

[0601] When enabled (suspCntEn=1), the suspend counter counts down. When it reaches zero, the PMA suspends operation (remains in stalled state) until suspCntEn is reset or p_SuspCnt is written with a new value (!=0). During stall, the scalar unit can reconfigure the PMA (instructions, constants, etc.). When suspCntEn is reset or p_StallCnt is written, the PMA will resume its operation with the new configuration. p_RstCtl defines which functions are reset on resume.

[0602] The features that can be reset are:

[0603] 1)DPU microcontroller.

[0604] 2)DPA program memory.

[0605] 3)BU program memory.

[0606] 4)SB program memory.

[0607] 5) Address generator.

[0608] 6)BU reads the buffer.

[0609] 7) GU iterator (1 bit)

[0610] Reset control on suspend

[0611]

[0612] Event counter p_EventCnt

[0613] A simple counter that counts with the DPU clock (does not count during p_stall). The counter is preset via configuration. Whenever the counter is empty, an event is signaled to the scalar unit. This counter is readable via the configuration bus.

[0614] Suspend and event increment register p_SuspEventInc

[0615] The lower half (15..0) is used to increment the pending counter, both during and simultaneously with its normal decrement. The upper half (31..16) also increments the event counter.

[0616] Figure 33 Method 3300 according to an embodiment of the present invention is shown.

[0617] Method 3300 may begin by selecting a source pixel from a group of source pixels at step 3310. The selected source pixel will be referred to as a "source pixel."

[0618] Step 3310 may be followed by step 3320 of performing a warp calculation process for each source pixel in the group of source pixels, including:

[0619] 1) Calculate (3321) or receive warp parameters for a selected source pixel. The warp parameters may include first and second weights (Wx, Wy) and coordinates (x, y) of a given destination pixel that should be processed during the warp calculation. The first and second weights are received by a first set of processing units (DPUs) in a processing unit array (DPA).

[0620] 2) Request 3322 the neighboring target pixels (which include the given target pixel) from a memory unit such as a collection unit. The collection unit can receive 4 coordinates and convert them into 16 target pixels - four groups of neighboring target pixels - in various modes of operation.

[0621] 3) Neighboring destination pixels associated with the source pixel are received (3323) by a second set of processing units.

[0622] 4) Calculating (3324) a warping result by the second set of processing units in response to the values ​​of the adjacent target pixels and a pair of weights; and providing the warping result to the memory module.

[0623] Steps 3321, 3322, 3323, and 3324 may be performed in a pipeline manner.

[0624] Step 3320 is followed by step 3330 of checking if the warp is calculated for all source pixels of the group. If not - end the warp calculation.

[0625] Step 3326 may include relaying the values ​​of some adjacent target pixels among the processing units of the second group.

[0626] Any reference to any of the terms "comprise," "comprises," "comprising," "including," "may include," and "includes" may be applied to the terms "consists," "consisting," and "and consisting essentially of." For example—any method describing steps may include more steps than shown in a figure, only the steps shown in a figure, or substantially only the steps shown in a figure. The same applies to components of a device, processor, or system, and instructions stored in any non-transitory computer-readable storage medium.

[0627] The present invention may also be implemented in a computer program for running on a computer system, the computer program comprising at least code portions for executing the steps of the method according to the present invention when run on a programmable device (such as a computer system) or for enabling the programmable device to perform the functions of the apparatus or system according to the present invention. The computer program may cause the storage system to assign the hard disk drive to the hard disk drive group.

[0628] A computer program is a list of instructions for a particular application and / or operating system. A computer program may include, for example, one or more of a subroutine, a function, a procedure, an object method, an object implementation, an executable application, an applet, a servlet, source code, object code, a shared library / dynamically loaded library, and / or other sequence of instructions designed to be executed on a computer system.

[0629] The computer program may be stored internally on a non-transitory computer-readable medium. All or some of the computer programs may be provided on a computer-readable medium that is permanently, removably, or remotely coupled to an information processing system. The computer-readable medium may include, for example, but not limited to, any number of the following: magnetic storage media, including hard disks and tape storage media; optical storage media, such as optical disk media (e.g., CD-ROMs, CD-Rs, etc.) and digital video disk storage media; non-volatile memory storage media, including semiconductor-based memory cells, such as flash memory, EEPROM, EPROM, ROM; ferromagnetic digital memory; MRAM; volatile storage media, including registers, buffers or caches, main memory, RAM, etc.

[0630] A computer process typically consists of an executing (running) program or part of a program, current program values ​​and state information, and resources used by the operating system to manage the execution of the process. An operating system (OS) is software that manages the sharing of computer resources and provides programmers with an interface for accessing those resources. The OS processes system data and user input and responds by allocating and managing tasks and internal system resources as services to the system's users and programs.

[0631] A computer system may, for example, include at least one processing unit, associated memory, and a plurality of input / output (I / O) devices.When executing a computer program, the computer system processes information according to the computer program and produces resulting output information via the I / O devices.

[0632] In the foregoing specification, the invention has been described with reference to specific examples of embodiments of the invention. It will, however, be evident that various modifications and changes may be made therein without departing from the broader spirit and scope of the invention as set forth in the appended claims.

[0633] Furthermore, the terms "front," "back," "top," "bottom," and the like in the specification and claims, if any, are used for descriptive purposes and not necessarily for describing permanent relative positions. It is understood that the terms so used are interchangeable under appropriate circumstances such that the embodiments of the invention described herein are, for example, capable of operation in other orientations than those illustrated or otherwise described herein.

[0634] The connection discussed herein can be any type of connection suitable for transmitting signals to and from related nodes, units or devices, for example, via an intermediate device. Therefore, unless implied or otherwise stated, a connection can be, for example, a direct connection or an indirect connection. Connections can be illustrated or described with reference to a single connection, multiple connections, unidirectional connections or bidirectional connections. However, different embodiments can change the implementation of the connection. For example, a separate unidirectional connection can be used instead of a bidirectional connection, or vice versa. Moreover, multiple connections can be replaced by a single connection that transmits multiple signals in serial or time-division multiplexing mode. Similarly, a single connection carrying multiple signals can be separated into various different connections that carry subgroups of these signals. Therefore, there are many options for transmitting signals.

[0635] Although specific conductivity types or polarities of potentials have been described in the embodiments, it should be understood that the conductivity types and polarities of potentials may be reversed.

[0636] Each signal described herein can be designed as either positive logic or negative logic. In the case of a negative logic signal, the signal is active low when its logically true state corresponds to a logic level zero. In the case of a positive logic signal, the signal is active high when its logically true state corresponds to a logic level zero. Note that any signal described herein can be designed as either a negative logic signal or a positive logic signal. Therefore, in alternative embodiments, those signals described as positive logic signals can be implemented as negative logic signals, and those signals described as negative logic signals can be implemented as positive logic signals.

[0637] Furthermore, when referring to a signal, status bit, or the like assuming its logically true or logically false state, the terms "assertion" or "setting" and "negation" (or "deactivation" or "clearing") are used herein for its logically true or logically false state, respectively. If the logically true state is a logic level 1, the logically false state is a logic level 0. If the logically true state is a logic level 0, the logically false state is a logic level 1.

[0638] Those skilled in the art will recognize that the boundaries between logic blocks are merely illustrative and that alternative embodiments may merge logic blocks or circuit elements, or impose an alternative decomposition of functionality on various logic blocks or circuit elements. Therefore, it should be understood that the architectures described herein are merely exemplary and that in fact many other architectures that achieve the same functionality may be implemented.

[0639] Any arrangement of components that achieve the same functionality is effectively "associated" so that the desired functionality is achieved. Thus, any two components combined herein to achieve a particular functionality can be considered to be "associated" with each other so that the desired functionality is achieved, regardless of architecture or intermediary components. Likewise, any two components so associated can also be considered to be "operably connected" or "operably coupled" to each other so that the desired functionality is achieved.

[0640] Furthermore, those skilled in the art will recognize that the boundaries between the above-described operations are illustrative only. Multiple operations may be combined into a single operation, a single operation may be distributed among additional operations, and operations may be performed at least partially overlapping in time. Furthermore, alternative embodiments may include multiple instances of a particular operation, and the order of the operations may be changed in various other embodiments.

[0641] Also for example, in one embodiment, the illustrated examples may be implemented as circuits located on a single integrated circuit or within the same device. Alternatively, the examples may be implemented as any number of separate integrated circuits or separate devices interconnected with each other in a suitable manner.

[0642] Also for example, the examples or portions thereof may be implemented as software or code representations of physical circuitry, or as software or code representations of logical representations convertible into physical circuitry, such as in any appropriate type of hardware description language.

[0643] Furthermore, the present invention is not limited to physical devices or units implemented in non-programmable hardware, but may also be employed in programmable devices or units capable of performing the desired device functions by operating according to appropriate program code, such as mainframes, minicomputers, servers, workstations, personal computers, notepads, personal digital assistants, electronic games, automotive and other embedded systems, cellular telephones and various other wireless devices, generally referred to in this application as "computer systems."

[0644] However, other modifications, changes, and substitutions are possible. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.

[0645] In the claims, any reference signs placed between parentheses should not be construed as limiting the claim. The word "comprising" does not exclude the presence of other elements or steps beyond those listed in the claim. In addition, the terms "a" or "an" as used herein are defined as one or more than one. In addition, the use of introductory phrases such as "at least one" and "one or more" in the claims should not be interpreted as implying that any particular claim containing the introduced claim element by introducing another claim element through the indefinite article "a" or "an" should be limited to an invention containing only one of that element, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an". The same applies to the use of definite articles. Unless otherwise specified, terms such as "first" and "second" are used to arbitrarily distinguish between the elements described by such terms. Therefore, these terms are not necessarily intended to indicate a temporal or other priority of these elements. The fact that certain measures are cited in different claims does not indicate that a combination of these measures cannot be used to advantage.

[0646] While certain features of the present invention have been shown and described herein, many modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is therefore intended that the appended claims cover all such modifications and changes that fall within the true spirit of the invention.

Claims

1. A data processing module, comprising a data processor array; wherein: each data processor unit in a plurality of data processors in the data processor array is directly coupled to one or more data processors in the data processor array, is indirectly coupled to one or more other data processors in the data processor array, and includes a relay channel for relaying data between relay ports of the data processors, Wherein, the relay channel of each data processing unit includes: (i) a first path consisting of a single multiplexer, (ii) a second path consisting of a multiplexer and a flip-flop; (iii) a third path consisting of two multiplexers, and (iv) The fourth path consists of two multiplexers and a flip-flop.

2. The data processing module according to claim 1, wherein: The first path, the second path, the third path, and the fourth path of the relay channel of each of the plurality of data processors are coupled in parallel.

3. The data processing module according to claim 1, wherein: Each data processor of the plurality of data processors comprises a core; wherein the core includes an arithmetic logic unit and memory resources; The cores of the plurality of data processors are coupled to each other via a configurable network.

4. The data processing module according to claim 3, wherein: Each data processor of the plurality of data processors includes a plurality of data flow components of the configurable network.

5. The data processing module according to claim 1, wherein: Each data processor of the plurality of data processors includes a first non-rebroadcast input port directly coupled to a first set of neighbors.

6. The data processing module according to claim 5, wherein: The first set of neighbors is formed by data processors located within a distance of less than four data processors from the data processor.

7. The data processing module according to claim 5, wherein: The first non-rebroadcast input port of the data processor is directly coupled to the rebroadcast ports of the data processors of the first set of neighbors.

8. The data processing module according to claim 7, wherein: The data processor also includes a second non-rebroadcast input port directly coupled to the non-rebroadcast ports of the data processors of the first set of neighbors.

9. The data processing module according to claim 5, wherein: The first non-rebroadcast input port of the data processor is directly coupled to the non-rebroadcast ports of the data processors of the first set of neighbors.

10. The data processing module according to claim 5, wherein: The first group of neighbors consists of eight data processors.

11. The data processing module according to claim 5, wherein: The first repeating port of each data processor of the plurality of data processors is directly coupled to a second set of neighbors.

12. The data processing module according to claim 11, wherein: For each data processor of the plurality of data processors, the second set of neighbors is different from the first set of neighbors.

13. The data processing module according to claim 11, wherein: For each data processor of the plurality of data processors, the second set of neighbors includes data processing units that are farther away from the data processor than any data processor belonging to the first set of neighbors.

14. The data processing module according to claim 1, wherein: In addition to the plurality of data processors, the processor array further comprises at least one other data processor.

15. The data processing module according to claim 1, wherein: The data processors in the data processor array are arranged in rows and columns.

16. The data processing module according to claim 15, wherein: One or more data processors in each row are coupled to each other in a cyclic manner.

17. The data processing module according to claim 15, wherein: The data processors in each row are controlled by a shared microcontroller.

18. The data processing module according to claim 15, wherein: Each of the plurality of data processors includes a configuration instruction register; wherein the instruction register is arranged to receive a configuration instruction during a configuration process and store the configuration instruction in the configuration instruction register; wherein the data processors of a given row are controlled by a given shared microcontroller; wherein each data processor in the given row is arranged to receive selection information for selecting a selected configuration instruction from the given shared microcontroller, and configure the data processor to operate according to the selected configuration instruction under specific conditions.

19. The data processing module according to claim 18, wherein: The specific condition is satisfied when the data processor is arranged to respond to the selection information; wherein the specific condition is not satisfied when the data processor is arranged to ignore the selection information.

20. The data processing module according to claim 1, wherein: Each of the plurality of data processors comprises a controller, an arithmetic logic unit, a register file and a configuration instruction register; wherein the instruction register is arranged to receive a configuration instruction during a configuration process and store the configuration instruction in the configuration instruction register; wherein the controller is arranged to receive selection information for selecting a selected configuration instruction and configure the data processor to operate according to the selected configuration instruction.

21. The data processing module according to claim 20, wherein: Each data processor of the plurality of data processors includes up to three configuration instruction registers.

22. A method for operating a processing module comprising a data processor array; wherein: The operations include processing data by data processors in the array; wherein each data processor unit in a plurality of data processors in the data processor array is directly coupled to one or more data processors in the data processor array, is indirectly coupled to one or more other data processors in the data processor array, and relays data between relay ports of the data processors using one or more relay channels of one or more data processors, Wherein, the relay channel of each data processing unit includes: (i) a first path consisting of a single multiplexer, (ii) a second path consisting of a multiplexer and a flip-flop; (iii) a third path consisting of two multiplexers, and (iv) The fourth path consists of two multiplexers and a flip-flop.

23. The method according to claim 22, wherein The first path, the second path, the third path, and the fourth path of the relay channel of each of the plurality of data processors are coupled in parallel.

24. The method according to claim 22, wherein Each data processor of the plurality of data processors comprises a core; wherein the core includes an arithmetic logic unit and memory resources; The cores of the plurality of data processors are coupled to each other via a configurable network.

25. The method according to claim 24, wherein Each data processor of the plurality of data processors includes a plurality of data flow components of the configurable network.

26. The method according to claim 22, wherein Each data processor of the plurality of data processors includes a first non-rebroadcast input port directly coupled to a first set of neighbors.

27. The method according to claim 26, wherein The first set of neighbors is formed by data processors located within a distance of less than four data processors from the data processor.

28. The method according to claim 26, wherein The first non-rebroadcast input port of the data processor is directly coupled to the rebroadcast ports of the data processors of the first set of neighbors.

29. The method according to claim 28, wherein The data processor also includes a second non-rebroadcast input port directly coupled to the non-rebroadcast ports of the data processors of the first set of neighbors.

30. The method of claim 26, wherein: The first non-rebroadcast input port of the data processor is directly coupled to the non-rebroadcast ports of the data processors of the first set of neighbors.

31. The method of claim 26, wherein: The first group of neighbors consists of eight data processors.

32. The method of claim 26, wherein: The first repeating port of each data processor of the plurality of data processors is directly coupled to a second set of neighbors.

33. The method according to claim 32, wherein For each data processor of the plurality of data processors, the second set of neighbors is different from the first set of neighbors.

34. The method according to claim 33, wherein For each data processor of the plurality of data processors, the second set of neighbors includes data processing units that are farther away from the data processor than any data processor belonging to the first set of neighbors.

35. The method of claim 22, wherein: In addition to the plurality of data processors, the processor array further comprises at least one other data processor.

36. The method of claim 22, wherein: The data processors within the data processor array are arranged in rows and columns.

37. The method according to claim 36, wherein One or more data processors in each row are coupled to each other in a cyclic manner.

38. The method of claim 36, comprising controlling the data processors in each row by a shared microcontroller.

39. The method according to claim 36, wherein Each data processor of the plurality of data processors includes a configuration instruction register; wherein the method comprises: Configuration instructions are received via the instruction register during the configuration process Storing the configuration instruction in the configuration instruction register; Controlling a given row of data processors via a given shared microcontroller; receiving, by each data processor in the given row, selection information for selecting the selected configuration instruction from the given shared microcontroller; and The data processor is configured under certain conditions to operate according to the selected configuration instructions.

40. The method of claim 39, wherein The specific condition is satisfied when the data processor is arranged to respond to the selection information; wherein the specific condition is not satisfied when the data processor is arranged to ignore the selection information.

41. The method of claim 22, wherein: Each of the plurality of data processors includes a controller, an arithmetic logic unit, a register file, and a configuration instruction register; wherein the method comprises: receiving configuration instructions via the instruction register during a configuration process; Storing the configuration instruction in the configuration instruction register; receiving, by the controller, selection information for selecting the selected configuration instruction; and The data processor is configured to operate according to the selected configuration instructions.

42. The method of claim 22, wherein: Each data processor of the plurality of data processors includes up to three configuration instruction registers.

43. The method according to claim 22, wherein: A data processor is indirectly coupled to another data processor when the data processor is coupled to the other data processor through one or more intermediary data processors.

44. A non-transitory computer-readable medium storing instructions that, when executed by a processing module, are configured to cause the processing module to perform a method comprising: processing data by a data processor in a data processor array of the processing module; wherein each data processor unit in a plurality of data processors in the data processor array is directly coupled to one or more data processors in the data processor array and is indirectly coupled to one or more other data processors in the data processor array; and relaying data between relay ports of one or more data processors using one or more relay channels of said data processors, The relay channel of each data processing unit includes: (i) a first path consisting of a single multiplexer, (ii) a second path consisting of a multiplexer and a flip-flop, (iii) a third path consisting of two multiplexers, and (iv) a fourth path consisting of two multiplexers and a flip-flop.

45. The non-transitory computer readable medium of claim 44, wherein: Each data processor in the data processor array is arranged in rows and columns.

46. ​​The non-transitory computer readable medium of claim 45, storing instructions for controlling the data processors in each row by a shared microcontroller.

47. The non-transitory computer readable medium of claim 44, wherein: Each of the plurality of data processors includes a configuration instruction register; wherein the non-transitory computer-readable medium stores instructions for: receiving, by the instruction register, configuration instructions during a configuration process; Storing the configuration instruction in the configuration instruction register; Controlling a given row of data processors via a given shared microcontroller; receiving, by each data processor in the given row, selection information for selecting the selected configuration instruction from the given shared microcontroller; and The data processor is configured under certain conditions to operate according to the selected configuration instructions.

48. The non-transitory computer readable medium of claim 47, wherein: The specific condition is satisfied when the data processor is arranged to respond to the selection information; wherein the specific condition is not satisfied when the data processor is arranged to ignore the selection information.

49. The non-transitory computer readable medium of claim 48, wherein: Each of the plurality of data processors includes a controller, an arithmetic logic unit, a register file, and a configuration instruction register; wherein the non-transitory computer-readable medium stores instructions for: receiving configuration instructions during a configuration process via the instruction register; Storing the configuration instruction in the configuration instruction register; receiving, by the controller, selection information for selecting the selected configuration instruction; and The data processor is configured to operate according to the selected configuration instructions.

Citation Information

Patent Citations

  • System and method for generating communications arrangements for routing data in a massively parallel processing system

    US5247694A

  • Low-overhead operating systems

    US8327187B1

Cited By

  • Tracking near-identical memory addresses and reducing memory access requests

    US12657156B2