Image processors and methods for processing images
By utilizing image processors and methods to compute distortion and parallax results using processing unit arrays and data processor arrays, the need for high-throughput, small-area image processors is addressed, image processing efficiency is improved, and image processing for driver assistance systems and autonomous vehicles is supported.
Patent Information
- Application Number
- CN202210482955.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-02-11
- Filing Date
- 2016-06-09
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2036-06-09
AI Technical Summary
Existing technologies struggle to provide high-throughput, small-area image processors to support the image processing needs of driver assistance systems and autonomous vehicles.
An image processor and method are employed to calculate the warping result through a processing unit array, including receiving weights and neighboring source pixel values, calculating the warping result, and providing the result in a memory module. The absolute difference and disparity are calculated using a data processor array, and data access is optimized using a collection unit and a comparator array.
It achieves efficient image processing, improves the throughput and area utilization of the image processor, and supports the image processing needs of driver assistance systems and autonomous vehicles.
Smart Images

Figure CN115100019B_ABST
Abstract
Description
[0001] This application is a divisional application of the application filed on June 9, 2016, with application number 201680045334.8 and invention title "Image Processor and Method for Processing Images".
[0002] Cross-references to related applications
[0003] This application claims priority to the following patents: U.S. Provisional Patent Serial No. 62 / 173,389, filed June 10, 2015; U.S. Provisional Patent Serial No. 62 / 173,392, filed June 10, 2015; U.S. Provisional Patent Serial No. 62 / 290,383, filed February 2, 2016; U.S. Provisional Patent Serial No. 62 / 290,389, filed February 2, 2016; and U.S. Provisional Patent Serial No. 62 / 290,392, filed February 2, 2016. The entire contents of the following patent applications are incorporated herein by reference: U.S. Provisional Patent Serial No. 62 / 290,395, filed February 2, 2016; U.S. Provisional Patent Serial No. 62 / 290,400, filed February 2, 2016; U.S. Provisional Patent Serial No. 62 / 293,145, filed February 9, 2016; U.S. Provisional Patent Serial No. 62 / 293,147, filed February 9, 2016; and U.S. Provisional Patent No. 62 / 293,908, filed February 11, 2016. background
[0004] Over the past few years, camera-based driver assistance systems (DAS) have entered the market and are being developed rapidly for autonomous vehicles. DAS include lane departure warning (LDW), automatic high beam control (AHC), pedestrian detection, and forward collision warning (FCW). These driver assistance systems can use real-time image processing to detect multiple patches from multiple image frames captured by cameras installed in the vehicle.
[0005] There is an increasing need to provide high throughput low footprint image processors to support DAS and / or autonomous vehicles. Invention Overview
[0006] The system and method described in the claims and specification are provided.
[0007] Any combination of any subject matter in any claim may be provided.
[0008] Any method and / or combination of method steps disclosed in any of the accompanying drawings and / or in the specification may be provided.
[0009] Any combination of any unit, device, and / or component disclosed in any of the accompanying drawings and / or the specification may be provided. Non-limiting examples of such units include collection units, image processors, etc.
[0010] It may provide any combination of methods and / or method steps according to this application, any combination of any image processor and / or image processor components, any combination of any image processor with any collection unit and / or any processing module.
[0011] According to embodiments of the present invention, a method for calculating a warp result can be provided. This method may include performing a warp calculation process for each target pixel in a set of target pixels. The warp calculation process may include receiving a first weight and a second weight associated with the target pixel through a first set of processing units in a processing unit array; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating the warp result through the second set of processing units based on the values of the neighboring source pixels and a pair of weights; and providing the warp result to a memory module.
[0012] The calculation of the distortion result may include relaying the values of some neighboring source pixels between the processing units in the second group.
[0013] The calculation of the distortion result may include relaying the intermediate results calculated by the second group and the values of some neighboring source pixels between the processing units of the second group.
[0014] The calculation of the distortion result may include: calculating a first difference between a first pair of adjacent source pixels and a second difference between a second pair of adjacent source pixels through a first processing unit of the second group; providing the first difference to a second processing unit of the second group; and providing the second difference to a third processing unit of the second group.
[0015] The calculation of the distortion result may further include: calculating a first modified weight in response to the first weight through the fourth processing unit of the second group; providing the first modified weight from the fourth processing unit to the second processing unit of the second group; and calculating a first intermediate result through the second processing unit of the second group based on the first difference, the first neighboring source pixel and the first modified weight.
[0016] The calculation of the distortion result may further include: providing a second difference from the third processing unit of the second group to the sixth processing unit of the second group; providing a second neighboring source pixel from the fifth processing unit of the second group to the sixth processing unit of the second group; and calculating a second intermediate result based on the second difference, the second neighboring source pixel, and the first modified weight by the sixth processing unit of the second group.
[0017] The calculation of the distortion result may further include: providing a second intermediate result from the sixth processing unit of the second group to the seventh processing unit of the second group; providing a first intermediate result from the second processing unit of the second group to the seventh processing unit of the second group; and calculating a third intermediate result by the seventh processing unit of the second group based on the first and second intermediate results.
[0018] The calculation of the distortion result may further include: providing a third intermediate result from the seventh processing unit of the second group to the eighth processing unit of the second group; providing a second intermediate result from the sixth processing unit of the second group to the ninth processing unit of the second group; providing a second intermediate result from the ninth processing unit of the second group to the eighth processing unit of the second group; providing a second modification weight from the third processing unit of the second group to the eighth processing unit of the second group; and calculating the distortion result by the eighth processing unit of the second group based on the second intermediate result, the third intermediate result, and the second modification weight.
[0019] This method may include performing multiple twist calculations associated with subgroups of target pixels in parallel via an array.
[0020] The method may include: extracting, in parallel from a collection unit, neighboring source pixels associated with each target pixel in a subgroup of pixels; wherein the collection unit may include a set associate cache and may be arranged to access a memory module that may include multiple independently accessible storage banks.
[0021] The method may include: receiving a first twist parameter and a second twist parameter for each target pixel in a subgroup of pixels; wherein the first twist parameter and the second twist parameter may include a first weight and a second weight, as well as location information indicating the location of neighboring source pixels associated with the target pixel.
[0022] The method may include providing the collection unit with the location information of each target pixel in a subgroup of pixels.
[0023] This method may include converting location information into addresses of neighboring source pixels through a collection unit.
[0024] The method may include calculating a first twist parameter and a second twist parameter for each target pixel of a subgroup of pixels using a third set of processing units of the array; wherein the first twist parameter and the second twist parameter may include a first weight and a second weight, as well as position information indicating the position of neighboring source pixels associated with the target pixel.
[0025] This method senses the first and second weights from the third group to the first group.
[0026] The method may include providing the collection unit with the location information of each target pixel in a subgroup of pixels.
[0027] This method may include converting location information into addresses of neighboring source pixels through a collection unit.
[0028] According to one embodiment of the present invention, a method for calculating a distortion result can be provided. The method may include receiving a first weight and a second weight simultaneously for each target pixel in a subgroup of pixels through a first set of processing units in a processing unit array; simultaneously providing a collection unit with location information indicating the location of neighboring source pixels associated with each target pixel in the subgroup of pixels; simultaneously receiving neighboring source pixels associated with each target pixel in the subgroup of pixels through the array and from the collection unit; wherein different sets of the array receive neighboring source pixels associated with different target pixels in the subgroup of pixels; and simultaneously calculating distortion results associated with different target pixels through the different sets of the array.
[0029] The method may include: receiving or calculating a first twist parameter and a second twist parameter for each target pixel in a subgroup of pixels; wherein the first twist parameter and the second twist parameter may include a first weight and a second weight, as well as location information indicating the location of neighboring source pixels associated with the target pixel.
[0030] According to one embodiment of the present invention, a method for calculating a distortion result can be provided, the method comprising: repeating the following steps for each subgroup of target pixels in a target pixel group: receiving neighboring source pixels associated with each target pixel in the subgroup of target pixels via a processing unit array; and calculating a distortion result for the target pixels from the subgroup of target pixels via the array; wherein the calculation may include calculating intermediate results and relaying at least some of the intermediate results between processing units in the array.
[0031] Each processing unit in the array can be directly coupled to one set of processing units in the array and indirectly coupled to another set of processing units in the array. The terms "processing unit" and "data processor" can be used interchangeably.
[0032] According to embodiments of the present invention, an image processor configured to calculate a warping result can be provided. The image processor can be configured to perform a warping calculation process for each target pixel in a group of target pixels. The warping calculation process may include receiving a first weight and a second weight associated with the target pixel through a first group of processing units in an array of processing units of the image processor; receiving values of neighboring source pixels associated with the target pixel through a second group of processing units in the array; calculating the warping result based on the values of the neighboring source pixels and a pair of weights through the second group; and providing the warping result to a memory module.
[0033] According to embodiments of the present invention, an image processor can be provided, which can be configured to calculate distortion results. The image processor may include an array of processing units, which can be configured to simultaneously receive a first weight and a second weight for each target pixel in a subgroup of pixels through a first group of processing units in the array; simultaneously provide location information indicating the location of neighboring source pixels associated with each target pixel in the subgroup of pixels to a collection unit of the image processor; simultaneously receive neighboring source pixels associated with each target pixel in the subgroup of pixels through the array and from the collection unit; wherein different groups in the array receive neighboring source pixels associated with different target pixels in the subgroup of pixels; and simultaneously calculate distortion results associated with different target pixels through different groups in the array.
[0034] According to embodiments of the present invention, an image processor may be provided, which may be configured to calculate a distortion result. The image processor may be configured to repeat the following steps for each subgroup of target pixels in a group of target pixels: receiving neighboring source pixels associated with each target pixel in the subgroup of target pixels through an array of processing units of the image processor; and calculating a distortion result for the target pixels from the subgroup of target pixels through the array; wherein the calculation may include calculating intermediate results and relaying at least some of the intermediate results between processing units in the array.
[0035] According to embodiments of the present invention, a method for calculating disparity can be provided, the method comprising calculating a set of sums of absolute differences (SADs) by a first set of data processors in a data processor array; wherein the set of SADs may be associated with a subgroup of source pixels and target pixels; wherein each SAD may be calculated based on a previously calculated SAD and based on the absolute difference between another currently calculated source pixel and a target pixel belonging to a subgroup of target pixels; and determining the best matching target pixel in the subgroup of target pixels by a second set of data processors in the array in response to the value of the set of SADs.
[0036] The given SAD in this set of SADs reflects the absolute difference between a given rectangular source pixel array and a given rectangular target pixel array; wherein the previously calculated SAD may include (a) a first previously calculated SAD that reflects the absolute difference between (i) a rectangular source pixel array that differs from the given rectangular source pixel array by a first source pixel column and a second source pixel column and (ii) a rectangular target pixel array that differs from the given rectangular target pixel array by a first target pixel column and a second target pixel column; and (b) a second previously calculated SAD that reflects the absolute difference between the first source columns.
[0037] For a given SAD, another source pixel can be the lowest source pixel in the second source pixel column, and the target pixel belonging to a subgroup of the target pixels can be the lowest target pixel in the second target pixel column.
[0038] The method may include: calculating an intermediate result by subtracting (a) a second previously calculated SAD and (b) the absolute difference between (i) a target pixel that can be located at the top of the second target pixel column and (ii) a source pixel that can be located at the top of the second source pixel column from a first previously calculated SAD; and adding the absolute difference between the lowest target pixel in the second target pixel column and the lowest source pixel in the second source pixel column to the intermediate result, to calculate a given SAD.
[0039] The method may include: for a given SAD, storing in a data processor array a first previously computed SAD, a second previously computed SAD, a target pixel that can be positioned at the top of a second target pixel column, and a source pixel that can be positioned at the top of a second source pixel column.
[0040] The calculation of a given SAD can be performed before extracting the lowest target pixel of the second target pixel column and the lowest source pixel of the second source pixel column.
[0041] The subgroup of target pixels may include target pixels that can be stored sequentially in the memory module; wherein, the calculation of a set of SADs can be performed before the subgroup of target pixels is extracted from the memory module.
[0042] Extraction of subgroups of target pixels from the memory module can be performed by a collection unit that may include a content-addressable memory cache.
[0043] The target pixel subgroup belongs to a set of target pixels that may include multiple subgroups of target pixels; wherein, the method may include repeating the following steps for each target pixel subgroup: calculating a set of SADs for each target pixel subgroup by means of a first set of processing units; and finding the best matching target pixel in the set of target pixels by means of a second set of data processors of the array in response to the value of the set of SADs for each subgroup of target pixels.
[0044] The method may include: calculating multiple sets of SADs that can be associated with multiple source pixels and multiple target pixel subgroups via a first set of data processors in a data processor array; wherein each SAD in the multiple sets of SADs can be calculated based on the absolute difference between a previously calculated SAD and a currently calculated SAD; and finding the best matching target pixel via a second set of data processors in the array and the value of the SAD that can be associated with the source pixel for the source pixel.
[0045] Multiple sets of SADs can include subgroups of SADs, and each subgroup of SADs can be associated with multiple subgroups of target pixels in multiple source pixel and multiple target pixel subgroups.
[0046] Multiple source pixels can belong to columns of a rectangular pixel array and can be adjacent to each other.
[0047] The computation of multiple sets of SADs can include the parallel computation of SADs in different subgroups of SADs.
[0048] This method may include calculating the SADs belonging to the same subgroup of SADs in a sequential manner.
[0049] Multiple source pixels can be a pair of source pixels.
[0050] Multiple source pixels can be four source pixels.
[0051] Different subgroups of SAD can be computed by different first subgroup data processors in the data processor array.
[0052] The method may include sequentially calculating the SADs belonging to the same subgroup of the SAD; and sequentially extracting target pixels associated with different SADs in the same subgroup of the SAD to a data processor array.
[0053] According to embodiments of the present invention, an image processor may be provided, which may include a data processor array and may be configured to compute disparity by computing a set of sums of absolute differences (SADs) through a first set of data processors of the data processor array; wherein the set of SADs may be associated with a subgroup of source pixels and target pixels; wherein each SAD may be computed based on a previously computed SAD and based on the currently computed absolute difference between another source pixel and a target pixel belonging to a subgroup of target pixels; and a second set of data processors of the array may determine the best matching target pixel in the subgroup of target pixels in response to the value of the set of SADs.
[0054] According to embodiments of the present invention, a collection unit may be provided, which may include an input interface that can be arranged to receive multiple requests for retrieving multiple requested data units; a cache memory that may include multiple entries can be configured to store multiple tags and multiple cached data units; wherein each tag may be associated with a cached data unit and may indicate a set of memory units in a memory module that are different from the cache memory and store cached data units; a comparator array that can be arranged to simultaneously compare multiple tags with multiple requested memory group addresses to provide comparison results; wherein each requested memory group address may indicate a set of memory units in a memory module that store a requested data unit among the multiple requested data units; a contention assessment unit; a controller that can be arranged to: (a) classify the multiple requested data units into cached data units and uncached data units that can be stored in the cache memory based on the comparison results; and (b) send information about the cached and uncached data units to the contention assessment unit; wherein the contention assessment unit may be arranged to check for the occurrence of at least one contention; and an output interface that can be arranged to request any uncached data unit from the memory module in a contention-free manner.
[0055] The comparator array can be arranged to compare multiple tags with multiple requested memory group addresses simultaneously during a single collection cell clock cycle; and the contention evaluation unit can be arranged to check for the occurrence of at least one contention during a single collection cell clock cycle.
[0056] The contention assessment unit can be arranged to re-examine the occurrence of at least one contention in response to a new tag in the cache memory.
[0057] The collection units can be arranged to operate in a pipeline manner; wherein the duration of each stage of the pipeline can be one collection unit clock cycle.
[0058] Each group of memory cells in a row of a memory bank comes from multiple independently accessible memory banks; wherein the contention assessment unit can be arranged to determine a potential contention when two uncached data cells belong to different rows in the same memory bank.
[0059] The cache memory can be a fully associative memory cache.
[0060] The collection unit may include an address translator, which may be arranged to translate location information contained in multiple requests into multiple requested memory group addresses.
[0061] Multiple requested data units may belong to an array of data units; wherein, the location information includes the coordinates of the multiple requested data units within the array of data units.
[0062] The contention assessment unit may include multiple groups of nodes; wherein each group of nodes may be arranged to assess contention between multiple requested memory group addresses and tags in multiple tags.
[0063] According to embodiments of the present invention, a method for responding to multiple requests for retrieving multiple requested data units can be provided. The method may include: receiving multiple requests for retrieving multiple requested data units through an input interface of a collection unit; storing multiple tags and multiple cached data units through a cache memory that may include multiple entries; wherein each tag may be associated with a cached data unit and may indicate a set of memory units in a memory module that are different from the cache memory and store cached data units; simultaneously comparing the multiple tags with the addresses of the multiple requested memory groups through a comparator array to provide a comparison result; wherein each requested memory group address may indicate a set of memory units in a memory module that store the requested data units among the multiple requested data units; classifying the multiple requested data units into cached data units that can be stored in the cache memory and uncached data units based on the comparison result by a controller; sending information about the cached and uncached data units to a contention assessment unit; checking for the occurrence of at least one contention through the contention assessment unit; and requesting any uncached data units from the memory module in a contention-free manner through an output interface.
[0064] According to embodiments of the present invention, a processing module may be provided, which may include a data processor array; wherein each data processor unit of a plurality of data processors in the data processor array may be directly coupled to some data processors in the data processor array, may be indirectly coupled to some other data processors in the data processor array, and may include a relay channel for relaying data between relay ports of the data processors.
[0065] Each of the multiple data processors can exhibit essentially zero latency in its relay channel.
[0066] Each of the multiple data processors may include a core; wherein the core may include an arithmetic logic unit and memory resources; wherein the cores of the multiple data processors may be coupled to each other through a configurable network.
[0067] Each of the multiple data processors may include multiple data flow components of a configurable network.
[0068] Each of the multiple data processors may include a first non-transfer input port that can be directly coupled to the first set of neighbors.
[0069] The first group of neighbors can be formed by data processors, which can be located within a distance of less than four data processors.
[0070] The first non-rebroadcast input port of the data processor can be directly coupled to the rebroadcast port of the first group of neighboring data processors.
[0071] The data processor may also include a second non-transfer input port, which can be directly coupled to the non-transfer ports of the first set of neighboring data processors.
[0072] The first non-transfer input port of the data processor can be directly coupled to the non-transfer ports of the first group of neighboring data processors.
[0073] The first group of neighbors can consist of eight data processors.
[0074] The first relay port of each of the multiple data processors can be directly coupled to the second group of neighbors.
[0075] For each of the multiple data processors, the second group of neighbors is different from the first group of neighbors.
[0076] For each of the multiple data processors, the second group of neighbors may include data processing units that are further away from the data processor than any data processor belonging to the first group of neighbors.
[0077] In addition to multiple data processors, the processor array may also include at least one other data processor.
[0078] Data processors in a data processor array can be arranged in rows and columns.
[0079] Some data processors in each row can be coupled to each other in a cyclical manner.
[0080] The data processors in each row can be controlled by a shared microcontroller.
[0081] Each of the multiple data processors may include a configuration instruction register; wherein the instruction register may be arranged to receive configuration instructions during the configuration process and to store the configuration instructions in the configuration instruction register; wherein the data processors in a given row may be controlled by a given shared microcontroller; wherein each data processor in a given row may be arranged to receive selection information for selecting a selected configuration instruction from the given shared microcontroller, and to configure the data processor to operate according to the selected configuration instruction under specific conditions.
[0082] The specific condition can be satisfied when the data processor is configured to respond to selection information; wherein the specific condition may not be satisfied when the data processor is configured to ignore selection information.
[0083] Each of the multiple data processors may include a controller, an arithmetic logic unit, a register file, and a configuration instruction register; wherein the instruction register may be arranged to receive configuration instructions during the configuration process and store the configuration instructions in the configuration instruction register; wherein the controller may be arranged to receive selection information for selecting a selected configuration instruction and configure the data processor to operate according to the selected configuration instruction.
[0084] Each of the multiple data processors may include up to three configuration instruction registers.
[0085] According to embodiments of the present invention, a method for operating a processing module that may include a data processor array can be provided; wherein the operation may include processing data through data processors in the array; wherein each data processor unit of a plurality of data processors in the data processor array may be directly coupled to some data processors in the data processor array, may be indirectly coupled to some other data processors in the data processor array, and may relay data between relay ports of the processing data processors using one or more relay channels of one or more data processors.
[0086] According to embodiments of the present invention, an image processor may be provided, which may include a data processor array, a first microcontroller, a buffer unit, and a second microcontroller; wherein the data processor in the array may be arranged to receive data processor configuration instructions during a data processor configuration process; wherein the buffer unit may be arranged to receive buffer unit configuration instructions during a buffer unit configuration process; wherein the first microcontroller may be arranged to control the operation of the data processor by providing data processor selection information to the data processor; wherein the data processor may be arranged to select a selected data processor configuration instruction in response to the data processor selection information and perform one or more data processing operations according to the selected data processor configuration instruction; wherein the second microcontroller may be arranged to control the operation of the buffer unit by providing buffer unit selection information to the buffer unit; wherein the buffer unit may be arranged to select a selected buffer unit configuration instruction in response to at least a portion of the buffer unit selection information and perform one or more buffer unit operations according to the selected buffer unit configuration instruction; and wherein the size of the data processor selection information may be a portion of the size of the data processor configuration instruction.
[0087] The data processors in the array can be arranged according to data processor groups, where different data processor groups can be controlled by different first microprocessors.
[0088] A data processor group can be a row of data processors.
[0089] Data processors in the same group receive the same data processor selection information in parallel.
[0090] A buffer unit may include multiple memory resource groups; different memory resource groups may be coupled to different data processor groups.
[0091] The image processor may include a second microcontroller; wherein different second microcontrollers may be arranged to control different groups of memory resources.
[0092] Different memory resource groups can be different shift register groups.
[0093] Different shift register groups can be coupled to multiple buffer groups that can be arranged to receive data from memory modules.
[0094] Multiple buffer groups can be independent of the control of a second microcontroller.
[0095] Buffer unit selection information selects connectivity between multiple memory resource groups and multiple data processor groups.
[0096] Each data processor may include an arithmetic logic unit and a data flow component; wherein, the data processor configuration instructions define the opcode of the arithmetic logic unit and define the data flow to the arithmetic logic unit via the data flow component.
[0097] The image processor may also include a memory module that may include multiple storage units; wherein, buffer units may be arranged to retrieve data from the memory module and send the data to the data processor array.
[0098] The first microcontroller shares the program memory.
[0099] Each first microcontroller may include a control register that stores a first instruction address, multiple header instructions, and multiple loop instructions.
[0100] The image processor may include a memory module that can be coupled to a buffer unit; wherein the memory module may include a storage buffer, a load storage unit, and multiple storage banks; wherein the storage buffer may be controlled by a third microcontroller.
[0101] The memory buffer can be arranged to receive memory buffer configuration instructions during the memory buffer configuration process; wherein, the third microcontroller can be arranged to control the operation of the memory buffer by providing memory buffer selection information to the memory buffer.
[0102] The image processor may include a memory buffer that can be controlled by a third microprocessor.
[0103] The third microcontroller, the first microcontroller, and the second microcontroller can have the same structure.
[0104] According to embodiments of the present invention, an image processor may be provided, which may include a plurality of configurable circuits and a plurality of microcontrollers; wherein the plurality of configurable circuits may include memory circuits and a plurality of data processors; wherein each configurable circuit may be arranged to store up to a finite number of configuration instructions; wherein the plurality of microcontrollers may be arranged to control the plurality of configurable circuits by repeatedly providing selection information to the plurality of configurable circuits, the selection information being used to select a selected configuration instruction from the finite number of configuration instructions by each configurable circuit.
[0105] The multiple configurable circuits may include a memory module that may include multiple storage banks; and a buffer unit for exchanging data between the memory module and the data processor.
[0106] The size of the selected information should not exceed two bits.
[0107] According to embodiments of the present invention, a method for configuring an image processor that may include a plurality of configurable circuits and a plurality of microcontrollers can be provided; wherein the plurality of configurable circuits may include memory circuits and a plurality of data processors; wherein the method may include storing up to a limited number of configuration instructions in each configurable circuit; controlling the plurality of configurable circuits by repeatedly providing selection information to the plurality of configurable circuits through the plurality of microcontrollers, the selection information being used by each configurable circuit to select a selected configuration instruction from the limited number of configuration instructions.
[0108] According to embodiments of the present invention, a method for operating an image processor can be provided, the method comprising: providing an image processor, the image processor including a data processor array; a memory module including multiple storage units; a buffer unit; a collection unit; and multiple microcontrollers; controlling the data processor array via a portion of the multiple microprocessors, the memory module, and the buffer unit; retrieving data from the memory module via the buffer unit; sending data to the data processor array via the buffer unit; receiving multiple requests via the collection unit for retrieving multiple requested data units from the memory module; and sending the multiple requested data units to the data processor array via the collection unit.
[0109] According to embodiments of the present invention, a method for configuring an image processor can be provided. The image processor may include an array of data processors, a first microcontroller, a buffer unit, and an array of second microcontrollers. The method may include providing data processor configuration instructions to data processors in the array during a data processor configuration process; providing buffer unit configuration instructions to buffer units during a buffer unit configuration process; controlling the operation of the data processor by providing data processor selection information to the data processor via the first microcontroller; selecting a selected data processor configuration instruction in response to the data processor selection information, and performing one or more data processing operations according to the selected data processor configuration instruction; controlling the operation of the buffer unit by providing buffer unit selection information to the buffer unit via the second microcontroller; selecting a selected buffer unit configuration instruction in response to at least a portion of the buffer unit selection information, and performing one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information may be a portion of the size of the data processor configuration instructions.
[0110] According to embodiments of the present invention, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium storing instructions for calculating a warp result, the instructions, once executed by a processing unit array, causing the execution of the following steps: performing a warp calculation process for each target pixel in a set of target pixels, the warp calculation process including: receiving a pair of weights including a first weight and a second weight associated with the target pixel through a first set of processing units in the processing unit array; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating a warp result through the second set based on the values of the neighboring source pixels and the pair of weights; and providing the warp result to a memory module.
[0111] According to embodiments of the present invention, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium storing instructions for calculating distortion results, the instructions, once executed by a processing unit array, causing the following steps to be performed: simultaneously receiving a first weight and a second weight for each target pixel in a pixel subgroup through a first group of processing units in the processing unit array; simultaneously providing a collection unit with location information indicating the positions of neighboring source pixels associated with each target pixel in the pixel subgroup for each target pixel in the pixel subgroup; simultaneously receiving neighboring source pixels associated with each target pixel in the pixel subgroup through the array and from the collection unit; wherein different groups in the array receive neighboring source pixels associated with different target pixels in the pixel subgroup; and simultaneously calculating distortion results associated with the different target pixels through different groups of the array.
[0112] According to embodiments of the present invention, a non-transitory computer-readable medium may be provided, the non-transitory computer-readable medium storing instructions for calculating a distortion result, the instructions, once executed by a processing unit array, causing the execution of the following steps: repeating the following steps for each target pixel subgroup in a set of target pixels: receiving, via the processing unit array, neighboring source pixels associated with each target pixel in the target pixel subgroup; and calculating, via the array, a distortion result for the target pixels from the target pixel subgroup; wherein the calculation includes calculating intermediate results and relaying at least some of the intermediate results between processing units in the array.
[0113] According to embodiments of the present invention, a non-transitory computer-readable medium may be provided, the non-transitory computer-readable medium storing instructions for calculating parallax, the instructions, once executed by a data processor array, causing the following steps to be performed: calculating a set of sums of absolute differences (SADs) by a first set of data processors in the data processor array; wherein the set of SADs is associated with a source pixel and a target pixel subgroup; wherein each SAD is calculated based on a previously calculated SAD and based on the currently calculated absolute difference between another source pixel and a target pixel belonging to the target pixel subgroup; and determining the best matching target pixel in the target pixel subgroup by a second set of data processors in the array in response to the value of the set of SADs.
[0114] According to embodiments of the present invention, a non-transitory computer-readable medium may be provided, which stores instructions for responding to multiple requests for retrieving multiple requested data units, the instructions, once executed by a collection unit, causing the following steps to be performed: receiving multiple requests for retrieving multiple requested data units through an input interface of the collection unit; storing multiple tags and multiple cached data units through a cache memory comprising multiple entries; wherein each tag is associated with a cached data unit and indicates a set of memory units in a memory module that are different from the cache memory and store cached data units; simultaneously comparing the multiple tags with multiple requested memory group addresses through a comparator array to provide comparison results; wherein each requested memory group address indicates a set of memory units in a memory module that store a requested data unit among the multiple requested data units; classifying the multiple requested data units into cached data units and uncached data units stored in the cache memory based on the comparison results by a controller; sending information about cached and uncached data units to a contention assessment unit; checking for the occurrence of at least one contention by the contention assessment unit; and requesting any uncached data unit from the memory module in a contention-free manner through an output interface.
[0115] According to embodiments of the present invention, a non-transitory computer-readable medium may be provided for storing instructions for operating a processing module, which, once executed by the processing module, cause the following steps to be performed: processing data by data processors in a data processor array in the processing module; wherein each data processor unit of a plurality of data processors in the data processor array is directly coupled to some data processors in the data processor array, indirectly coupled to some other data processors in the data processor array, and relays data between relay ports of the data processors using one or more relay channels of one or more data processors.
[0116] According to embodiments of the present invention, a non-transitory computer-readable medium may be provided storing instructions for configuring an image processor comprising a plurality of configurable circuits and a plurality of microcontrollers; wherein the plurality of configurable circuits include memory circuitry and a plurality of data processors, and wherein, once executed by the image processor, the instructions cause the following steps to be performed: storing up to a finite number of configuration instructions in each configurable circuit; controlling the plurality of configurable circuits by repeatedly providing selection information to the plurality of configurable circuits through the plurality of microcontrollers, the selection information being used to select a chosen configuration instruction from the finite number of configuration instructions by each configurable circuit.
[0117] A non-transitory computer-readable medium storing instructions for operating an image processor, the image processor including a data processor array; a memory module including a plurality of memory banks; a buffer unit; a collection unit; and a plurality of microcontrollers; wherein execution by the image processor results in the following steps: sending data to the data processor array via the buffer unit; receiving, via the collection unit, a plurality of requests for retrieving a plurality of requested data units from the memory module; and sending the plurality of requested data units to the data processor array via the collection unit.
[0118] A non-transitory computer-readable medium storing instructions for configuring an image processor, the image processor including a data processor array, a first microcontroller, a buffer unit, and a second microcontroller; wherein the plurality of configurable circuits include memory circuitry and a plurality of data processors, wherein the instructions, once executed by the image processor, cause the following steps to be performed: providing data processor configuration instructions to data processors in the array during a data processor configuration process; providing buffer unit configuration instructions to buffer units during a buffer unit configuration process; controlling the operation of the data processors by providing data processor selection information to the data processors via the first microcontroller; selecting a selected data processor configuration instruction by the data processors in response to the data processor selection information, and performing one or more data processing operations according to the selected data processor configuration instruction; controlling the operation of the buffer units by providing buffer unit selection information to the buffer units via the second microcontroller; selecting a selected buffer unit configuration instruction by the buffer units in response to at least a portion of the buffer unit selection information, and performing one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information is a portion of the size of the data processor configuration instructions.
[0119] According to embodiments of the present invention, a method may be provided, comprising:
[0120] A warp calculation process is performed for each target pixel in a set of target pixels. The warp calculation process includes: receiving a pair of weights, including a first weight and a second weight associated with the target pixel, through a first set of processing units in an array of processing units; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating a warp result through the second set of processing units based on the values of the neighboring source pixels and the pair of weights; and providing the warp result to a memory module.
[0121] A set of absolute difference sums (SADs) is calculated by a first set of data processors in the array of data processors; wherein the set of SADs is associated with a source pixel and a target pixel subgroup; wherein each SAD is calculated based on a previously calculated SAD and based on the currently calculated absolute difference between another source pixel and a target pixel belonging to the target pixel subgroup; and the best matching target pixel in the target pixel subgroup is determined by a second set of data processors in the array in response to the value of the set of SADs.
[0122] According to an embodiment of the present invention, a method may be provided, comprising: (a) performing a warp calculation process for each target pixel in a set of target pixels, the warp calculation process comprising: receiving a pair of weights including a first weight and a second weight associated with the target pixel through a first set of processing units in a processing unit array; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating a warp result through the second set of processing units based on the values of the neighboring source pixels and the pair of weights; and providing the warp result to a memory module; and (b) receiving a plurality of requests for retrieving a plurality of requested data units through an input interface of a collection unit; storing a plurality of tags and a plurality of cached data units through a cache memory including a plurality of entries; wherein each tag is associated with a cached data unit and indicates storage. The memory module contains a set of memory cells that are different from the cache memory and store the cached data units; a comparator array simultaneously compares the plurality of tags with the plurality of requested memory group addresses to provide a comparison result; wherein each requested memory group address indicates a set of memory cells in the memory module that stores the requested data units among the plurality of requested data units; the controller classifies the plurality of requested data units into cached data units and uncached data units stored in the cache memory based on the comparison result; and sends information about cached and uncached data units to a contention assessment unit; the contention assessment unit checks for the occurrence of at least one contention; and requests any uncached data units from the memory module in a contention-free manner through an output interface.
[0123] According to an embodiment of the present invention, a method may be provided comprising: (a) performing a warping calculation process for each of a set of target pixels, the warping calculation process comprising: receiving a pair of weights including a first weight and a second weight associated with the target pixel through a first set of processing units in a processing unit array; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating a warping result through the second set based on the values of the neighboring source pixels and the pair of weights; and providing the warping result to a memory module; and (b) operating a processing module comprising a data processor array; wherein the operation comprises processing data through data processors in the array; wherein each data processor unit of a plurality of data processors in the data processor array is directly coupled to some data processors in the data processor array, indirectly coupled to some other data processors in the data processor array, and relaying data between relay ports of the data processors using one or more relay channels of one or more data processors.
[0124] According to an embodiment of the present invention, a method may be provided comprising: (a) performing a warping calculation process for each target pixel in a set of target pixels, the warping calculation process comprising: receiving a pair of weights through a first set of processing units in an array of processing units, the pair of weights including a first weight and a second weight associated with the target pixel; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating a warping result through the second set of processing units based on the values of the neighboring source pixels and the pair of weights; and providing the warping result to a memory module; and (b) configuring an image processor including a plurality of configurable circuits and a plurality of microcontrollers; wherein the plurality of configurable circuits include memory circuits and a plurality of data processors; wherein configuring the image processor includes storing up to a finite number of configuration instructions in each configurable circuit; controlling the plurality of configurable circuits through the plurality of microcontrollers by repeatedly providing selection information to the plurality of configurable circuits, the selection information being used to select a selected configuration instruction from the finite number of configuration instructions through each configurable circuit.
[0125] According to embodiments of the present invention, a method may be provided, comprising: (a) performing a warping calculation process for each target pixel in a set of target pixels, the warping calculation process comprising: receiving a pair of weights through a first set of processing units in a processing unit array, the pair of weights including a first weight and a second weight associated with the target pixel; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating a warping result through the second set based on the values of the neighboring source pixels and the pair of weights; and providing the warping result to a memory module; and (b) operating an image processor, wherein operating the image processor comprises providing an image processor including a data processor array; a memory module including a plurality of storage media; a buffer unit; a collection unit; and a plurality of microcontrollers; controlling the data processor array through the plurality of microcontrollers, a portion of the memory module, and the buffer unit; retrieving data from the memory module through the buffer unit; sending the data to the data processor array through the buffer unit; receiving a plurality of requests for retrieving a plurality of requested data units from the memory module through the collection unit; and sending the plurality of requested data units to the data processor array through the collection unit.
[0126] According to an embodiment of the present invention, a method may be provided comprising: (a) performing a warp calculation process for each target pixel in a set of target pixels, the warp calculation process comprising: receiving a pair of weights through a first set of processing units in an array of processing units, the pair of weights including a first weight and a second weight associated with the target pixel; receiving values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; calculating a warp result through the second set of processing units based on the values of the neighboring source pixels and the pair of weights; and providing the warp result to a memory module; and (b) configuring an image processor including an array of data processors, a first microcontroller, a buffer unit, and a second microcontroller; wherein configuring the image processor comprises: providing data processor configuration to the data processor in the array during a data processor configuration process. The process includes: setting instructions; providing buffer unit configuration instructions to the buffer unit during the buffer unit configuration process; controlling the operation of the data processor via the first microcontroller by providing data processor selection information to the data processor; selecting a selected data processor configuration instruction by the data processor in response to the data processor selection information, and performing one or more data processing operations according to the selected data processor configuration instruction; controlling the operation of the buffer unit via the second microcontroller by providing buffer unit selection information to the buffer unit; selecting a selected buffer unit configuration instruction by the buffer unit in response to at least a portion of the buffer unit selection information, and performing one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information is a portion of the size of the data processor configuration instruction. Brief description of the attached diagram
[0127] The subject matter considered to be the present invention is specifically pointed out and clearly claimed in the concluding section of the specification. However, when combined with the appendix... Figure 1 When reading this invention, one can best understand its methods and organization of operation, as well as its objects, features, and advantages, by referring to the following detailed description, wherein:
[0128] Figure 1 A method according to an embodiment of the present invention is shown;
[0129] Figure 2 An image processor according to an embodiment of the present invention is shown;
[0130] Figure 3 An image processor according to an embodiment of the present invention is shown;
[0131] Figure 4 A portion of an image processor according to an embodiment of the present invention is shown;
[0132] Figure 5A clock tree according to an embodiment of the present invention is shown;
[0133] Figure 6 A memory module according to an embodiment of the present invention is shown;
[0134] Figure 7 The mapping between the LSU of the memory module and the memory bank of the memory module is shown according to an embodiment of the present invention;
[0135] Figure 8 A storage buffer according to an embodiment of the present invention is shown;
[0136] Figure 9 and Figure 10 The instructions Row, Sel are shown according to an embodiment of the present invention;
[0137] Figure 11 A buffer unit according to an embodiment of the present invention is shown;
[0138] Figure 12 A collection unit according to an embodiment of the present invention is shown;
[0139] Figure 13 It is a timing diagram showing the process including address translation, cache hit / miss, contention, and information output;
[0140] Figure 14 This illustrates a contention assessment unit according to an embodiment of the present invention;
[0141] Figure 15 and Figure 16 A data processing unit according to an embodiment of the present invention is shown;
[0142] Figure 17 This illustrates a distortion calculation method according to an embodiment of the present invention;
[0143] Figure 18 and Figure 19 A data processor array for performing warp calculations according to an embodiment of the present invention is shown;
[0144] Figure 20 The distortion parameters output from various data processors according to embodiments of the present invention are shown;
[0145] Figure 21 A data processor group for performing warp calculations according to an embodiment of the present invention is shown;
[0146] Figure 22 This illustrates a distortion calculation method according to an embodiment of the present invention;
[0147] Figure 23 A group of processing units according to an embodiment of the present invention is shown;
[0148] Figure 24 The first subgroup of source pixels, the first subgroup of target pixels, the second subgroup of source pixels, and the second subgroup of target pixels are shown.
[0149] Figure 25 The subgroup SG(B) of the source pixels with center pixel SB is shown;
[0150] Figure 26 The corresponding subgroup TG(B) of the target pixel (not shown) with center pixel TB is shown;
[0151] Figure 27 A method according to an embodiment of the present invention is shown;
[0152] Figure 28 The illustration shows eight source pixels and thirty-two target pixels processed by DPA according to an embodiment of the present invention;
[0153] Figure 29 An array of source pixels is shown according to an embodiment of the present invention;
[0154] Figure 30 An array of target pixels is shown according to an embodiment of the present invention;
[0155] Figure 31 Multiple sets of data processors (DPUs) according to embodiments of the present invention are shown;
[0156] Figure 32 Eight groups of data processors (DPUs) are shown according to an embodiment of the present invention—each group comprising four DPUs; and
[0157] Figure 33 A distortion calculation method according to an embodiment of the present invention is shown.
[0158] Detailed description of the attached figures
[0159] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, those skilled in the art will understand that the invention can be practiced without these specific details. In other instances, well-known methods, processes, and components have not been described in detail so as not to obscure the invention.
[0160] The subject matter considered to be the present invention is specifically pointed out and clearly claimed in the concluding section of the specification. However, when combined with the appendix... Figure 1 When reading this invention, one can best understand its methods and organization of operation, as well as its purpose, features, and advantages, by referring to the following detailed description.
[0161] It should be understood that, for the sake of simplicity and clarity, the components shown in the figures are not necessarily drawn to scale. For example, for clarity, the dimensions of some components may be enlarged relative to others. Furthermore, reference figures may be repeated in the figures where deemed appropriate to indicate corresponding or similar components.
[0162] Because the embodiments shown in this invention can be largely implemented using electronic components and circuits known to those skilled in the art, the details will not be explained to a greater extent than those that must be considered as described above in order to understand and comprehend the basic concepts of the invention, and so as not to obscure or distract from the teachings of the invention.
[0163] Any references to methods in this specification should be modified as necessary to apply to systems capable of performing the method, and should also be modified as necessary to apply to non-transitory computer-readable media storing instructions that, once executed by a computer, result in the execution of the method. For example, any method steps provided in this application can be executed by a system. In this sense, a system can be an image processor, a collection unit, or any component of an image processor. A non-transitory computer-readable medium storing instructions that, once executed by a computer, result in the execution of each method provided in this application.
[0164] Any references to the system and any other components in this specification shall be modified as necessary to apply to methods executable by a memory device, and shall be modified as necessary to apply to non-transitory computer-readable media storing instructions executable by a memory device. For example, methods and / or method steps performed by the image processor described in this application may be provided.
[0165] Any reference in the specification to a non-transitory computer-readable medium shall be modified as necessary to apply to a system capable of executing instructions stored in a non-transitory computer-readable medium, and shall be appropriately modified to apply to a method executable by a computer that reads instructions stored in a non-transitory computer-readable medium.
[0166] Any combination of any drawings, any part of the specification, and / or any module or unit listed in any claim may be provided. In particular, any combination of any claimed features may be provided.
[0167] A pixel can be an image element obtained by a camera or a processed image element.
[0168] The terms “row” and “line” are used interchangeably.
[0169] The term "car" is used as a non-restrictive example of a vehicle.
[0170] For the sake of simplicity, some of the figures and some of the following text include numerical examples (e.g., bus width, number of memory rows, number of registers, register length, data unit size, instruction size, number of components per unit or module, number of microprocessors, number of data processors per row and / or per column in the array). Each numerical example is merely a non-limiting example.
[0171] Figure 1 A method 90 according to an embodiment of the present invention is shown.
[0172] System 90 could be a DAS, part of an autonomous vehicle control module, etc.
[0173] System 90 can be installed in vehicle 10. At least some components of system 90 are located within the vehicle.
[0174] System 90 may include a first camera 81, a first processor 83, a storage unit 85, a human-machine interface 86, and an image processor 100. These components may be coupled to each other via a bus or network 82 or through any other arrangement.
[0175] System 90 may include an additional camera and / or an additional processor and / or an additional image processor.
[0176] The first processor 83 can determine which task the image processor 100 should perform and instruct the image processor 100 to operate accordingly.
[0177] It should be noted that the image processor 100 may be part of the first processor 83, and it may be part of any other system.
[0178] Human-machine interface 86 may include a display, speaker, one or more light-emitting diodes, microphone, or any other type of human-machine interface. The human-machine interface can communicate with the driver's mobile device, the vehicle's multimedia system, etc.
[0179] Figure 2 An image processor 100 according to an embodiment of the present invention is shown.
[0180] The main port 101 and the slave port 103 provide interfaces between the image processor 100 of the system 90 and any other components.
[0181] Image processor 100 includes:
[0182] 1) Direct memory access (DMA) for accessing external memory resources such as memory cell 85.
[0183] 2) Controller, such as, but not limited to, scalar unit 104.
[0184] 3) Scalar Unit (SU) Program Memory 106.
[0185] 4) Scalar Unit (SU) Data Memory 108.
[0186] 5) Memory module (MM) 200.
[0187] 6) MM control unit 290.
[0188] 7) Collection Unit (GU) 300.
[0189] 8) Buffer Unit (BU) 400.
[0190] 9) BU control unit 490.
[0191] 10) Data Processing Array (DPA) 500.
[0192] 11) DPA control unit 590.
[0193] 12) Configure bus 130.
[0194] 13) Multiplexers and buffers 110, 112, 114, 116, 118, 120.
[0195] 14) Buses 132, 133, 134, 135, 136 and 137.
[0196] 15) PMA status and configuration buffer 109.
[0197] The image processor 100 also includes multiple microcontrollers. For simplicity, these microcontrollers... Figure 4 As shown in the image.
[0198] DMA 102 is coupled to multiplexers 112, 114, and 120. Scalar unit 104 is coupled to buffer 118 and multiplexer 112. Buffer 118 is coupled to multiplexer 116. Buffer 110 is coupled to multiplexers 112, 116, and 114. Multiplexer 112 is coupled to SU program memory 106. Multiplexer 114 is coupled to SU data memory 108.
[0199] Memory cell 200 is coupled to collection cell 300 via (unidirectional) bus 132, to buffer cell 400 via (unidirectional) bus 134, and to DPA 500 via (unidirectional) bus 133.
[0200] Collection unit 300 is coupled to buffer unit 400 via (unidirectional) bus 135 and to DPA 500 via (unidirectional) bus 137. Buffer unit 400 is coupled to DPA 500 via (unidirectional) bus 136.
[0201] Image processor units can be coupled to each other via other buses, via an additional few buses, via interconnects and / or networks, via other buses of different widths and orientations.
[0202] It is important to note that Figure 2 The collection unit 300, buffer unit 400, DPA 500, memory module 200, scalar unit 104, SU program memory 106, SU data memory 108 and any other multiplexers and / or buffers may be coupled to each other in other ways, via additional and / or other common buses, networks, grids, etc.
[0203] Scalar unit 104 can control other components of image processor 100 to perform tasks. Scalar unit 104 can receive (e.g., from...) Figure 1 The first processor 83 executes instructions for which tasks and can retrieve relevant instructions from the SU program memory 106.
[0204] Scalar unit 104 can determine which programs the microcontrollers within SB control unit 290, BU control unit 490, and DPA control unit 590 will execute.
[0205] The program executed by the microcontroller of the SB control unit 290 controls the storage buffer (not shown) of the memory module 200. The program executed by the microcontroller of the BU control unit 490 controls the buffer unit 400. The program executed by the microcontroller of the DPA control unit 590 controls the data processing unit of the DPA 500.
[0206] Any microcontroller can control any module or unit by providing short selection information (e.g., 2-3 bits, less than one byte, or any number of bits smaller than the number of bits of the selected configuration instruction) for selecting a configuration instruction already stored in the controlled module or unit. This allows for reduced throughput and rapid configuration changes (since configuration changes may require selection between different configuration registers already stored in the relevant unit or module).
[0207] It should be noted that the number of control units and their distribution among the components of the image processor may differ. Figure 2 Those shown.
[0208] Memory module 200 is the highest-level memory resource of image processor 100. Buffer unit 400 and collection unit 300 are lower-level memory resources of image processor 100 and can be configured to retrieve data from memory module 200 and provide the data to DPA 500. DPA 500 can send data directly to memory module 200.
[0209] The DPA 500 includes multiple data processors arranged to perform computational tasks, such as, but not limited to, image processing algorithms. Non-limiting examples of image processing algorithms include warp algorithms, parallax algorithms, etc.
[0210] Collection unit 300 includes a cache memory. Collection unit 300 is configured to receive requests from DPA 500 to retrieve multiple data units (such as pixels) from the cache memory or from memory cells and extract the requested pixel. Collection unit 300 can operate in a pipelined manner and has a limited number (e.g., three) of pipeline stages with very low latency (e.g., one (or less than five or ten) clock cycles). As shown below, the collection unit can also extract data units in an additional mode while using the address generator of the memory module to extract information.
[0211] Buffer unit 400 is configured to act as a buffer for data between memory module 200 and DPA 500. Buffer unit 400 can be arranged to provide data to multiple data processors of DPA 500 in parallel.
[0212] Configuration bus 130 is coupled to DMA 102, memory module 200, collection unit 300, buffer unit 400 and DPA500.
[0213] The DPA 500 demonstrates an architecture that supports both parallel and pipelined implementations. It exhibits flexible connectivity, enabling virtually every data processing unit (DPU) to be connected to every other DPU.
[0214] The unit of the image processor 100 is controlled by a compact microprocessor that can execute zero-latency loops and can implement nested loops.
[0215] Figure 3 An image processor 100 according to an embodiment of the present invention is shown.
[0216] Figure 3 Non-limiting examples of various bus widths and the contents of memory module 200, collection unit 300, buffer unit 400, and DPA500 are provided.
[0217] DPU 500 may include 6 rows and 6 columns of data processing units (DPUs) 510(0,0)-510(5,15).
[0218] The configuration bus 130 is 32 bytes wide.
[0219] Bus 132 is 8×64 bytes wide.
[0220] Bus 134 is 6×128 bytes wide.
[0221] Bus 135 is 2 × 128 bytes wide.
[0222] Bus 137 is 2×16×16 bytes wide.
[0223] Bus 133 is 6×2×16×16 bytes wide.
[0224] Bus 136 is 2×16×16 bytes wide.
[0225] The memory module 200 is shown as including an address generator, six load memory cells, 16 multi-port memory interfaces, and 16 independently accessible memory banks with 8-byte lines.
[0226] The collection unit 300 includes a cache memory comprising 18 registers, each 8 bytes in size.
[0227] The buffer unit 400 includes a 6-row, 4-column, 16-byte register and a 6-row, 16-column, 2:1 multiplexer.
[0228] Figure 4 A portion of an image processor 100 according to an embodiment of the present invention is shown.
[0229] The two memory buffers of memory module 200 can be controlled by SB control unit 290. SB control unit 290 may include SB program memory 292 and SB microcontrollers 291 and 292. SB program memory 293 stores instructions to be executed by SB microcontrollers 291 and 292. SB microcontrollers 291 and 292 can be fed information (stored in configuration register 298) (via configuration bus 130 and / or via scalar unit 104) indicating which instructions (among the instructions stored in SB program memory 293) should be executed.
[0230] Different register lines of the buffer unit 400 can be controlled via the BU control unit 490. The BU control unit 490 may include the BU program memory 497, the configuration register 498, and the BU microcontrollers 491-496.
[0231] The BU program memory 297 stores instructions to be executed by the BU microcontrollers 491-496. The BU microcontrollers 491-496 can be fed information (stored in the configuration register 498) indicating which instructions (of which instructions are to be executed) via the configuration bus 130 and / or via the scalar unit 104.
[0232] Different rows of DPUs in the DPA 500 can be controlled by the DPA control unit 590. The DPA control unit 590 may include a DPA program memory 597, a configuration register 598, and DPA microcontrollers 591-596.
[0233] DPA program memory 297 stores instructions to be executed by DPA microcontrollers 591-596. DPA microcontrollers 591-596 can be fed information indicating which instructions (of which instructions are stored in DPA program memory 597) to execute (via configuration bus 130 and / or via scalar unit 104).
[0234] It is important to note that microcontrollers can be grouped in other ways. For example, there can be one microprocessor group, two, three, or more microprocessor groups.
[0235] Figure 5 A clock tree according to an embodiment of the present invention is shown.
[0236] Input clock signal 2131 is fed to scalar unit 104. The scalar unit sends clk_mem 2132 to the memory banks 610-625 of the memory module, and clk 2133 to the buffer unit 400, collection unit 300, and load memory units (LSUs) 630-635 of the memory module 200. clk 2133 is converted into dpa_clk 2134 and sent to DPA 500.
[0237] Memory module
[0238] Figure 6 A memory module 200 according to an embodiment of the present invention is shown. Figure 7 This illustrates the mapping between the LSU of a memory module and the memory bank of the memory module according to an embodiment of the present invention.
[0239] The memory module 200 includes 16 independently accessible memory banks M0-M15 610-625, 6 load memory cells LSU0-LSU5 630-625, a size address generator AG0-AG5 640-645, and two memory buffers 650 and 660.
[0240] Memory banks M0-M15 610-625 are eight bytes wide (64 bits per row) and contain 1K rows to provide a total storage capacity of 96KB. Each memory bank may include (or may be coupled to) a multi-port memory interface for arbitrating requests sent to the memory bank.
[0241] exist Figure 6In this configuration, four clients are coupled to each storage unit (four arrows), and the multi-port storage interface must arbitrate the access requests that appear on these four inputs.
[0242] Multi-port memory interfaces can employ any arbitration scheme. For example, priority-based arbitration can be used.
[0243] Each LSU can select one of six addresses from the address generator and connect to four memory banks. Each access can access 16 bytes (from two memory banks), allowing the six LSUs to access 12 of the 16 memory banks at a time.
[0244] Figure 7 The mapping between control signals SysMemMap with different values is shown, as well as the mapping between LSU 630-635 and memory banks M0-M15 610-625.
[0245] Figure 6 The memory module 200 is shown to output data units to the collection unit via eight eight-byte wide buses (parts of bus 132) and to the buffer unit via six sixteen-byte wide buses (parts of bus 134).
[0246] Each address generator in AGO-AG5 640-645 can implement a four-dimensional (4D) iterator using the following variables and registers:
[0247] Baddr defines the base address in the memory.
[0248] The 'W' direction - variable wDepth defines the byte distance in the W direction with one step. Variable wCount defines the maximum value of the W counter - when this value is reached, zCounter is incremented and wCounter is cleared.
[0249] The 'Z' direction - zArea defines the byte distance in the Z direction by one step, and the variable zCount defines the maximum value of the Z counter. When this value is reached, the X counter increments and the Z counter is cleared.
[0250] The 'X' direction - variable xStep defines the step size (can be 1, 2, 4, 8, or 16 bytes). The variable xCount defines the maximum value of the X counter before the next 'Y'.
[0251] The 'Y' direction – the variable stride – defines the byte distance between the starting points of consecutive "rows". The variable yCount defines the maximum value of the Y counter.
[0252] A stopping condition is generated when all counters reach their maximum values.
[0253] The generated address is: Addr = BAddr + wCount * wDepth + xcounter * xstep + ycounter * Stride + zcounter * Area.
[0254] The variables are stored in registers that can be configured via the configuration bus 130.
[0255] Below is an example of an address generator configuration mapping:
[0256]
[0257] Storage can also be written in sizes other than 2, 4, 8, and 16 bytes.
[0258] Each LSU can perform load / store operations between two of the four connected memory banks (16 bytes each).
[0259] The accessed memory bank can depend on the chosen mapping (see, for example, see...). Figure 7 ) and address.
[0260] The data to be stored is prepared in one of the storage buffers 650 and 660 described below.
[0261] Each LSU can choose an address generated from one of the six address generators AG0-AG5.
[0262] Loading operation
[0263] Data read from the memory bank is stored in a buffer (not shown) in the loading memory unit and then (via bus 134) transmitted to buffer unit 400. This buffer helps avoid pauses caused by contention on the memory bank.
[0264] Storage operations
[0265] Data to be stored in the memory is prepared in one of storage buffers 650 and 660. There are two storage buffers (storage buffer 0 650 and storage buffer 1 660). Each storage buffer can request to write one of the LSUs between one and four 16-byte words.
[0266] Each LSU can therefore receive up to 8 simultaneous requests, which are granted one after another in a predetermined order: 1) Buffer 0 - word 0, 2) Buffer 0 - word 1, ... 4) Buffer 1 - word 0, ... 8) Buffer 1 - word 3.
[0267] When a storage buffer is configured to operate in conditional storage mode, the storage buffer can either ignore (not send to the storage) or process (send to the storage) data.
[0268] When a storage buffer is configured to operate in distributed mode, a portion of the data units received by it can be regarded as addresses associated with the storage of the remaining data units.
[0269] LSU execution priority. Storage operations take precedence over load operations, ensuring that storage will not pause due to contention with loads. Because load operations use buffers, contention is typically swallowed, preventing pauses.
[0270] storage buffer
[0271] Storage buffers 650 and 660 are controlled by storage buffer microcontrollers 291 and 292.
[0272] During the configuration process, each of the storage buffers 650 and 660 receives (and stores) three configuration instructions (sb_instr[1]-sb_instr[3]). The configuration instructions (also known as storage buffer configuration instructions) of different storage buffers can be different from each other or can be the same.
[0273] During configuration, each memory buffer microcontroller receives the address of the instruction to be executed by that microcontroller. The first and last program counters (PCs) indicate the first and last instructions to be read from the memory buffer program memory 293. The location of the program memory for each memory buffer microcontroller is also defined in the following configuration example:
[0274] Storage buffer configuration
[0275]
[0276] Storage buffer microcontroller configuration mapping:
[0277]
[0278] The microcontroller instructions for the storage buffer can be either execute instructions or do-loop instructions. They have the following formats:
[0279] Execute command:
[0280] [8:0]RLC: Repeat Instruction Counter: 1...495: Immediately -496...511:
[0281] Indirect Counter Register
[0282] [10:9]sel: Store buffer instruction selection. The trigger stores when it is non-zero.
[0283] [13:11] Row: DPA row selection
[0284]
[14] Retain
[0285]
[15] Type 0
[0286] Do loop:
[0287] [8:0]RLC: Repeating Cycle Counter: 0...495: Immediately -496...511:
[0288] Indirect Counter Register
[0289] [13:9] Length: Cycle length: 1...16
[0290]
[14] Mode: 0: counting cycle, 1: counting period
[0291]
[15] Type 1
[0292] Figure 8 A storage buffer 660 according to an embodiment of the present invention is shown.
[0293] The storage buffer 660 has four multiplexers 661-664, four buffer words 0-3 671-674, and four demultiplexers 681-684.
[0294] Buffer word 0 671 is coupled between multiplexer 661 and demultiplexer 681. Buffer word 1 672 is coupled between multiplexer 662 and demultiplexer 682. Buffer word 2 673 is coupled between multiplexer 663 and demultiplexer 683. Buffer word 3 674 is coupled between multiplexer 664 and demultiplexer 684.
[0295] Each of the multiplexers 661-664 has four inputs for receiving different lines of bus 133 and is controlled by the control signals Row, Sel.
[0296] Each of the multiplexers 681-684 has six outputs for providing data to any one of LSU0-LSU5 and is controlled by control signals En, LSU.
[0297] Storage buffer configuration instructions control the operation of storage buffers and can even generate commands Row, Sel, and En, LSU.
[0298] The following is an example of a configuration command format:
[0299] Storage buffer instruction encoding
[0300] [2:0] lsu0 writes word 0 via LSU#lsu0
[0301] [3] En0 Write word 0
[0302] [4] Conditional
[0303] [5] Dispersed
[0304] [8:6] lsu 1 writes word 1 via LSU#lsu 1
[0305] [9] Enl writes word 1
[0306]
[10] Conditional
[0307]
[11] Dispersed
[0308] [14:12] lsu2 writes word 2 via LSU#lsu2
[0309]
[15] En2 Write word 2
[0310]
[16] Conditional
[0311]
[17] Dispersed
[0312] [20:18] lsu 3 Write word 3 via LSU#lsu3
[0313]
[21] En3 Write word 3
[0314]
[22] Conditional
[0315]
[23] Dispersed
[0316] [28:24]sel Data Selection
[0317] The five bits called "data select" are actually the instruction Row, Sel, and... Figure 9 and Figure 10 The values between 0 and 28 are shown in the diagram. They are mapped to different ports of the DPA 500's DPU. Figure 9 In the diagram, D# and E# represent the outputs D and E of DPU[row,#], respectively, where 'row' is the row selection of the currently selected instruction, and D'# is the output D of DPU[row+1,#].
[0318] Buffer unit
[0319] Figure 11 A buffer unit 400 according to an embodiment of the present invention is shown.
[0320] The buffer unit 400 includes a read buffer (RB) uniformly labeled 402, a register file (RF) 404, an internal network of the buffer unit 408, multiplexer control circuits 471-476, an output multiplexer 406, BU configuration registers 401(1)-401(5) (each storing two configuration instructions), a history configuration buffer 405, and a BU read buffer configuration register 403.
[0321] The BU microcontroller can select which configuration instruction to read for each row (from the two configuration instructions stored in each BU configuration register in 401(1)-401(5)).
[0322] There are six line multiplexers, and they include multiplexers 491(0)-491(15) and 491'(0)-491'(15), multiplexers 492(0)-492(15) and 492'(0)-492'(15), multiplexers 493(0)-493(15) and 493'(0)-493'(15), multiplexers 494(0)-494(15) and 494'(0)-494'(15), and multiplexers 495(0)-495(15) and 495'(0)-495'(15).
[0323] To simplify the explanation, Figure 11 Only the multiplexer control circuits 471 and 476 are shown.
[0324] The internal network 408 of the buffer unit couples the read buffer 402 to the register file 404.
[0325] The four read buffers 415, 416, 417 and 417 of the first row (via the internal network 408 of the buffer cell) are coupled to the four registers R3 413, R2 412, R1 411 and R0 410 of the first row.
[0326] The four read buffers 425, 426, 427 and 427 in the second row (via the internal network 408 of the buffer cell) are coupled to the four registers R3 423, R2 422, R1 421 and R0 420 in the second row.
[0327] The four read buffers 435, 436, 437 and 437 in the third row (via the internal network 408 of the buffer cell) are coupled to the four registers R3 433, R2 432, R1 431 and R0 430 in the third row.
[0328] The four read buffers 445, 446, 447 and 447 in the fourth row (via the internal network 408 of the buffer cell) are coupled to the four registers R3 443, R2 442, R1 441 and R0 440 in the fourth row.
[0329] Different lines of the register file and corresponding lines of the multiplexer are controlled by different BU microcontrollers in 491-495.
[0330] The DPU microcontrollers also control multiplexers for different rows. Specifically, each DPU microcontroller (591-596) controls its corresponding DPU row and sends control instructions (MuxCtl) to the corresponding multiplexer row (via multiplexer control circuits 471-476). Each multiplexer control circuit stores the last (e.g., sixteen) MuxCtl instructions (instruction history), and the history configuration buffer 405 stores selection information for determining which MuxCtl instruction to send to the multiplexer row.
[0331] The multiplexer control circuit 471 controls the first row of multiplexers and includes a FIFO 481(1) for storing MuxCtl instructions sent from the DPU microcontroller 491, and includes a control multiplexer 481(2) to select which stored MuxCtl instruction to retrieve from the FIFO 481(1) and send to the first row of multiplexers including multiplexers 491(0)-491(15) and 491'(0)-491'(15).
[0332] The multiplexer control circuit 476 controls the sixth row multiplexer and includes a FIFO 486(1) for storing MuxCtl instructions sent from the DPU microcontroller 496 and includes a control multiplexer 486(2) to select which stored MuxCtl instruction to retrieve from the FIFO 482(1) and send to the sixth row multiplexer including multiplexers 495(0)-495(15) and 495'(0)-495'(15).
[0333] The register file can be controlled by the BU microcontroller. Operations performed on the register file can include (a) shifting from the most significant bit to the least significant bit, where the shift jump is a power of 2 bytes (or any other value), and (b) loading one or two registers from the read buffer. The contents of the file registers can be manipulated. For example, the contents can be interleaved and / or interlaced. Some examples are provided in the instruction set provided below.
[0334] The buffer configuration map includes the address of the configuration buffer, which stores the configuration instructions for the buffer unit and an indication of which commands are retrieved by the buffer unit microcontrollers (BUuC0-BuuC5). The latter refers to the first and last PCs of the BU instruction pair (each 32-bit MuxConfig instruction includes two separate buffer unit configuration instructions):
[0335] Buffer unit configuration mapping
[0336]
[0337]
[0338] Instructions executed by the BU microcontroller include bits [0:8] for loop control or instruction repetition, and bits [9:13] containing the value controlling the execution of the instruction. This is true for both register file commands and do loop commands.
[0339] Instruction encoding:
[0340] RFL represents the register file line and RBL represents the read buffer line.
[0341] RRR[r / r1:r0] represents register number "r" (16 bytes) / registers from register r1 to t0 (RFL or RBL). RRR[r][b / b1:b0] represents register number "r" byte b / registers from bytes b1 to b0 (RFL or RBL).
[0342] RF instructions: (Each cycle, the BU microcontroller can send bits 9-14 to the BU)
[0343] [8:0]RIC: Repeat Instruction Counter: 1...495: Immediate -496...511: Indirect Counter Register
[0344] [11:9] Shift: 0: NOP, 1: 1B, 2: 2B, 3: 4B, 4: 8B, 5: 1R, 6: 2R, 7: 1L, 8-15: Res
[0345] [14:12] Loading:
[0346] 0: NOP
[0347] 1: Single-precision type A: RFL[3] <= RBL[0]
[0348] 2: Single-precision type B: RFL[2] <= RBL[0]
[0349] 3: Double precision type: RFL[2]<=RBL[0], RFL[3]<=RBL[1]
[0350] 4: Interleaved byte type: RFL[2:3]<={RBL[1]
[15] ,RBL[0]
[15] ,…,RBL[0][0]}
[0351] 5: Alternating short integers: RFL[2:3] <= {RBL[1][15:14], RBL[0][15:14], ..., RBL[1][3:2], RBL[0][3:2], RBL[1][1:0], RBL[0][1:0]}
[0352] 6: Interleaved byte type &0: RFL[2:3]<={0, RBL[0]
[15] , ..., 0, RBL[0][0]}
[0353] 7: Alternating short integers &0: RFL[2:3]<={0, RBL[0][15:14], ..., 0, RBL[0][1:0]}
[0354]
[15] Type: 0
[0355] Do loop:
[0356] [8:0]RLC: Repeating Cyclic Counter: 0...495: Immediate -496...511: Indirect Counter Register
[0357] [13:9] Length: Cycle length: 1...16
[0358]
[14] Mode: 0: counting cycle, 1: counting period
[0359]
[15] Type 1
[0360] Load from the buffer.
[0361] The RB load operation from the LSU is configuration-controlled and self-triggered by its state and the state of the associated LSU. The BU read buffer configuration register (also known as RBSrcCnf) 403 specifies which LSU each RB line is loaded from.
[0362] The configuration instructions stored in the BU read buffer configuration register 403 have the following format:
[0363] [3:0] Read buffer 0LSU source:
[0364] [7:4] Read buffer 1LSU source
[0365] [11:8] Read buffer 2LSU source
[0366] [15:12] Read buffer 3LSU source
[0367] [19:16] Read buffer 4LSU source
[0368] [23:20] Read buffer 5LSU source
[0369] "Read buffer" refers to the rows of the read buffer. The four bits of each read buffer row can have the following meanings: 0-7: no load, 8: LSU0, 9: LSU1, 10: LSU2, 11: LSU3, 12: LSU4, 13: LSU5, 14: GU0, 15: GU1 (GUI swaps 8 bytes loaded as d8..d15, d0..d7 replaces d0..d15 or the last 8 short integers of GU in short integer mode are output).
[0370] Multiplexer configuration
[0371] The multiplexer is configured as follows (this example involves the first line, and my directive MuxCtl is represented as MuxCt10):
[0372] MuxCtl0[Row][4:0]: BRfpSelB: / / [0-7]: BSelB0, 8-19: {RSelB0, BSelB0}, 31: fpB0,
[0373] MuxCtl0[Row][7:5]:BSelAb / / Reg[4][BSelAb*2+31]
[0374] MuxCtl0[Row][8]: fpA0 / / floating-point mode of MuxA(col<<2)
[0375] MuxCtl0[Row][11:9]:BSelBb / / Reg[5][BSelAb*2+31]
[0376] MuxCtl0[Row][15:12]: Reserved
[0377] MuxCtl operates Muxes by selecting registers from the register file and ports of the DPU in the following ways:
[0378] PxlA[Row,Col]=! fpA? Reg[Row][Col*2]:REG[Row][Col*4]
[0379] PxlFpB[Row, Col]=Reg[Row][Col*4+2]
[0380] PxlB[Row, Col]=BSelB==0x1f?
[0381] PxlFpB[Row, Col]:
[0382] BSelB<8?
[0383] REG[Row][BSelB+Col*2]:
[0384] Reg[((BSelB-8) / 2+Row)%6][(BSelB&1)+Col*2]
[0385] PxlAb[Row, 0]=Reg[4][BSelAb*2+32]
[0386] PxlBb[Row, 0]=Reg[5][BSelBb*2+32]
[0387] The selection of history can be accomplished (by FIFO 471-476) by reading the contents of the four-bit history configuration buffer 405 that stores the selection information of the multiplexer for each line.
[0388] Collection Unit
[0389] Figure 12 A collection unit 300 according to an embodiment of the present invention is shown.
[0390] The collection unit 300 includes an input buffer 301, an address converter 302, a cache memory 303, an address tag comparator 304, a contention evaluation unit 306, a controller 307, a memory interface 308, an iterator 310, and a configuration register 311.
[0391] Collection unit 300 is configured to collect up to 16 bytes or short integer pixels from eight memory banks MB0..MB7 or MB8..MB15 via a fully associative cache memory (CAM) 303 comprising 16 8-byte registers, depending on the bit gu_Ctrl
[15] of configuration register 311. Pixel addresses or coordinates are generated or received from the array according to the pattern gu_Ctrl[3:0] of configuration register 311.
[0392] The collection unit 300 can access the memory of the memory module 200 by using the memory cell number and row address (duplets) that indicate the memory cell number and row address within the memory cell.
[0393] The collection unit can receive or generate X and Y coordinates representing the location of a requested pixel in the image, instead of an address. The collection unit includes an address converter 302 for converting the X and Y coordinates into addresses.
[0394] Iterator 310 can operate in one of two modes - (a) internal iterator only and (b) address generator using memory module.
[0395] When operating in mode (a), iterator 310 can generate sixteen addresses using the following control parameters (stored in configuration register 311):
[0396] 1) AddBase - 16 base addresses.
[0397] 2) AddStep - Iterator address step size.
[0398] 3) xCount: A counter for the maximum number of steps before the stride. The stride is performed from the previous step or the base coordinates.
[0399] 4) AddStride: stride (or y-step).
[0400] When operating in mode (b), the iterator provides AddBase to the address generator, and the address generator uses that address to perform iteration.
[0401] For example, during parallax calculation, an iterator pattern can be useful, where collection units can retrieve data units from the source and target images (especially pixels that are close to each other).
[0402] Another operating mode involves receiving the address of the requested data (the address could be the X, Y coordinates to be converted by the address converter 302 or a memory address such as a binary), checking whether the requested data unit is stored in the cache memory, and if not, retrieving the data unit from the memory module 200.
[0403] Another mode of operation involves receiving the address of a requested data unit, transforming that address into more requested addresses, and then retrieving the contents of those more requested addresses from cache memory 303 or from a memory cell. This mode of operation could be useful, for example, during warp calculations, where the collection unit can receive the address of a pixel and obtain the pixel and several other adjacent data units.
[0404] The cache memory 303 stores tags as binaries, and these tags are used to determine whether a requested data unit is within the cache memory 303. These binaries are also used to detect contention—when multiple requested data units reside in different rows of the same memory within the same cycle.
[0405] Figure 13 It is a timing diagram showing the process including address translation, cache hit / miss, contention, and information output.
[0406] Coordinates (X, Y) are accepted or generated in cycle 0 and converted into binaries in cycle 1. Memory access is computed within the same cycle to prepare for the access in cycle 2. If several coordinates address the same memory at different addresses (as is the case in this example), contention occurs, the corresponding pixel extraction is deferred to the next cycle, and is asserted to stop in cycle 1. The coordinates causing the contention exit in cycle 1 and access the memory in cycle 2. The delay between coordinates and pixels is 5 cycles plus the number of stopping cycles. In extreme cases, accessing 16 pixels can result in 15 stopping cycles. For warp operations, since the accessed pixels are close to each other, 16 pixels are extracted in an average of 1.0–1.4 cycles, depending on the type of warp.
[0407] Configuration mapping
[0408]
[0409]
[0410] Return to reference Figure 12 Input buffer 301 is coupled to address converter 302. Address tag comparator 304 receives input from cache memory 303 and address converter 302 (if address translation is required) or from input buffer 301. Address tag comparator 304 sends output signals indicating the comparison (such as cache miss, cache hit, and if a hit, where the hit occurred) to controller 307, contention evaluation unit 306, and memory interface 308. Iterator 310 is coupled to input buffer 301 and memory interface 308.
[0411] Input interfaces (such as input buffer 301) are arranged to receive multiple requests for retrieving multiple requested data units.
[0412] The cache memory 303 includes entries (such as sixteen entries or rows) for storing multiple tags (each tag may be a pair) and multiple cached data units.
[0413] Each tag is associated with a cached data unit and represents a set of memory units (such as rows) in a memory module (such as a storage bank) that are different from the cache memory and store the cached data units.
[0414] Address tag comparator 304 includes a comparator array arranged to simultaneously compare multiple tags and multiple requested memory group addresses to provide a comparison result.
[0415] The address tag comparator consists of K×J nodes – covering each tag pair and the requested memory address. The (k,j)th node 304(k,j) compares the address of the kth requested address with the jth tag.
[0416] If, for example, a requested bank / row address does not match any tag, the address tag comparator 304 will send a miss signal.
[0417] The controller 307 may be arranged to (a) classify multiple requested data units into data units stored in cache memory 303 and uncached data units (not stored in cache memory 303) based on the comparison results; and (b) send information about a data unit to the contention evaluation unit 306 when at least one uncached data unit exists.
[0418] The contention assessment unit 306 is configured to check for the occurrence of at least one contention.
[0419] The memory interface 308 is configured to request any uncached data units from the memory module in a contention-free manner.
[0420] When one or more uncached data units are retrieved by the collection unit, they are stored in the cache memory. In the cycle following the cycle in which contention is detected, the address tag comparator 304 can receive the same requested data units from the given cycle from the input buffer or from the address converter. The previously uncached data units now stored in the cache memory 303 will change the comparison result made by the address tag comparator 304, and will retrieve only the uncached memory data units that were not requested in the previous iteration from the memory module.
[0421] The contention assessment unit can include multiple groups of nodes. Figure 14 Examples are provided in the document.
[0422] The number of node groups is the maximum number of memory banks that a collection unit can access simultaneously (e.g., eight).
[0423] Each group of nodes is arranged to evaluate contention associated with a single memory bank. For example, in Figure 14 There are eight groups of nodes – used for comparisons between the sixteen requested addresses and the rows of eight memory banks.
[0424] The nodes in the first group (305(0,0)-306(0,15)) are connected in series. The leftmost node in the first group, 306(0,0), receives the input signal (used, memory 0, row) and the signal (valid 0, address 0).
[0425] Address 0 is the first of 16 requested addresses (for a data unit), and a valid 0 indicates whether the first requested address is valid – the first requested address refers to either a cached data unit (invalid) or an uncached data unit (valid).
[0426] The input signal (used, memory bank 0, row) indicates whether memory bank 0 is used and the row requested by the first node in the group. The input signal (used, memory bank 0, row) fed to the leftmost node 306 (0, 0) indicates that memory bank 0 is not used.
[0427] For example, if the leftmost node 306(0,0) (or any other node) determines that a previously unused storage bank (currently not associated with any uncached data unit) (used to retrieve the address of a valid data unit associated with the node) should be used, the node changes the signal (used, storage bank 0, row) to indicate that the storage bank is used - and also updates "row" to the row requested by the node.
[0428] If any node in the first group (of nodes 306(0,1)-306(0,15)) receives a valid address indicating memory bank 0, that node compares the row of the requested address with the row indicated by (used, memory bank 0, row). If the row values do not match, the node outputs a contention signal.
[0429] The same process can be executed simultaneously by any group of nodes.
[0430] If there are J labels, there are J groups of nodes connected in series, and each group may contain K nodes.
[0431] Therefore, each node in the group is arranged to (a) receive an access request indication (e.g., a signal (used, storage, row) indicating whether any previous node in the group has requested access to a storage bank identified by a bin (valid, address) and (b) update the access request indication to indicate whether the group is requesting access to its corresponding storage bank.
[0432] Data Processing Array (DPA) 500.
[0433] The DPA 500 includes 96 DPUs arranged in six rows, with each row containing sixteen DPUs.
[0434] Each row of the DPU can be controlled by a separate DPA microcontroller.
[0435] Figure 15 and Figure 16 A DPU 510 according to an embodiment of the present invention is shown.
[0436] The DPU 510 includes:
[0437] 1) Arithmetic Logic Unit (ALU) 540
[0438] 2) Register file 550 includes sixteen registers 550(0)-550(15).
[0439] 3) Two output multiplexers, MuxD 534 and MuxE 535.
[0440] 4) Multiple input multiplexers MuxIn0 570, MuxIn1 571, MuxIn2 572, MuxIn3 573, MuxA 561, MuxB 562, MuxCl 563, MuxCh 564, MuxF 526 and MuxG 527.
[0441] 5) Internal multiplexers MuxH 529 and MuxG 528.
[0442] 6) Triggers 565 and 566.
[0443] 7) Registers RegA 531, RegB 532, RegCl 533 and RegCh 534.
[0444] The multiplexers mentioned above are non-restrictive examples of data flow components.
[0445] Each input multiplexer is coupled to the input port of DPU 510 and can be coupled to other DPUs, to the collection unit 300, to the memory module 200, to the buffer unit 400, or to the output port of the DPU.
[0446] The input multiplexers MuxA 561, MuxB 562, MuxC1 563, MuxCh 564, and MuxH 529 also include input terminals coupled to bus 581. Bus 581 is also coupled to the output terminals of MuxIn0 570, MuxIn1 571, MuxIn2 572, and MuxIn3 573.
[0447] Registers RegA 531, RegB 532, RegCl 533 and RegCh 534 are connected between the input multiplexers MuxA 561, MuxB 562, MuxCl 563, and MuxCh 564 (one register for each multiplexer) and the ALU 540, and feed data to the ALU.
[0448] Using one set of multiplexers to receive the output of the other set increases the number of sources of data that can be fed to the ALU 540.
[0449] The output of ALU 540 is coupled to the input of register file 550. The first register file Reg0 550(0) is also connected to ALU 540 as an input.
[0450] The output of register file 550 is coupled to output multiplexers MuxD 534 and MuxE 535. The outputs of output multiplexers MuxD 534 and MuxE 535 are coupled to output ports D (521) and E (522), respectively, and (via MuxG'528) to output port G 523 and flip-flop 566.
[0451] Register RegH 539 is connected between MuxH 529 and MuxG 527. MuxF 526 is directly connected to port F522 and trigger 565, thus providing a low-latency relay channel. MuxG 527 is coupled to MuxG'528.
[0452] MuxIn0, MuxIn1, MuxIn2, and MuxIn3 can achieve short-routing to other DPUs: (a) MuxIn0 and MuxIn1 obtain input from the D output of the 8 DPUs, and (b) MuxIn2 and MuxIn3 obtain input from the E output of the same 8 DPUs.
[0453] The other five input multiplexers, MuxA, MuxB, MuxCl, MuxCh, and MuxG, can implement the following routing:
[0454] MuxA can obtain its input from the buffer unit and from MuxIn0..MuxIn3.
[0455] Each of MuxB, MuxCl, and MuxCh can obtain its input from buffer units MuxIn0..MuxIn3 and from internal registers of the register file (such as R14 or R15 of the register file).
[0456] Most ALU operations produce a short integer result (Out0). Some operations produce a word result or two short integer results ({Out1, Out0}). The output is stored at a constant location in register file 550: R(0) <= Out0, and R(1) <= Out1 (for operations that produce two short integers).
[0457] DPU 510 and other DPUs in the same row are controlled by a shared row PDA microcontroller, which generates a selection information stream for choosing between configuration instructions (see configuration register 511) stored in the DPU.
[0458] Configuration register (also known as dpu_Ctrl) 511 can store the following:
[0459]
[0460]
[0461] The four configuration registers 511(1)-511(3) are called dpu_Inst[0]-dpu_Inst[1].
[0462] As described above, each DPU of the DPA 500 is directly coupled to some DPUs of the PMA and directly coupled (via one or more intermediate DPUs) to some other data processors in the data processor array. Each DPU has a relay channel (between ports F and G) for relaying data between the DPU's relay ports (ports F and G). This simplifies and reduces connections while providing sufficient connectivity and flexibility to perform image processing tasks efficiently.
[0463] The relay channel for each of the multiple data processors (specifically the path between the output G of port F and port G) exhibits essentially zero latency. This allows the use of a PDU as a zero-latency relay channel, thereby indirectly coupling between DPUs, and also allows data to be broadcast to multiple DPUs by using relay channels between different DPUs.
[0464] See again Figure 16 Each DPU comprises a core. The core includes an ALU 540 and memory resources such as a register file 550. The cores of multiple DPUs are coupled to each other via a configurable network. The configurable network includes data flow components such as multiplexers MuxA-MuxCh, MuxD-MuxE, MuxIn0-MuxIn3, MuxF-MuxH, and MuxG'. These data flow components can be included within the DPU (e.g., ...). Figure 16 (as shown), but it can be located at least partially outside the DPU.
[0465] The DPU may include non-rebroadcast input ports directly coupled to the first group of neighbors. For example, non-rebroadcast input ports may include input ports A, B, C1, Ch, In0, In1, In2, and In3. Their connectivity to the first group of neighbors is listed in the examples below. The first group of neighbors may include, for example, eight neighbors.
[0466] The first group of neighbors is formed by DPUs located within a distance of less than four DPUs (circular distance). The distance and direction are cyclic. For example, the D ports of MuxIn0 coupled to the DPU are: (a) in the same row but one column to the left (D(0, -1)), (b) in the same row but one column to the right (D(0, +1)), (c) in the same column but in the row above (D(-1, 0)), (d) in the row above and one column to the left (D(-1, -1)), (e) in the row above and one column to the right (D(-1, +1)), (f) in the two rows above and in the same column (D(-2, 0)), (g) in the two rows above and one column to the left (D(-2, -1)), (h) in the two rows above but one column to the right (D(-2, +1)).
[0467] The first non-rebroadcast input port of a data processor can be directly coupled to the rebroadcast port of the first group of neighboring data processors. For example, see port A, which is directly coupled to the F port of the same DPU and also to the F port of the DPU in the same row but one column to the left (F / Fd(0, -1)).
[0468] The first relay port (e.g., ports G and F) can be directly coupled to the second group of neighbors. For example, the input multiplexer F (coupled to port F) can be coupled to the G output and delayed G output (Gd) of the DPU's G port, which are (a) the next row and in the same column (G / Gd(+1,0)), (b) the next two rows and in the same column G / Gd(+2,0), (c) the next three rows and in the same column (G / Gd(+3,0)), (d) the same row but one column to the left (G / Gd(0,+1)), (e) the same row but two columns to the left G / Gd(0,+2), (f) the same row but four columns to the left (G / Gd(0,+4)), and (g) the same row but eight columns to the left (G / Gd(0,+8)).
[0469] The following configuration example provides the locations of the different configuration buffers for the PMA and DPU microcontroller configuration register 598 (including registers p_FLIP0-p_FLIPS5 - one for each DPU microcontroller):
[0470]
[0471]
[0472] Configuration registers 511(0)-511(3) can store up to four configuration instructions. Configuration instructions can be 64 bits long and can be read by one or two read operations.
[0473] The configuration instructions control the selection of the multiplexer (A, B, C, D, E, F, and G), and shift the register file by 550 bits.
[0474] 1) Shift according to step 1: For n in [0,15]: For (i=15;i>0;i--)Ri<=R(i-1).
[0475] 2) Shift according to step 2: For n in [0,2,4...14]: For (i=7; i>0; i--){R(2i+1), R(2i)}<={R(2i-1), R(2i-2)}
[0476] These fields can be applied at different times:
[0477] 1) The input multiplexers A, B, C, F and the output multiplexer G are controlled without delay.
[0478] 2) The ALU controls a delay of one clock cycle.
[0479] 3) Output multiplexers D and E control delay of two clock cycles.
[0480] The following table describes the different fields of the DPU configuration command:
[0481] 4:0 Enter A to select
[0482] 9:5 Enter B to select
[0483] 14:10 Enter Cl to select
[0484] 19:15 Enter Ch to select
[0485] 23:20 Output D selection
[0486] 27:24 Output E selection
[0487] 31:28 Input / Output F Selection
[0488] 36:32 Output G selection
[0489] 41:37 RegOpCode:
[0490] 40:37 Shift destination (considered as register file):
[0491] The last change of register position in a shift operation
[0492] 41. Shifting steps:
[0493] 0: Shift by 1 (short integer result)
[0494] 1: Shift by 2 (for integer results)
[0495] 42. Interleave the ALU outputs into R0 and R1.
[0496] 44:43 wrConst: 0: NOP; 1: Regl5<=ALU low; 2: Regl5:
[0497] ALU high; 3: {Regl4, Regl5}<=ALU
[0498] 56:45 ALU Control: (AluCtl_t)
[0499] [5:0]: DPU ALU opcode (DpuAluOpCode_t)
[0500] [7:6]: Post Modifier: Rounds FP operations, and performs post-shifts for others.
[0501] Rounding mode: 0: Round, 1: Int, 2: Floor, 3: Ceil
[0502] Left shift: 0, 1, 2, or 3 bits
[0503] [8]: Vectorial: 0: Regular – 1: Vector
[0504] [9]: mode_a: 0: unsigned – 1: signed
[0505]
[10] : mode_b: 0: unsigned – 1: signed
[0506]
[11] : acc: Accumulator mode: Select Cacc and skip register C
[0507] 57: wrReg1: Write Reg1 high level from ALU.
[0508] 63:58 Reserved
[0509] Each of MuxIn0-MuxIn3 is connected to multiple buses - as shown below:
[0510] MuxIn0:
[0511] οMuxIn0[0]<=D(0,-1)
[0512] οMuxIn0[1]<=D(0,+1)
[0513] οMuxIn0[2]<=D(-1,0)
[0514] οMuxIn0[3]<=D(-1,-1)
[0515] οMuxIn0[4]<=D(-1,+1)
[0516] οMuxIn0[5]<=D(-2,0)
[0517] οMuxIn0[6]<=D(-2,-1)
[0518] οMuxIn0[7]<=D(-2,+1)
[0519] MuxIn1:
[0520] οMuxIn1[0]<=D(0,-1)
[0521] οMuxIn1[1]<=D(0,+1)
[0522] οMuxIn1[2]<=D(-1,0)
[0523] οMuxIn1[3]<=D(-1,-1)
[0524] οMuxIn1[4]<=D(-1,+1)
[0525] οMuxIn1[5]<=D(-2,0)
[0526] οMuxIn1[6]<=D(-2,-1)
[0527] οMuxIn1[7]<=D(-2,+1)
[0528] MuxIn2:
[0529] οMuxIn2[0]<=E(0,-1)
[0530] οMuxIn2[1]<=E(0,+1)
[0531] οMuxIn2[2]<=E(-1,0)
[0532] οMuxIn2[3]<=E(-1,-1)
[0533] οMuxIn2[4]<=E(-1,+1)
[0534] οMuxIn2[5]<=E(-2,0)
[0535] οMuxIn2[6]<=E(-2,-1)
[0536] οMuxIn2[7]<=E(-2,+1)
[0537] MuxIn3:
[0538] οMuxIn3[0]<=E(0,-1)
[0539] οMuxIn3[1]<=E(0,+1)
[0540] οMuxIn3[2]<=E(-1,0)
[0541] οMuxIn3[3]<=E(-1,-1)
[0542] οMuxIn3[4]<=E(-1,+1)
[0543] οMuxIn3[5]<=E(-2,0)
[0544] οMuxIn3[6]<=E(-2,-1)
[0545] οMuxIn3[7]<=E(-2,+1)
[0546] Input multiplexers MuxA, MuxB, MuxC, MuxF, output multiplexers MuxD, MuxE, output F, delayed output Fd, output G, and delayed output Gd provide the connectivity listed in this section. The symbol X(N, M) represents the output X(row+N%6, col+M%16) of the DPU. The DPUs in each row are connected to each other in a circular manner, and the DPUs in each column are connected to each other in a circular manner. It should be noted that the various multiplexers listed below have multiple (e.g., sixteen) inputs, and the following list provides the connections to each of these inputs. For example, A[0]-A
[15] are the sixteen inputs of MuxA. In the following text, 6% represents modulo 6 operation, and 16% represents modulo 16 operation. R14 and 51R are the last two registers in the register file.
[0547] Input A multiplexer:
[0548] οA[0]<=0
[0549] οA[1]<=nu_PxlA
[0550] οA[2]<=D(0,0)
[0551] οA[3]<=E(0,0)
[0552] οA[4]<=MuxIn0
[0553] οA[5]<=MuxIn1
[0554] οA[6]<=MuxIn2
[0555] οA[7]<=MuxIn3
[0556] οA[8]<=PxlAb
[0557] οA[9]<=PxlBb
[0558] οA
[10] <=0
[0559] οA
[11] <=F(0,0)
[0560] οA
[12] <=F / Fd(0,-1)
[0561] οA
[13] <=R15
[0562] οA
[14] <=R14
[0563] οA
[15] <=PxlBi
[0564] • Input B multiplexer:
[0565] οB[0]<=0
[0566] οB[1]<=nu_PxlB
[0567] οB[2]<=D(0,0)
[0568] οB[3]<=E(0,0)
[0569] οB[4]<=MuxIn0
[0570] οB[5]<=MuxIn1
[0571] οB[6]<=MuxIn2
[0572] οB[7]<=MuxIn3
[0573] οB[8]<=PxlAb
[0574] οB[9]<=PxlBb
[0575] οB
[10] <=0
[0576] οB
[11] <=F(0,0)
[0577] οB
[12] <=F / Fd(0,-1)
[0578] οB
[13] <=R15
[0579] οB
[14] <=R14
[0580] οB
[15] <=PxlBi
[0581] Input C1 multiplexer:
[0582] οCl[0] <= 0
[0583] οCl[1] <= nu_PxlA
[0584] οCl[2] <= D(0, 0)
[0585] οCl[3] <= E(0, 0)
[0586] οCl[4] <= MuxIn0
[0587] οCl[5] <= MuxIn1
[0588] οCl[6] <= MuxIn2
[0589] οCl[7] <= MuxIn3
[0590] οCl[8] <= PxlAb
[0591] οCl[9] <= PxlBb
[0592] οCl
[10] <= 0
[0593] οCl
[11] <= F(0, 0)
[0594] οCl
[12] <= F / Fd(0, -1)
[0595] οCl
[13] <= R15
[0596] οCl
[14] <= R14
[0597] οCl
[15] <= PxlBi
[0598] · Input Ch multiplexer:
[0599] οCh[0] <= 0
[0600] οCh[1] <= nu_PxlB
[0601] οCh[2] <= D(0, 0)
[0602] οCh[3] <= E(0, 0)
[0603] οCh[4] <= MuxIn0
[0604] οCh[5] <= MuxIn1
[0605] οCh[6] <= MuxIn2
[0606] οCh[7] <= MuxIn3
[0607] οCh[8]<=PxlAb
[0608] οCh[9]<=PxlBb
[0609] οCh
[10] <=0
[0610] οCh
[11] <=F(0,0)
[0611] οCh
[12] <=F / Fd(0,-1)
[0612] οCh
[13] <=R15
[0613] οCh
[14] <=R14
[0614] οCh
[15] <=PxlBi
[0615] • Output D multiplexer: {Regs[0..15]}
[0616] • Output E-multiplexer: {Regs[0..15]}
[0617] • Input F multiplexer:
[0618] οF[0]: Use the F multiplexer defined in dpu_CSR
[0619] οF[l]<=G / Gd(+1,0)
[0620] οF[2]<=G / Gd(+2,0)
[0621] οF[3]<=G / Gd(+3,0)
[0622] οF[4]<=G / Gd(0,+1)
[0623] οF[5]<=G / Gd(0,+2)
[0624] οF[6]<=G / Gd(0,+4)
[0625] οF[7]<=G / Gd(0,+8)
[0626] • Output F: F multiplexer output
[0627] • Output Fd: F latch @clk
[0628] • Output G multiplexer:
[0629] οG[0]<= Using the G multiplexer defined in dpu_CSR
[0630] οG[l]<=D
[0631] οG[2]<=E
[0632] οG[3]<=F
[0633] οG[4]<=D(+1,0)
[0634] οG[5]<=E(+1,0)
[0635] οG[6]<=NU_A
[0636] οG[7]<=NU_B
[0637] οG[8]<=MuxG, MuxG<=MuxInD0
[0638] οG[9]<=MuxG, MuxG<=MuxInD1
[0639] οG
[10] <=MuxG, MuxG<=MuxInE0
[0640] οG
[11] <=MuxG, MuxG<=MuxInE1
[0641] οG
[12] <=PxlAb
[0642] οG
[13] <=PxlBb
[0643] οG
[14] <=PxlAbi
[0644] οG
[15] <=PxlBbi
[0645] • Output Gd: G latch @clk
[0646] Input (A, B, Cl, Ch, G) register configuration specifications (configuration bits are stored in the DPU's configuration register and are used to control the various components of the DPU).
[0647] The following example lists the values of the individual bits included in the DPU configuration instructions. The symbol X(N, M) represents the DPU output X(row+N%6, col+M%16).
[0648] Input A multiplexer:
[0649] οA[0]<=0
[0650] οA[1]<=nu_PxlA
[0651] οA[2]<=D(0,0)
[0652] οA[3]<=E(0,0)
[0653] οA[4]<=MuxIn0
[0654] οA[5]<=MuxIn1
[0655] οA[6]<=MuxIn2
[0656] οA[7]<=MuxIn3
[0657] οA[8]<=nu_PxlAb
[0658] οA[9]<=nu_PxlBb
[0659] οA
[10] <=nu_PxlAbi
[0660] οA
[11] <=F(0,0)
[0661] οA
[12] <=F / Fd(0,−1)
[0662] οA
[13] <=R15
[0663] οA
[14] <=R14
[0664] οA
[15] <=nu_PxlBbi
[0665] οA
[16] <=D(0,−1)
[0666] οA
[17] <=D(0,+1)
[0667] οA
[18] <=D(−1,0)
[0668] οA
[19] <=D(−1,−1)
[0669] οA
[20] <=D(−1,+1)
[0670] οA
[21] <=D(−2,0)
[0671] οA
[22] <=D(−2,−1)
[0672] οA
[23] <=D(−2,+1)
[0673] οA
[24] <=E(0,−1)
[0674] οA
[25] <=E(0,+1)
[0675] οA
[26] <=E(−1,0)
[0676] οA
[27] <=E(−1,−1)
[0677] οA
[28] <=E(−1,+1)
[0678] οA
[29] <E(-2,0)
[0679] οA
[30] <E(-2,-1)
[0680] οA
[31] <E(-2,+1)
[0681] ·
[0682] οB[0]<0
[0683] οB[1]<nu_PxlB
[0684] οB[2]<D(0,0)
[0685] οB[3]<E(0,0)
[0686] οB[4]<6MuxIn0
[0687] οB[5]<6MuxIn1
[0688] οB[6]<6MuxIn2
[0689] οB[7]<6MuxIn3
[0690] οB[8]<nu_PxlAb
[0691] οB[9]<nu_PxlBb
[0692] οB
[10] <nu_PxlAbi
[0693] οB
[11] <6F(0,0)
[0694] οB
[12] <6F / Fd(0,-1)
[0695] οB
[13] <IR15
[0696] οB
[14] <IR1
[0697] οB
[15] <nu_PxlBbi
[0698] οB
[16] <D(0,-1)
[0699] οB
[17] <D(0,+1)
[0700] οB
[18] <D(-1,0)
[0701] οB
[19] <D(-1,-1)
[0702] οB
[21] <D(-2,0)
[0703] οB
[22] <= D(-2, -1)
[0704] οB
[23] <= D(-2, +1)
[0705] οB
[24] <= E(0, -1)
[0706] οB
[25] <= E(0, +1)
[0707] οB
[26] <= E(-1, 0)
[0708] οB
[27] <= E(-1, -1)
[0709] οB
[28] <= E(-1, +1)
[0710] οB
[29] <= E(-2, 0)
[0711] οB
[30] <= E(-2, -1)
[0712] οB
[31] <= E(-2, +1)
[0713] · Input Cl multiplexer:
[0714] οCl[0] <= 0
[0715] οCl[1] <= nu_PxlA
[0716] οCl[2] <= D(0, 0)
[0717] οCl[3] <= E(0, 0)
[0718] οCl[4] <= MuxIn0
[0719] οCl[5] <= MuxIn1
[0720] οCl[6] <= MuxIn2
[0721] οCl[7] <= MuxIn3
[0722] οCl[8] <= nu_PxlAb
[0723] οCl[9] <= nu_PxlBb
[0724] οCl
[10] <= nu_PxlAbi
[0725] οCl
[11] <= F(0, 0)
[0726] οCl
[12] <= F / Fd(0, -1)
[0727] οCl
[13] <= R15
[0728] οCl
[14] <= R14
[0729] οCl
[15] <= nu_PxlBbi
[0730] οCl
[16] <= D(0, -1)
[0731] οCl
[17] <= D(0, +1)
[0732] οCl
[18] <= D(-1, 0)
[0733] οCl
[19] <= D(-1, -1)
[0734] οCl
[21] <= D(-2, 0)
[0735] οCl
[22] <= D(-2, -1)
[0736] οCl
[23] <= D(-2, +1)
[0737] οCl
[24] <= E(0, -1)
[0738] οCl
[25] <= E(0, +1)
[0739] οCl
[26] <= E(-1, 0)
[0740] οCl
[27] <= E(-1, -1)
[0741] οCl
[28] <= E(-1, +1)
[0742] οCl
[29] <= E(-2, 0)
[0743] οCl
[30] <= E(-2, -1)
[0744] οCl
[31] <= E(-2, +1)
[0745] · Input Ch multiplexer:
[0746] οCh[0] = 0
[0747] οCh[1] = nu_PxlB
[0748] οCh[2] = D(0, 0)
[0749] οCh[3] = E(0, 0)
[0750] οCh[4] = MuxIn0
[0751] οCh[5]<=MuxIn1
[0752] οCh[6]<=MuxIn2
[0753] οCh[7]<=MuxIn3
[0754] οCh[8]<=nu_PxlAb
[0755] οCh[9]<=nu_PxlBb
[0756] οCh
[10] <=nu_PxlAbi
[0757] οCh
[11] <=F(0,0)
[0758] οCh
[12] <=F / Fd(0,-1)
[0759] οCh
[13] <=R15
[0760] οCh
[14] <=R14
[0761] οCh
[15] <=nu_PxlBbi
[0762] οCh
[16] <=D(0,-1)
[0763] οCh
[17] <=D(0,+1)
[0764] οCh
[18] <=D(−1,0)
[0765] οCh
[19] <=D(-1,-1)
[0766] οCh
[21] <=D(−2,0)
[0767] οCh
[22] <=D(-2,-1)
[0768] οCh
[23] <=D(-2,+1)
[0769] οCh
[24] <=E(0,-1)
[0770] οCh
[25] <=E(0,+1)
[0771] οCh
[26] <=E(-1,0)
[0772] οCh
[27] <=E(-1,-1)
[0773] οCh
[28] <=E(-1,+1)
[0774] οCh
[29] <=E(−2,0)
[0775] οCh
[30] <=E(-2,-1)
[0776] οCh
[31] <=E(-2,+1)
[0777] • Output G multiplexer:
[0778] οG[0]<= Using the G multiplexer defined in dpu_CSR
[0779] οG[1]<=D
[0780] οG[2]<=E
[0781] οG[3]<=F
[0782] οG[4]<=D(+1,0)
[0783] οG[5]<=E(+1,0)
[0784] οG[6]<=NU_A
[0785] οG[7]<=NU_B
[0786] οG[8]<=MuxG, MuxG<=MuxInD0
[0787] οG[9]<=MuxG, MuxG<=MuxInD1
[0788] οG
[10] <=MuxG, MuxG<=MuxInE0
[0789] οG
[11] <=MuxG, MuxG<=MuxInE1
[0790] οG
[12] <=PxlAb
[0791] οG
[13] <=PxlBb
[0792] οG
[14] <=PxlAbi
[0793] οG
[15] <=PxlBb
[0794] οG
[16] <=D(0,-1)
[0795] οG
[17] <=D(0,+1)
[0796] οG
[18] <=D(-1,0)
[0797] οG
[19] <=D(-1,-1)
[0798] οG
[21] <=D(-2,0)
[0799] οG
[22] <=D(-2,-1)
[0800] οG
[23] <=D(-2,+1)
[0801] οG
[24] <=E(0,-1)
[0802] οG
[25] <=E(0,+1)
[0803] οG
[26] <=E(-1,0)
[0804] οG
[27] <=E(-1,-1)
[0805] οG
[28] <=E(-1,+1)
[0806] οG
[29] <=E(-2,0)
[0807] οG
[30] <=E(-2,-1)
[0808] οG
[31] <=E(-2,+1)
[0809] The inputs (A, B, Cl, Ch, G) are multiplexed to (MuxIn0, ..., MuxIn3) and have two additional inputs. To preserve the configuration bits, dynamic resource allocation of the (MuxIn0, ..., MuxIn3) configuration bits D(N, M) and E(N, M) of (A, B, Cl, Ch, G) can be used. The allocation can be performed as follows: if one or two of the inputs (A, B, Cl, Ch, G) configurations have the form D(N, M), then the first input is assigned to MuxIn0, and the second (if present) is assigned to MuxIn1. For convenience, the inputs are represented by inputD1 and inputD2, and the configurations are represented by D1(N, M) and D2(N, M), respectively. The control of the MuxIn0 and MuxIn1 multiplexers can be based on D1(N, M) and D2(N, M), respectively. Additionally, inputD1 and inputD2 multiplexers control MuxIn0 and MuxIn1. If one or two of the inputs (A, B, Cl, Ch, G) configuration have the form E(N, M), then the same dynamic allocation is applied to MuxIn2 and MuxIn3.
[0810] For example:
[0811] Assume the following configuration:
[0812] ·A=nu_PxlA
[0813] B = D(0, -1)
[0814] Cl = E(0, -1)
[0815] Ch = D(-1, 0)
[0816] ·G=MUXIn0
[0817] Then, the multiplexer will be controlled as follows:
[0818] MuxIn0 = D(0, -1)
[0819] MuxIn1 = D(-1, 0)
[0820] MuxIn2 = E(0, -1)
[0821] ·muxA=nu_PxlB
[0822] ·muxB=MuxIn0
[0823] ·muxCl=MuxIn2
[0824] ·muxCh=MuxIn1
[0825] ·muxG=MuxIn0
[0826] ALU opcodes
[0827] Integer arithmetic
[0828] 0.nop
[0829] 1.addc: {Out1, Out0}<=A+B+C
[0830] addc(v): Out0<=A1+B1+Cl; Out1=Ah+Bh+Ch
[0831] 2.adds: Out0<=(A+B)>>C; Out1<=(A+B)[C-1:0]
[0832] 3.addl:{Out1,Out0}<={B,A}+C
[0833] addl(v):Out0<=A+Cl; Out1<=B+Ch
[0834] 4.addrv: {Out1, Out0}<=A1+B1+Ah+Bh+C
[0835] 5.add4: {Out1, Out0}<=A+B+Cl+Ch
[0836] 6.subb: {Out1, Out0}<=A-B+C
[0837] subb(v):Out0<=Al-Bl+C1;Out1=Ah-Bh+Ch
[0838] 7.subc:{Out1,Out0}<=A+BC
[0839] subc(v):Out0<=A1+B1–C1;Out1=Ah+Bh-Ch
[0840] 8.subl:{Out1,Out0}<={B,A}-C
[0841] subl(v):Out0<=A–C1;Out1<=B-Ch
[0842] 9.subrv:{Out1,Out0}<=A1+B1+Ah+Bh-C
[0843] 10.mac:{Out1,Out0}<=A*B+C
[0844] mac(v):Out0<=A1*B1+C1;Out1=Ah*Bh+Ch
[0845] 11.macs:{Out1,Out0}<=A*BC
[0846] macs(v):Out0<=A1*B1–C1;Out1=Ah*Bh-Ch
[0847] 12.macrv:{Out1,Out0}<=(Al*B1)+(Ah*Bh)+C
[0848] 13.shift:{Out1,Out0}<=(A<<B)> >C
[0849] shift(v):Out0<=(A1<<B1)> >C1;Out1=(Ah<<Bh)> >Ch
[0850] 14.shiftrl:{Out0,Out1}<,{B,A}>>C
[0851] shiftrl(v):Out0<=A>>C1;Out1<=B>>Ch
[0852] 15.shiftll:{Out0,Out1}<={B,A}< <C
[0853] shiftll(v):Out0<=A< <C1;Out1<=B<<Ch
[0854] 16.absd:{Out1,Out0}<=|AB|+C
[0855] absd(v):Out0<=|A1-B1|+C1; Out1=|Ah-Bh|+Ch
[0856] 17.absdrv:{Out1,Out0}<=|A1–B1|+|Ah-Bh|+C
[0857] 18.absddrv:{Out1,Out0}<=|A1–B1|-|Ah-Bh|+C
[0858] 19. min: {Out1, Out0} <= min({B(val), A(Idx)}, C({val, Idx})) --> Output: Idx min
[0859] 20. equalone: {Out1, Out0} <= min({B(val), A(Idx)}, C({val, Idx})) --> Output: First operand
[0860] 21. equaltwo: {Out1, Out0} <= min({B(val), A(Idx)}, C({val, Idx})) --> Output: the second operand
[0861] 22.lessc:{Out1,Out0}<=(A <B)+C
[0862] lessc(v): Out0 <= (Al) <Bl)+Cl;Out1=(Ah<Bh)+Ch
[0863] 23.lesseqc:{Out1,OUt0}<=(A<=B)+C
[0864] lesseq(v): Out0<=(A1<=B1)+C1; Out1=(Ah<=Bh)+Ch
[0865] 24.equalc:{Out1,Out0}<=(A==B)+C
[0866] equalc(v): Out0<=(A1==B1)+C1; Out1=(Ah==Bh)+Ch
[0867] 25.nequalc: {Out1, Out0}<=(A!=B)+C
[0868] nequalc(v): Out0<=(A1!=B1)+C1; Out1=(Ah!=Bh)+Ch
[0869] 26.lesscrv:{Out1,Out0}<=(A1 <B1)+(Ah<Bh)+C
[0870] 27.lesseqcrv: {Out1, OUt1}<=(A1<=B1)+(Ah<=Bh)+C
[0871] 28.equalcrv: {Out1, Out0}<=(A1==B1)+(Ah==Bh)+C
[0872] 29.nequalcrv: {Out1, Out0}<=(A1!=B1)+(Ah!=Bh)+C
[0873] Floating-point arithmetic
[0874] The C exponent bias parameter in the conversion can be viewed as a 7-bit signed integer (the other 25 bits of C can be ignored, and the sign bit is extended).
[0875] 30.short2fp: {Out1, Out0}<=A*2**C
[0876] 31.word2fp:{Out1,Out0}<={B,A}*2**C
[0877] 32 lessf{Out1, Out0}<={B, A} <C
[0878] 33 lesseqf{Out1, Out0}<={B, A}<=C
[0879] 34.fp2word:{Out1,Out0}<={B,A}*2**C
[0880] 35.mulf: {Out1, Out0}<={B, A}*C
[0881] 36.addf:{Out1,Out0}<={B,A}+C
[0882] 37.subf:{Out1,Out0}<={B,A}-C
[0883] 38.int_div: {Out1, Out0}<=1 / A
[0884] addition operation
[0885] 39.or{Out0, Out1} <= {B, A}|Cn
[0886] 40.xor{Out0, Out1} <= {B, A} ^ AC
[0887] 41.and{Out0, Out1} <= {B, A}&C
[0888] 42.equal32{Out1, Out0} <= {B, A} == C
[0889] 43.nequal32{Out1, Out0} <= {B, A}! = C
[0890] 44.less32{Out1, Out0} <= {B, A} < C
[0891] 45.lesseq32{Out1, Out0} <= {B, A} <= C
[0892] 46.abs32{Out1, Out0} <= |{B, A} - C|
[0893] 47.max32{Out1, Out0} <= max({B, A}, C) --> Output: 32max
[0894] 48.minf{Out1, Out0} <= min({B, A}, C) --> Output: fp min
[0895] 49.maxf{Out1, Out0} <= max({B, A}, C) --> Output: fp max <unk>
[0896] 50.mac4{Out1, Out0} <= (A1*B1)+(Ah*Bh)+(C0*C2)+(C1*C3)+acc
[0897] mac4(v)Out0 <= (Al*B1)+(Ah*Bh)+acc0; Out1 = (C0*C2)+(C1*C3)+acc1
[0898] 51.shift_sat{Out1, Out0} <= Sat({B, A}, C): rnd = byteMode: 0 shortMode = 1|notRELU: 2
[0899] 52.add_sat{Out1, Out0} <= Sat({A + B}, C)
[0900] It should be noted that there is an "unk" in the translation of line 32 which might be an error in the original text. You may want to double-check the original content for accuracy.add_sat(v){Out1, Out0}<=Sat({Al+Bl}, C)
[0901] DPU microcontroller instruction encoding
[0902] Execute command:
[0903] [8:0]rc: Repeat counter:
[0904] [0..495]: Immediate value
[0905] [496..511]: Counter [rc-496]
[0906] [10:9]sel: DPU instruction selection
[0907]
[11] Collection: Activate the collection unit (microcontroller 0 only)
[0908] [14:12] Retain
[0909]
[15] Type 0
[0910] Do loop:
[0911] [8:0]RLC: Repeating Cycle Counter:
[0912] [0..495]: Immediate value
[0913] [496..511]: Counter [rc-496]
[0914] [13:9] Length: loop length: [1..32].
[0915]
[14] Mode: 0: counting cycle, 1: counting period
[0916]
[15] Type 1
[0917] DPUs can be configured one at a time (each DPU has a unique unicast address) or they can be configured in broadcast mode - there are addresses that can reflect the rows and / or columns of DPUs sharing the same rows and / or columns, and this allows the configuration information to be broadcast.
[0918] Data processing array: linear mapping (programming a single DPU configuration register)
[0919]
[0920]
[0921] Data processing array: Broadcast mapping (simultaneous programming of configuration registers for multiple DPUs)
[0922]
[0923] It should be noted that any image processing algorithm can be executed iteratively by the image processor. Results for some pixels are processed by the DPA 500. Some results may be stored in the DPA for a period of time before being sent to the memory module. This time period is typically set based on the size of the PMA's memory resources and the number of source or target pixels processed by the DPA during a given task. These results may be retrieved from the memory module when needed again. For example, when the DPA 500 performs calculations on some source pixels of the source image, these results may be stored for a period of time (e.g., when performing calculations related to neighboring source pixels) before being sent to memory. When the results are needed further, they can be retrieved from the memory module.
[0924] Twisted computing
[0925] Distortion calculations may be applied for various reasons. For example, to compensate for imbalances in image acquisition.
[0926] Twist calculations can be performed by the DPA 500.
[0927] According to embodiments of the present invention, a warp calculation is applied to each target pixel (a pixel of the target image) in a set of target pixels. The target pixel set may include the entire target image or a portion of the target image. Typically, the target image is virtually divided into multiple windows, and each window is a set of target pixels.
[0928] The warp calculation can receive or compute a corresponding set of source pixels. The source pixels in the corresponding set are processed during the warp calculation. Typically, the selection of the source pixels in the corresponding set is fed into the PMA, and can depend on, for example, the desired warp function.
[0929] The distortion value of the target pixel is calculated by applying weights (Wx, Wy) to its neighboring source pixels associated with it. The weights and coordinates (x, y) of at least one neighboring source pixel are defined by distortion parameters (X', Y').
[0930] Figure 17 A method 1700 according to an embodiment of the present invention is shown.
[0931] Method 1700 can begin at step 1710, selecting a target pixel from a set of target pixels. The selected target pixel will be referred to as the "target pixel".
[0932] Step 1710 can be followed by step 1720, which performs a warp calculation process for each target pixel in a set of target pixels, including:
[0933] 1) Calculate (1721) or receive warp parameters for the selected target pixel. The warp parameters may include a first weight and a second weight (Wx, Wy) and coordinates (x, y) of a given source pixel that should be processed during warp calculation. The first weight and the second weight are received by the first set of processing units (DPUs) in the processing unit array (DPA).
[0934] 2) Request (1722) neighboring source pixels (which include the given source pixel) from a memory cell such as a collection unit. The collection unit can receive 4 coordinates in various operating modes and convert them into 16 source pixels - four groups of neighboring source pixels.
[0935] 3) Receive (1723) neighboring source pixels associated with the target pixel through the second set of processing units.
[0936] 4) Calculate (1724) the distortion result in response to the values of neighboring source pixels and a pair of weights through the second set of processing units; provide the distortion result to the memory module.
[0937] Steps 1721, 1722, 1723 and 1724 can be executed in a pipeline manner.
[0938] refer to Figure 18 The first group of processing units is represented as 505 and may include the leftmost four DPUs in the top first row of DPA 500. The second group of processing units is represented as 501 and may include the rightmost two columns of DPA 500.
[0939] Step 1720 is followed by step 1730, which checks whether the warp is calculated for all target pixels in the group. If not, the warp calculation ends.
[0940] Step 1726 may include relaying the values of some adjacent source pixels between the processing units in the second group.
[0941] Figure 18 and Figure 19 The output signal of DPU (0, 4) (X' of group 504) is shown to be sent to DPU (0, 15) and then relayed to DPU (1, 15). It should be noted that in... Figure 18 and Figure 19 In the process, PMA calculates the distortion function of four pixels in parallel:
[0942] 1) DPU(0,3), DPU(1,3) and DPU in group 501 are involved in calculating the distortion of the first pixel.
[0943] 2) DPU(0,2), DPU(1,2) and the DPU of group 502 are involved in calculating the distortion of the second pixel.
[0944] 3) DPU(0,1), DPU(1,1) and the DPU of group 503 are involved in calculating the distortion of the third pixel.
[0945] 4) DPU(0,0), DPU(1,0) and the DPU of group 504 are involved in calculating the distortion of the third pixel.
[0946] Step 1726 may include relaying the intermediate results calculated by the second group to the values of some neighboring source pixels between the processing units of the second group.
[0947] Figure 20 It is shown that the twisting parameters (X' of groups 501-504 and Y' of groups 501-504) are sent from the DPUs (510(0,0)-510(0,3) and 510(1,0)-510(1,3)) of groups 505 and 506 to groups 501, 502, 503 and 504.
[0948] Figure 21 and Figure 22 The distortion calculations performed by DPUs 510(0,15)-510(3,15) and DPUs 510(0,14)-510(5,14) of group 501 according to an embodiment of the present invention are shown.
[0949] Figure 21 The distortion calculation includes the following steps (some of which are performed in parallel). Steps 1751-1762 are also included. Figure 22 As shown in the image.
[0950] The first processing unit (DPU 510(5,14)) of the second group calculates (1751) the first difference (P0-P2) between the first pair of adjacent source pixels and the second difference (P1-P3) between the second pair of adjacent source pixels.
[0951] The first difference is provided (1752) to the second processing unit (DPU 510(1,14)) of the second group, and the second difference is provided to the third processing unit of the second group.
[0952] The fourth processing unit (DPU 510(1,15)) of the second group responds to the first weight calculation (1753) and modifies the first weight Wy'.
[0953] The first modified weight is provided from the fourth processing unit (1754) to the second processing unit of the second group (DPU 510(1,14)).
[0954] The second processing unit of the second group calculates (1755) the first intermediate result (Var0) based on the first difference (P0-P2), the first neighboring source pixel (P0), and the first modified weight (Wy'). Var0 = (P0-P0)*Wy'-P0.
[0955] The second difference (P1-P3) is provided from the third processing unit of the second group to the sixth processing unit (DPU 510(0,15)) of the second group.
[0956] The second adjacent source pixel (P1) is provided (1757) from the fifth processing unit (DPU 510(0,14)) of the second group to the sixth processing unit (DPU 510(2,14)) of the second group.
[0957] The second intermediate result Var1 is calculated by the sixth processing unit of the second group based on the second difference, the second adjacent source pixel and the first modified weight. Var1 = (P1-P3)*Wy'-P1.
[0958] The second intermediate result Var1 from the sixth processing unit of the second group is provided (1759) to the seventh processing unit of the second group (DPU 510(2,15)), and the first intermediate result Var0 is provided from the second processing unit of the second group to the seventh processing unit of the second group.
[0959] The third intermediate result Var2, relative to the first and second intermediate results, is calculated by the seventh processing unit of the second group (1760). Var2 = Var0 - Var1.
[0960] The third intermediate result is provided from the seventh processing unit (1761) of the second group to the eighth processing unit (DPU 510(3,15)) of the second group. The second intermediate result is provided from the sixth processing unit of the second group to the ninth processing unit (DPU 510(3,14)) of the second group.
[0961] The second intermediate result is provided from the ninth processing unit of the second group (1762) to the eighth processing unit of the second group. The second modified weight (Wx') is provided from the third processing unit of the second group to the eighth processing unit of the second group.
[0962] The warp result (1763) is calculated by the eighth processing unit of the second group based on the second and third intermediate results and the second modified weights. Warp_result = Var2*Wx' + Var1.
[0963] like Figure 20 As shown, DPU 510 (5, 14) can receive pixels P0, P1, P2, and P3 from the collection unit. When DPA500 processes four pixels at a time, groups 501, 502, 503, and 504 receive 16 pixels from the collection unit (in parallel).
[0964] It should be noted that the DPA 500 also receives (e.g., from the collection unit) the distortion parameters X', Y' associated with each pixel.
[0965] According to one embodiment of the invention, the distortion parameters of each pixel can be calculated by the DPU of the DPA - for example, when the distortion parameters can be expressed by a mathematical formula such as a polynomial.
[0966] Figure 23 A set of DPUs 507 is shown that compute X' and Y', and these computed X' and Y' can be fed to sets 505 and 506.
[0967] It should be noted that Figures 18 to 22 Only non-restrictive grouping schemes are shown. Twist calculations can be performed by DPU groups of other shapes and sizes.
[0968] Parallax
[0969] Disparity calculation aims to find the best matching target pixel for a source pixel. A search can be performed against all source pixels in the source image and against all target pixels in the target image – however, this is not always the case, and disparity may be applied only to some source pixels in the source image and / or some target pixels in the target image.
[0970] Parallax calculation involves more than just comparing the difference between a single source pixel and a single target pixel; it also compares subgroups of source pixels with subgroups of target pixels. The comparison can include calculating functions such as the sum of absolute differences (SAD) between source pixels and their corresponding target pixels.
[0971] The source pixel can be located at the center of a source pixel subgroup, and the target pixel can be located at the center of a target pixel subgroup. Other locations for the source pixel and / or target pixel can be used.
[0972] Subgroups of source pixels and subgroups of target pixels can be rectangular (or can have any other shape) and can include N rows and N columns, where N can be an odd positive integer greater than three.
[0973] Most of the parallax calculations likely benefit from previous computer parallax calculations. These are shown in... Figure 24 and 25 It was provided in China.
[0974] Figure 24 The first subgroup 1001 of 5×5 source pixels S(1,1)-S(5,5), the first subgroup 1002 of 5×5 target pixels T(1,1)-T(5,5), the second subgroup 1003 of 5×5 source pixels S(1,2)-S(5,6), and the second subgroup 1004 of 5×5 target pixels T(1,2)-T(5,6) are shown.
[0975] Source pixels S(3,3) and S(3,4) are located at the center of the first subgroup 1001 and the second subgroup 1003 of the source pixels. Target pixels T(3,3) and T(3,4) are located at the center of the first subgroup 1002 and the second subgroup 1004 of the target pixels.
[0976] The SAD associated with S(3,3) and T(3,3) is equal to:
[0977] SAD(S(3,3),T(3,3))=SUM(|S(i,j)-T(i,j)|)-for indices i and j between 1 and 5.
[0978] The SAD associated with S(3,4) and T(3,4) is equal to:
[0979] SAD(S(3,4),T(3,4))=SUM(|S(i,j)-T(i,j)|)-index i between 2 and 6 and index j between 1 and 5.
[0980] Assume that SAD is calculated from left to right. Under this assumption, the calculation of -SAD(S(3,4),T(3,4)) may benefit from the calculation of SAD(S(3,3),T(3,3)).
[0981] Specifically: SAD(S(3,4), T(3,4)) = SAD(S(3,3), T(3,3)) - SAD(rightmost column of the first subgroup of source and target pixels) + SAD(leftmost column of the second subgroup of source and target pixels).
[0982] Since the source and target images are two-dimensional, and it is assumed that the source pixels are scanned from left to right (per piece) and from top to bottom, the calculation of SAD is more efficient.
[0983] Figure 25 A subgroup SG(B) of source pixels with center pixel SB is shown. Figure 26 The corresponding subgroup TG(B) of the target pixel (not shown) with center pixel TB is shown.
[0984] Calculate SUD for the source pixels in the row above the row of SB and for the pixels to the left of SB in the same row.
[0985] Pixel SA is the center of subgroup SG(A) and the left neighbor of pixel SB. Target pixel TA is the left neighbor of pixel SB and the center of subgroup TG(A).
[0986] Pixel SC is the center of subgroup SG(C) and the upper neighbor of pixel SB. Target pixel TC is the upper neighbor of pixel SB and the center of subgroup TG(C).
[0987] The leftmost column of SG(A) is represented as 1110. The rightmost column of SG(C) is represented as 1114. The current rightmost column of SG(B) is represented as 1115. The rightmost lowest pixel of SG(B) (also known as the new source pixel NSP) is represented as 1116. The old pixel at the top of the current rightmost column of SG(B) (belonging to SG(C)) (also known as the old source pixel NSP) is represented as 1112.
[0988] The leftmost column of TG(A) is represented as 1110'. The rightmost column of TG(C) is represented as 1114'. The current rightmost column of TG(B) is represented as 1115'. The rightmost lowest pixel of TG(B) (also known as the new target pixel NTP) is represented as 1116'. The old pixel at the top of the current rightmost column of TG(B) (belonging to TG(C)) (also known as the old target pixel NTP) is represented as 1112'.
[0989] The SAD of (SB, TB) can be calculated as follows:
[0990] SAD(SA,TA).
[0991] -SAD(SG(A) leftmost column, SG(B) leftmost column).
[0992] +SAD(the rightmost column of SG(C), the rightmost column of TG(C)).
[0993] The absolute difference between the rightmost lowest source pixel and target pixel of +SG(B) and TG(B).
[0994] -SG(C) and TG(C) are the absolute differences between the top source and target pixels in the rightmost column.
[0995] Figure 27 A method 2600 according to an embodiment of the present invention is shown.
[0996] Method 2600 may begin at step 2610, selecting source pixels and selecting a subgroup of target pixels. The subgroup of target pixels may be a portion of the entire target image.
[0997] Step 2610 can be followed by step 2620, where a set of absolute differences (SAD) is calculated by the first set of data processors in the data processor array.
[0998] This group of SADs is associated with a subgroup of target pixels, including the target pixels selected in step 2610, and a source pixel. Different SADs for this group are calculated based on different target pixels within the subgroups of (the same) source and target pixels.
[0999] Calculating the SAD of the same source pixels reduces the amount of data extracted into DPA.
[1000] A subgroup of target pixels may include target pixels that are stored sequentially in a memory module. The subgroup of target pixels is extracted from the memory module before calculating the SAD of this group. Extraction of the subgroup of target pixels from the memory module is performed by a collection unit that includes a content-addressable memory cache.
[1001] Each SAD is calculated based on the previously calculated SAD and the absolute difference between other source pixels and other target pixels belonging to the subgroup of the target pixel. Figure 25 An example of such a calculation is provided.
[1002] Step 2620 may be followed by step 2630, in which the second set of data processors in the array finds the best matching target pixel in the subgroup of the target pixel in response to the value of the set of SAD.
[1003] Steps 2620 and 2630 may include storing the computation results in the data processor array—the SAD of the entire rectangular pixel array, the SAD of each column, etc. It should be noted that the register file depth of each DPU can be long enough to store the SAD of the rightmost column of the previous rectangular array. For example, if there are 15 columns in SG(A), then the DPU's register file 550 should be at least 15.
[1004] After storing the previous SAD, for a given SAD, store the first previously calculated SAD, the second previously calculated SAD, the target pixel at the top of the second target pixel column, and the source pixel at the top of the second source pixel column.
[1005] Referring to step 2620 - the previously calculated SAD can reflect the absolute difference between the following two items: (i) the rectangular source pixel array differing from the given rectangular source pixel array by the first source pixel column and the second source pixel column, and (ii) the rectangular target pixel array differing from the given rectangular target pixel array by the first target pixel column and the second target pixel column. For example - SAD(SA, TA).
[1006] The previously calculated SAD can be reflected in the absolute difference between the first source column and the first source column. For example, -SAD(the leftmost column of SG(A), the leftmost column of SG(B)).
[1007] Step 2620 may include:
[1008] 1) By subtracting (a) the previously calculated second SAD (e.g., -SAD(the leftmost column of SG(A), the leftmost column of SG(B)) and (b) the absolute difference between the following two items from the first previously calculated SAD (e.g., SAD(SA, TA)): (i) the target pixel at the top of the second target pixel column and (ii) the source pixel at the top of the second source pixel column (e.g., the absolute difference between -OSP1112 and OTP 1112').
[1009] 2) Add the absolute difference between the lowest target pixel in the second target pixel column and the lowest source pixel in the second source pixel column to the intermediate result (e.g., the absolute difference between -NSP 1116 and NTP 1116').
[1010] It should be noted that finding the best matching target pixel may involve an iterative process, and steps 2610, 2620 and 2630 may be repeated multiple times for different subgroups of pixels and by comparing the results of these multiple iterations, the best matching target pixel in the group of target pixels can be found.
[1011] It should also be noted that the processing unit array can perform multiple disparity calculations in parallel (for different source pixels and / or for different target pixels).
[1012] Figure 28 The illustration shows eight source pixels and thirty-two target pixels processed by DPA according to an embodiment of the present invention. Figure 29 A source pixel array is shown according to an embodiment of the present invention. Figure 30 A target pixel array is shown according to an embodiment of the present invention.
[1013] Calculate the SAD associated with the following pixels: source pixels (SP0, SP1, SP2 and SP3) and (SP'0, SP'1, SP'2 and SP'3), 4×8 target pixels (including the leftmost column of TP0, TP1, TP2 and TP3), and another 4×8 target pixels (including the leftmost column of TP'0, TP'1, TP'2 and TP'3).
[1014] Source pixels SP0, SP1, SP2, and SP3 belong to the same column, and their SAD is calculated in a pipelined manner:
[1015] 1) Calculate SAD for SP0 and a target pixel.
[1016] 2) Use the previous calculations when calculating SP1 and the SAD of a target pixel.
[1017] 3) Use the previous calculations when calculating SP2 and the SAD of a target pixel.
[1018] 4) Use previous calculations when calculating SP3 and the SAD of a target pixel.
[1019] While calculating the SAD of source pixels SP0, SP1, SP2, and SP3, PMA also calculates the SAD of SP'0, SP'1, SP'2, and SP'3. SP'0, SP'1, SP'2, and SP'3 belong to the same column, and their SADs are calculated in a pipelined manner.
[1020] 1) Calculate SAD for SP'0 and a target pixel.
[1021] 2) Use the previous calculations when calculating SP'1 and the SAD of a target pixel.
[1022] 3) Use the previous calculations when calculating SP'2 and the SAD of a target pixel.
[1023] 4) Use the previous calculations when calculating SP'3 and the SAD of a target pixel.
[1024] The DPA 500 can compute the SAD of each source pixel and multiple other target pixels in parallel.
[1025] For example, assuming the first row of 4×8 target pixels includes TP0 and seven shifted target pixels (TP0, Ts1P0, Ts2P0, Ts2P0, Ts3P0, Ts4P0, Ts5P0, Ts6P0, Ts7P0), the calculation of SAD for SP0 can include calculating SAD for each of SP0 and TP0, Ts1P0, Ts2P0, Ts2P0, Ts3P0, Ts4P0, Ts5P0, Ts6P0, Ts7P0.
[1026] When calculating any SAD, it is necessary to calculate the absolute difference of the new pixels. Figure 29 Four new source pixels, NSO, NS1, NS2, and NS3, are shown (used to calculate SAD associated with SP0, SP1, SP2, and SP3, as well as only one target pixel column).
[1027] Figure 30 32 new target pixels are shown:
[1028] 1) Used to calculate the new target pixel SAD for SP0 and eight different target pixels - NT0, Ns1T0, Ns2T0, Ns3T0, Ns4T0, Ns5T0, Ns6T0, Ss7T0.
[1029] 2) Used to calculate new target pixels for SP1 and the SAD of eight different target pixels - NT1, Ns1T1, Ns2T1, Ns3T1, Ns4T1, Ns5T1, Ns6T1, Ss7T1.
[1030] 3) Used to calculate the new target pixels for SP2 and the SAD of eight different target pixels - NT2, Ns1T2, Ns2T2, Ns3T2, Ns4T2, Ns5T2, Ns6T2, Ss7T2.
[1031] 4) Used to calculate the new target pixel for SP3 and the SAD of eight different target pixels - NT3, Ns1T3, Ns2T3, Ns3T3, Ns4T3, Ns5T3, Ns6T3, Ss7T3.
[1032] Figure 31 Groups 1131, 1132, 1133, 1134, 1135, 1136, 1137 and 1138 of eight DPUs are shown - each group includes 4 DPUs.
[1033] For each group 1131, 1132, 1133, and 1134, the SAD of pixels SP0, SP1, SP2, and SP3 is calculated—but for different target pixels (TP0, TP2, TP3, and TP4).
[1034] For each group 1135, 1136, 1137, and 1138, calculate the SAD of pixels SP'0, SP'1, SP'2, and SP'3 – but for different target pixels (TP0, TP2, TP3, and TP4).
[1035] Pixel group 1140 performs a minimization operation on the SAD calculated from groups 1131-1138.
[1036] Accordingly, method 2600 may include calculating multiple sets of SADs associated with multiple subgroups of multiple source pixels and target pixels by a first set of data processors in the data processor array; wherein each SAD in the multiple sets of SADs is calculated based on the absolute difference between a previously calculated SAD and a currently calculated SAD; and finding the best matching target pixel in response to the value of the SAD associated with the source pixel by a second set of data processors in the array and for the source pixel.
[1037] Multiple sets of SADs can include subgroups of SADs, each subgroup of SADs being associated with multiple subgroups of target pixels in multiple subgroups of source pixels and target pixels. For example, groups 1131-1138 compute different SAD subgroups.
[1038] Multiple source pixels can belong to columns of a rectangular pixel array and be adjacent to each other.
[1039] Computing multiple sets of SADs can include parallel computation of the SADs of different SAD subgroups.
[1040] The calculation may include calculating the SADs belonging to the same SAD subgroup in a sequential manner.
[1041] The following text describes some PMA states and configuration buffers according to embodiments of the present invention.
[1042] These status and configuration buffers 109 include a PMA control status register, a PMA pause enable control register, and a PMA pause event status register.
[1043] The control register allows, for example, a scalar unit to determine a predetermined operating period for the image processor. Alternatively or concurrently, the scalar unit can pause the image processor (without changing the state of the PMA) and program the program processor, sending control signals to the program processor and resuming the image processor's operation from the same point (except for the changes introduced by the scalar unit).
[1044] PMA Control Status Register (p_PmaCsr)
[1045] [2:0]lsuAddSel[0] Address selection LSU0
[1046] [5:3]lsuAddSel[l] Address selection LSU1
[1047] [8:6]lsuAddSel[2] Address selection LSU2
[1048] [11:9]lsuAddSel[3] Address selection LSU3
[1049] [14:12]lsuAddSel[4] Address selection LSU4
[1050] [17:15]lsuAddSel[5] Address selection LSU5
[1051] [23:18]Retain[26:24]agStopSel AGU generation stop condition selection
[1052]
[27] XorParity writes and reverses parity when enabled [29:28]sysMemMap system memory mapping
[1053]
[30] ProEn program enabled
[1054]
[31] suspCntEn suspend counter enabled
[1055] PMA HaltOnEvent (HaltOnEvent)
[1056] [15:0]mbParityEn memory[15:0]parity error pause enable
[16] spmparityEn SU program memory parity error pause enable
[17] sdmparityEn SU data memory parity error pause enable
[18] dmaIntEn DMA interrupt pause enable
[1057]
[19] suRdErrEn Scalar unit read error pauses enable
[1058]
[20] suDivZeroEn scalar unit divided by zero pauses enable
[1059]
[21] SuEnEn Scalar Unit Interrupt Pause Enabled
[1060]
[22] suHaltEn scalar unit is paused.
[1061] [31:23] Retain
[1062] PMA Pause Event Status Register (HoeStatus)
[1063] [15:0]mbParityErr storage [15:0] parity error
[1064]
[16] spmparityErr SU Program memory parity error
[1065]
[17] sdmparityErr SU Data storage parity error
[1066]
[18] dmaInt DMA interrupt
[1067]
[19] suRdErr scalar unit read error
[1068]
[20] suDivZero scalar unit divided by zero
[1069]
[21] suInt scalar unit interrupt
[1070]
[22] SuHalt scalar unit pause
[1071] [31:23] Retain
[1072] Suspend - Resume - Event Counter & Increment.
[1073] This feature enables suspend operations, allowing changes to certain configurations without clearing the computation pipeline, and then resuming operations. It is implemented through the following registers: (a) the suspend counter enable control bit (suspCntEn in p_PmaCsr), (b) the suspend counter (p_SuspCnt), and (c) the suspend reset control (p_RstCtl).
[1074] When enabled (suspCntEn = 1), a suspend counter counts down. When it reaches zero, the PMA suspends operations (remains stopped) until suspCntEn is reset or p_SuspCnt is written with a new value (!= 0). During the stop, the PMA can be reconfigured (instructions, constants, etc.) using scalar units. When suspCntEn is reset or p_StallCnt is written, the PMA resumes its operation with the new configuration. p_RstCtl defines which functions are reset upon resumption.
[1075] The resetting feature is:
[1076] 1) DPU microcontroller.
[1077] 2) DPA program memory.
[1078] 3) BU program memory.
[1079] 4) SB program memory.
[1080] 5) Address generator.
[1081] 6) BU read buffer.
[1082] 7) GU Iterator (1 bit)
[1083] Reset control while suspended
[1084]
[1085]
[1086] Event counter p_EventCnt
[1087] A simple counter, counted by the DPU clock (not counting during p_stall). The counter is preset via configuration. An event signal is sent to the scalar unit whenever the counter is empty. This counter is readable via the configuration bus.
[1088] Suspend and event increment register p_SuspEventInc
[1089] The lower half (15..0) is used to increment the pending counter, both during its normal decrement and simultaneously. The higher half (31..16) also increments the event counter simultaneously.
[1090] Figure 33 A method 3300 according to an embodiment of the present invention is shown.
[1091] Method 3300 can begin with step 3310, selecting a source pixel from the group of source pixels. The selected source pixel will be referred to as the "source pixel".
[1092] Step 3310 may be followed by step 3320, which performs a warp calculation process for each source pixel in the group of source pixels, including:
[1093] 1) Calculate (3321) or receive warp parameters for the selected source pixel. The warp parameters may include a first weight and a second weight (Wx, Wy) and coordinates (x, y) of a given target pixel that should be processed during warp calculation. The first weight and the second weight are received by a first set of processing units (DPUs) in the processing unit array (DPA).
[1094] 2) Request (3322) neighboring target pixels (including the given target pixel) from a memory unit such as a collection unit. The collection unit can receive 4 coordinates in various operating modes and convert them into 16 target pixels - four groups of neighboring target pixels.
[1095] 3) The second set of processing units receives (3323) the adjacent target pixels associated with the source pixel.
[1096] 4) The second set of processing units calculates the warping result (3324) in response to the values of adjacent target pixels and a pair of weights; and provides the warping result to the memory module.
[1097] Steps 3321, 3322, 3323, and 3324 can be executed in a pipeline manner.
[1098] Step 3320 is followed by step 3330, which checks whether the warp is calculated for all source pixels in the group. If not, the warp calculation ends.
[1099] Step 3326 may include relaying the values of some adjacent target pixels between the processing units of the second group.
[1100] Any reference to any of the terms “comprise,” “comprises,” “comprising,” “including,” “may include,” and “includes” may apply to the terms “consists,” “consisting,” and “and consisting essentially of.” For example, any method describing steps may include more steps than those shown in the figure, only those shown in the figure, or steps that are essentially only shown in the figure. This also applies to components of a device, processor, or system, and instructions stored in any non-transitory computer-readable storage medium.
[1101] The invention can also be implemented in a computer program for running on a computer system, the computer program including at least a code portion for performing steps of the method according to the invention when running on a programmable device (such as a computer system), or for enabling the programmable device to perform the functions of a device or system according to the invention. The computer program can cause a storage system to allocate hard disk drives to a group of hard disk drives.
[1102] A computer program is a list of instructions such as a particular application and / or operating system. A computer program may include, for example, one or more of the following: subroutines, functions, procedures, object methods, object implementations, executable applications, applets, service applets, source code, object code, shared libraries / dynamically loaded libraries, and / or other sequences of instructions designed to be executed on a computer system.
[1103] Computer programs may be internally stored on non-transitory computer-readable media. All or some computer programs may be set on computer-readable media that are permanently, removably, or remotely coupled to an information processing system. Computer-readable media may include, for example, but not limited to, any number of the following: magnetic storage media, including hard disk and magnetic tape storage media; optical storage media, such as optical disc media (e.g., CD-ROM, CD-R, etc.) and digital video disc storage media; non-volatile memory storage media, including semiconductor-based memory cells, such as flash memory, EEPROM, EPROM, ROM; ferromagnetic digital memory; MRAM; and volatile memory media, including registers, buffers or caches, main memory, RAM, etc.
[1104] A computer process typically includes an executing (running) program or a part of a program, current program values and status information, and resources used by the operating system to manage the execution of the process. The operating system (OS) is software that manages shared computer resources and provides programmers with interfaces for accessing these resources. The operating system processes system data and user input, and responds by allocating and managing tasks and internal system resources that serve users and programs on the system.
[1105] A computer system may include, for example, at least one processing unit, associated memory, and multiple input / output (I / O) devices. When a computer program is executed, the computer system processes information according to the computer program and generates the obtained output information via the I / O devices.
[1106] In the foregoing description, the invention has been described with reference to specific examples of embodiments thereof. However, it will be apparent that various modifications and changes may be made therein without departing from the broader spirit and scope of the invention as set forth in the appended claims.
[1107] Furthermore, the terms “front,” “rear,” “top,” “bottom,” etc., used in the specification and claims, if applicable, are for descriptive purposes and not necessarily for describing permanent relative positions. It should be understood that such terms are interchangeable where appropriate, enabling, for example, the embodiments of the invention described herein to operate in directions other than those shown or otherwise described herein.
[1108] The connections discussed herein can be any type of connection suitable for transmitting signals between relevant nodes, units, or devices, for example, via intermediate devices. Therefore, unless implied or otherwise stated, connections can be, for example, direct or indirect connections. Connections can be illustrated or described with reference to a single connection, multiple connections, unidirectional connections, or bidirectional connections. However, different embodiments can vary the implementation of the connection. For example, a single unidirectional connection can be used instead of a bidirectional connection, and vice versa. Moreover, multiple connections can be replaced by a single connection that transmits multiple signals serially or in a time-division multiplexing manner. Similarly, a single connection carrying multiple signals can be separated into various different connections carrying subgroups of these signals. Therefore, there are many options for signal transmission.
[1109] Although specific conductivity types or potential polarities have been described in the embodiments, it should be understood that conductivity types and potential polarities can be reversed.
[1110] Each signal described herein can be designed as either positive or negative logic. In the case of a negative logic signal, the signal is active low when the logical true state corresponds to logic level zero. In the case of a positive logic signal, the signal is active high when the logical true state corresponds to logic level zero. Note that any signal described herein can be designed as either a negative or positive logic signal. Therefore, in alternative embodiments, those signals described as positive logic signals can be implemented as negative logic signals, and those signals described as negative logic signals can be implemented as positive logic signals.
[1111] Furthermore, when it comes to presenting a signal, status bit, or similar device as being logically true or logically false, this document uses the terms “assert” or “set” and “negate” (or “invalidate” or “clear”) for its logically true or logically false state, respectively. If the logically true state is logic level 1, then the logically false state is logic level 0. If the logically true state is logic level 0, then the logically false state is logic level 1.
[1112] Those skilled in the art will recognize that the boundaries between logic blocks are merely illustrative, and alternative embodiments may combine logic blocks or circuit elements, or impose alternative decompositions of functionality on various logic blocks or circuit elements. Therefore, it should be understood that the architecture described herein is merely exemplary, and many other architectures that implement the same functionality can actually be implemented.
[1113] Any arrangement of components that perform the same function is effectively “associated” to achieve the desired function. Therefore, any two components combined in this paper to achieve a specific function can be considered “associated” with each other to achieve the desired function, regardless of the architecture or intermediate components. Similarly, any two such associated components can also be considered “operably connected” or “operably coupled” to each other to achieve the desired function.
[1114] Furthermore, those skilled in the art will recognize that the boundaries between the above operations are merely illustrative. Multiple operations may be combined into a single operation, a single operation may be distributed among additional operations, and operations may be performed with at least partial overlap in time. Additionally, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be varied in various other embodiments.
[1115] For example, in one embodiment, the illustrated example can be implemented as circuitry located on a single integrated circuit or within the same device. Alternatively, the example can be implemented as any number of separate integrated circuits or separate devices interconnected in a suitable manner.
[1116] For example, an example or part thereof may be implemented as a software or code representation of a physical circuit, or may be implemented as a software or code representation of a logic representation that can be converted into a physical circuit, such as in any suitable type of hardware description language.
[1117] Furthermore, the present invention is not limited to physical devices or units implemented in non-programmable hardware, but can also be applied to programmable devices or units that can perform desired device functions by operating according to appropriate program code, such as mainframes, minicomputers, servers, workstations, personal computers, notebooks, personal digital assistants, video games, automobiles and other embedded systems, cellular phones and various other wireless devices, which are generally referred to as "computer systems" in this application.
[1118] However, other modifications, changes, and substitutions are also possible. Therefore, the specification and drawings are considered illustrative rather than restrictive.
[1119] In the claims, any reference numerals placed between parentheses should not be construed as limiting the claims. The word “comprising” does not exclude the presence of other elements or steps besides those listed in the claims. Furthermore, the terms “a” or “an” as used herein are defined as one or more. Moreover, the use of introductory phrases such as “at least one” and “one or more” in the claims should not be construed as implying that the introduction of another claim element by the indefinite article “a” or “an” limits any particular claim containing that introductory claim element to an invention containing only one of that element, even when the same claim includes the introductory phrase “one or more” or “at least one” and indefinite articles such as “a” or “an”. This also applies to the use of definite articles. Unless otherwise stated, terms such as “first” and “second” are used to arbitrarily distinguish the elements described by such terms. Therefore, these terms are not necessarily intended to indicate the time or other priority of these elements. The indisputable fact that certain measures are referenced in mutually different claims does not indicate that a combination of these measures cannot be used advantageously.
[1120] While certain features of the invention have been shown and described herein, many modifications, substitutions, alterations, and equivalents will occur to those skilled in the art. Therefore, it should be understood that the appended claims are intended to cover all such modifications and alterations falling within the true spirit of the invention.
Claims
1. A method for calculating a distortion result, the method comprising: A warp calculation process is performed for each target pixel in a set of target pixels, the warp calculation process including: The processing units in the first group of the array of processing units receive a pair of weights, including a first weight and a second weight associated with the target pixel. The processing units in the second group of the array receive the values of neighboring source pixels associated with the target pixel; The distortion result is calculated using the second set of values based on the neighboring source pixels and the pair of weights; and The distortion result is provided to the memory module; The calculation of the distortion result includes: The first processing unit of the second group calculates the first difference between the first pair of adjacent source pixels and the second difference between the second pair of adjacent source pixels. The first difference is provided to the second processing unit of the second group; and The second difference is provided to the third processing unit of the second group; And wherein each processing unit in the first group of processing units and the second group of processing units includes a core, and the cores of the plurality of processing units are coupled to each other through a configurable network.
2. The method according to claim 1, wherein, The calculation of the distortion result includes relaying the value of one or more of the neighboring source pixels between the processing units in the second group.
3. The method according to claim 1, wherein, The calculation of the distortion result includes relaying the intermediate result calculated by the second group and the value of one or more of the neighboring source pixels between the processing units of the second group.
4. The method according to claim 1, wherein, The calculation of the distortion result also includes: The fourth processing unit of the second group calculates a first modified weight in response to the first weight and calculates a second modified weight in response to the second weight; The first modified weight is provided from the fourth processing unit to the second processing unit of the second group; The second processing unit of the second group calculates the first intermediate result based on the first difference, the first neighboring source pixel, and the first modified weight.
5. The method according to claim 4, wherein, The calculation of the distortion result also includes: The second difference is provided from the third processing unit of the second group to the sixth processing unit of the second group; The fifth processing unit of the second group provides a second adjacent source pixel to the sixth processing unit of the second group; and The sixth processing unit of the second group calculates the second intermediate result based on the second difference, the second neighboring source pixel, and the first modified weight.
6. The method according to claim 5, wherein, The calculation of the distortion result also includes: The second intermediate result is provided from the sixth processing unit of the second group to the seventh processing unit of the second group; The first intermediate result is provided from the second processing unit of the second group to the seventh processing unit of the second group; and The seventh processing unit of the second group calculates the third intermediate result based on the first intermediate result and the second intermediate result.
7. The method according to claim 6, wherein, The calculation of the distortion result also includes: The third intermediate result is provided from the seventh processing unit of the second group to the eighth processing unit of the second group; The second intermediate result is provided from the sixth processing unit of the second group to the ninth processing unit of the second group; The second intermediate result is provided from the ninth processing unit of the second group to the eighth processing unit of the second group; The second modified weight is provided from the fourth processing unit of the second group to the eighth processing unit of the second group; and The distortion result is calculated by the eighth processing unit of the second group based on the second intermediate result, the third intermediate result, and the second modified weight.
8. The method of claim 1, further comprising performing multiple twist calculation processes associated with a target pixel subgroup in parallel via the array.
9. The method of claim 8, comprising: The collection unit extracts neighboring source pixels associated with each target pixel in the pixel subgroup in parallel from the collection unit; wherein the collection unit includes a group-associative buffer and is arranged to access a memory module comprising multiple independently accessible storage blocks.
10. The method of claim 9, comprising: For each target pixel in the pixel subgroup, a first distortion parameter and a second distortion parameter are received; wherein the first distortion parameter and the second distortion parameter include a first weight and a second weight, as well as position information indicating the position of neighboring source pixels associated with the target pixel.
11. The method of claim 10, further comprising providing the collection unit with the location information of each target pixel in the pixel subgroup.
12. The method of claim 11, further comprising converting the location information into addresses of the neighboring source pixels through the collection unit.
13. The method of claim 9, comprising: The first and second distortion parameters are calculated by the third set of processing units in the array for each target pixel in the pixel subgroup; wherein the first and second distortion parameters include the first weight and the second weight, as well as position information indicating the position of the neighboring source pixels associated with the target pixel.
14. The method of claim 13, wherein the first weight and the second weight are sensed from the third group to the first group.
15. The method of claim 14, further comprising providing the collection unit with the location information of each target pixel in the pixel subgroup.
16. The method of claim 15, further comprising converting the location information into addresses of the neighboring source pixels via the collection unit.
17. The method according to claim 1, further comprising: A set of absolute differences and SADs are calculated by the processing units of the first group in the array of processing units; wherein the set of SADs is associated with a source pixel and a target pixel subgroup; wherein each SAD is calculated based on a previously calculated SAD and based on the currently calculated absolute difference between another source pixel and a target pixel from the target pixel subgroup; and the best matching target pixel in the target pixel subgroup is determined by the processing units of the second group in the array in response to the value of the set of SADs.
18. The method according to claim 1, further comprising: The collection unit receives multiple requests for retrieving multiple requested data units via its input interface; stores multiple tags and multiple cached data units via a cache memory comprising multiple entries; wherein each tag is associated with a cached data unit and indicates a set of memory units in the memory module that are different from the cache memory and store the cached data unit; simultaneously compares the multiple tags with multiple requested memory group addresses via a comparator array to provide a comparison result; wherein each requested memory group address indicates a set of memory units in the memory module that store the requested data unit among the multiple requested data units; classifies the multiple requested data units into cached data units and uncached data units stored in the cache memory based on the comparison result via a controller; sends information about cached and uncached data units to a contention assessment unit; checks for the occurrence of at least one contention via the contention assessment unit; and requests any uncached data units from the memory module in a contention-free manner via an output interface.
19. The method of claim 1, further comprising operating a processing module including a data processor array; wherein, The operation includes processing data via data processors in the array; wherein each data processor unit of the plurality of data processors in the data processor array is directly coupled to one or more data processors in the data processor array, indirectly coupled to one or more other data processors in the data processor array, and relays data between the relay ports of the data processors using one or more relay channels of one or more data processors.
20. The method of claim 1, further comprising configuring an image processor including a plurality of configurable circuits and a plurality of microcontrollers; wherein, The plurality of configurable circuits include memory circuitry and a plurality of data processors; wherein the configuration of the image processors includes storing up to a limited number of configuration instructions in each configurable circuit; the plurality of configurable circuits are controlled by the plurality of microcontrollers by repeatedly providing selection information to the plurality of configurable circuits, the selection information being used by each configurable circuit to select a chosen configuration instruction from the limited number of configuration instructions.
21. The method of claim 1, further comprising operating an image processor, wherein operating the image processor includes providing an image processor including a data processor array and a memory module including a plurality of storage units; A buffer unit; a collection unit; and multiple microcontrollers; the data processor array is controlled via the multiple microcontrollers, a portion of the memory module, and the buffer unit; Data is retrieved from the memory module through the buffer unit; The data is sent to the data processor array through the buffer unit; The collection unit receives multiple requests to retrieve multiple requested data units from the memory module. The collection unit sends the plurality of requested data units to the data processor array.
22. The method according to claim 1, further comprising configuring an image processor, the image processor comprising a data processor array, a first microcontroller, a buffer unit, and a second microcontroller; wherein, The configuration of the image processor includes providing data processor configuration instructions to the data processors in the array during the data processor configuration process; providing buffer unit configuration instructions to the buffer unit during the buffer unit configuration process; and controlling the operation of the data processor by providing data processor selection information to the data processor through the first microcontroller. The data processor selects a data processor configuration instruction in response to the data processor selection information, and performs one or more data processing operations according to the selected data processor configuration instruction; the second microcontroller provides buffer unit selection information to the buffer unit to control the operation of the buffer unit; the buffer unit selects a selected buffer unit configuration instruction in response to at least a portion of the buffer unit selection information, and performs one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information is a portion of the size of the data processor configuration instruction.
23. An image processor configured to compute a warp result, the image processor being configured to perform a warp calculation process for each of a set of target pixels: in, The first group of processing units in the processing unit array of the image processor is configured to receive a pair of weights, including a first weight and a second weight associated with the target pixel, during the distortion calculation process; The second group of processing units in the array is configured as follows: During the distortion calculation process, the values of neighboring source pixels associated with the target pixel are received; The distortion result is calculated using the second set of values based on the neighboring source pixels and the pair of weights; and The distortion result is provided to the memory module; The first processing unit of the second group is configured as follows: Calculate the first difference between the first pair of adjacent source pixels and the second difference between the second pair of adjacent source pixels; The first difference is provided to the second processing unit of the second group; and The second difference is provided to the third processing unit of the second group; And wherein each of the first group of processing units and the second group of processing units includes a core, and the cores of the plurality of processing units are coupled to each other through a configurable network.
24. The image processor according to claim 23, wherein, The processing units in the second group are configured to relay the values of one or more of the adjacent source pixels between the processing units.
25. The image processor according to claim 23, wherein, The processor units of the second group are configured to relay intermediate results calculated by the second group and the values of one or more of the adjacent source pixels between the processing units of the second group.
26. The image processor according to claim 23, wherein: The fourth processing unit of the second group is configured to calculate a first modified weight in response to the first weight and to calculate a second modified weight in response to the second weight; And provide the first modified weight from the fourth processing unit to the second processing unit of the second group; as well as The second processing unit of the second group is configured to calculate a first intermediate result based on the first difference, the first neighboring source pixel, and the first modified weight.
27. The image processor according to claim 26, wherein: The third processing unit of the second group is configured to provide the second difference to the sixth processing unit of the second group; The fifth processing unit of the second group is configured to provide a second adjacent source pixel to the sixth processing unit of the second group; as well as The sixth processing unit of the second group is configured to calculate a second intermediate result based on the second difference, the second neighboring source pixel, and the first modified weight.
28. The image processor according to claim 27, wherein: The sixth processing unit of the second group is configured to provide the second intermediate result to the seventh processing unit of the second group; The second processing unit of the second group is configured to provide the first intermediate result to the seventh processing unit of the second group; as well as The seventh processing unit of the second group is configured to calculate a third intermediate result based on the first intermediate result and the second intermediate result.
29. The image processor according to claim 28, wherein: The seventh processing unit of the second group is configured to provide the third intermediate result to the eighth processing unit of the second group; The sixth processing unit of the second group is configured to provide the second intermediate result to the ninth processing unit of the second group; The ninth processing unit of the second group is configured to provide the second intermediate result to the eighth processing unit of the second group; The fourth processing unit of the second group is configured to provide the second modified weight to the eighth processing unit of the second group; as well as The eighth processing unit of the second group is configured to calculate the distortion result based on the second intermediate result, the third intermediate result, and the second modified weight.
30. The image processor of claim 23, wherein the array is configured to perform multiple warp calculation processes associated with a target pixel subgroup in parallel.
31. The image processor of claim 30, further comprising a collection unit, wherein the array is configured to extract, in parallel, neighboring source pixels associated with each target pixel in the pixel subgroup; wherein, The collection unit includes a set-associative cache and is arranged to access a memory module comprising multiple independently accessible memory banks.
32. The image processor according to claim 31, wherein: The array is configured to receive a first twist parameter and a second twist parameter for each target pixel in the pixel subgroup; wherein the first twist parameter and the second twist parameter include a first weight and a second weight, as well as position information indicating the position of neighboring source pixels associated with the target pixel.
33. The image processor according to claim 32, wherein, The array is configured to provide the collection unit with the location information of each target pixel in the pixel subgroup.
34. The image processor according to claim 32, wherein, The collection unit is configured to convert the location information into the addresses of the neighboring source pixels.
35. The image processor according to claim 34, wherein, The third set of processing units in the array is configured to calculate a first twist parameter and a second twist parameter for each target pixel in the pixel subgroup; wherein the first twist parameter and the second twist parameter include a first weight and a second weight, as well as position information indicating the position of the neighboring source pixel associated with the target pixel.
36. The image processor according to claim 35, wherein, The third group is configured to send the first weight and the second weight to the first group.
37. The image processor according to claim 36, wherein, The array is configured to provide the collection unit with the location information of each target pixel in the pixel subgroup.
38. The image processor according to claim 37, wherein, The collection unit is configured to convert the location information into the address of the neighboring source pixel.
39. The image processor according to claim 23, wherein, A first set of data processors in the array of data processors is configured to calculate a set of absolute differences and SADs; wherein the set of SADs is associated with a source pixel and a target pixel subgroup; wherein each SAD is calculated based on a previously calculated SAD and based on the currently calculated absolute difference between another source pixel and a target pixel belonging to the target pixel subgroup; and wherein a second set of data processors in the array is configured to calculate the best matching target pixel in the target pixel subgroup in response to the value of the set of SADs.
40. The image processor according to claim 23, comprising a cache memory, a collection unit, a comparator array, a controller, a contention assessment unit, and an output interface; wherein, The collection unit includes an input interface configured to receive multiple requests for retrieving multiple requested data units; The cache memory is configured to store multiple entries, multiple tags, and multiple cached data units; wherein each tag is associated with a cached data unit and indicates a set of memory units in the memory module that are different from the cache memory and store the cached data unit; The comparator array is configured to provide a comparison result between the plurality of tags and the plurality of requested memory group addresses; Each requested memory group address indicates a group of memory cells in the memory module that stores the requested data cells among the plurality of requested data cells; The controller classifies the plurality of requested data units into cached data units and uncached data units stored in the cache memory based on the comparison result; and sends information about the cached and uncached data units to the contention assessment unit. The contention assessment unit is configured to check for the occurrence of at least one contention. And the output interface therein requests any uncached data units from the memory module in a contention-free manner.
41. The image processor according to claim 23, wherein, Each data processor unit in the plurality of data processors of the data processor array is directly coupled to one or more data processors in the data processor array, indirectly coupled to one or more other data processors in the data processor array, and relays data between the relay ports of the data processors using one or more relay channels of one or more data processors.
42. The image processor according to claim 23, comprising a plurality of configurable circuits and a plurality of microcontrollers; wherein, The plurality of configurable circuits include memory circuitry and a plurality of data processors; wherein the configuration of the image processors includes storing up to a limited number of configuration instructions in each configurable circuit; the plurality of configurable circuits are controlled by the plurality of microcontrollers by repeatedly providing selection information to the plurality of configurable circuits, the selection information being used by each configurable circuit to select a chosen configuration instruction from the limited number of configuration instructions.
43. The image processor of claim 23, comprising a memory module including a plurality of storage units; a buffer unit; a collection unit; and a plurality of microcontrollers; controlling a data processor array via the plurality of microcontrollers, a portion of the memory module, and the buffer unit; Data is retrieved from the memory module through the buffer unit; The data is sent to the data processor array through the buffer unit; The collection unit receives multiple requests to retrieve multiple requested data units from the memory module. The collection unit sends the plurality of requested data units to the data processor array.
44. The image processor according to claim 23, comprising a first microcontroller, a buffer unit, and a second microcontroller; in, The data processors in the data processor array are configured to receive data processor configuration instructions during the data processor configuration process; The buffer unit is configured to receive a buffer unit configuration instruction during the buffer unit configuration process; The first microcontroller is configured to control the operation of the data processor by providing data processor selection information to the data processor; The data processor is configured to select a selected data processor configuration instruction in response to the data processor selection information, and to perform one or more data processing operations according to the selected data processor configuration instruction; The second microcontroller is configured to control the operation of the buffer unit by providing buffer unit selection information to the buffer unit; The buffer unit is configured to select a selected buffer unit configuration instruction in response to at least a portion of the buffer unit selection information, and to perform one or more buffer unit operations according to the selected buffer unit configuration instruction; wherein the size of the data processor selection information is a portion of the size of the data processor configuration instruction.
45. A non-transitory computer-readable medium storing instructions, which, when executed by at least one processor, are configured to cause at least one processor to perform a method, the method comprising: A warp calculation process is performed for each target pixel in a set of target pixels, the warp calculation process including: The processing units in the first group of the array of processing units receive a pair of weights, including a first weight and a second weight associated with the target pixel. The processing units in the second group of the array receive the values of neighboring source pixels associated with the target pixel; The distortion result is calculated using the second set of values based on the neighboring source pixels and the pair of weights; and The distortion result is provided to the memory module; The calculation of the distortion result includes: The first processing unit of the second group calculates the first difference between the first pair of adjacent source pixels and the second difference between the second pair of adjacent source pixels. The first difference is provided to the second processing unit of the second group; and The second difference is provided to the third processing unit of the second group; And wherein each processing unit in the first group of processing units and the second group of processing units includes a core, and the cores of the plurality of processing units are coupled to each other through a configurable network.
46. A method for calculating a distortion result, the method comprising: The first set of processing units in the array of processing units simultaneously receives the first weight and the second weight for each target pixel in the pixel subgroup. The collection unit is simultaneously provided with location information indicating the location of neighboring source pixels associated with each target pixel in the pixel subgroup; The array simultaneously receives, from the collecting unit, neighboring source pixels associated with each target pixel in the pixel subgroup; wherein different groups of the array receive neighboring source pixels associated with different target pixels in the pixel subgroup; and The distortion results associated with the different target pixels are calculated simultaneously using the different groups of the array; The calculation of the distortion result includes: The first processing unit of the different groups calculates the first difference between the first pair of adjacent source pixels and the second difference between the second pair of adjacent source pixels; The first difference is provided to the second processing unit of the different groups; and The second difference is provided to the third processing unit of the different groups; And wherein, each of the first group of processing units and the processing units in different groups includes a core, and the cores of the multiple processing units are coupled to each other through a configurable network.
47. The method of claim 46, comprising: For each target pixel in the pixel subgroup, a first distortion parameter and a second distortion parameter are received or calculated; wherein the first distortion parameter and the second distortion parameter include a first weight and a second weight, as well as position information indicating the position of neighboring source pixels associated with the target pixel.
48. An image processor configured to compute a distortion result, the image processor comprising: An array of processing units, the array of processing units being configured as follows: The first set of processing units in the array of processing units simultaneously receives the first weight and the second weight for each target pixel in the pixel subgroup. The image processor's collection unit is simultaneously provided with location information indicating the location of neighboring source pixels associated with each target pixel in the pixel subgroup; The array simultaneously receives, from the collecting unit, neighboring source pixels associated with each target pixel in the pixel subgroup; wherein different groups of the array receive neighboring source pixels associated with different target pixels in the pixel subgroup; and The distortion results associated with the different target pixels are calculated simultaneously using the different groups of the array; The calculation of the distortion result includes: The first processing unit of the different groups calculates the first difference between the first pair of adjacent source pixels and the second difference between the second pair of adjacent source pixels; The first difference is provided to the second processing unit of the different groups; and The second difference is provided to the third processing unit of the different groups; And wherein, each of the first group of processing units and the processing units in different groups includes a core, and the cores of the multiple processing units are coupled to each other through a configurable network.
49. A non-transitory computer-readable medium storing instructions configured, when executed by at least one processor, to cause the at least one processor to perform a method, the method comprising: The first set of processing units in the array of processing units simultaneously receives the first weight and the second weight for each target pixel in the pixel subgroup. The collection unit is simultaneously provided with location information indicating the location of neighboring source pixels associated with each target pixel in the pixel subgroup; The array simultaneously receives, from the collecting unit, neighboring source pixels associated with each target pixel in the pixel subgroup; wherein different groups of the array receive neighboring source pixels associated with different target pixels in the pixel subgroup; and The distortion results associated with the different target pixels are calculated simultaneously using the different groups of the array; The calculation of the distortion result includes: The first processing unit of the different groups calculates the first difference between the first pair of adjacent source pixels and the second difference between the second pair of adjacent source pixels; The first difference is provided to the second processing unit of the different groups; and The second difference is provided to the third processing unit of the different groups; And wherein, each of the first group of processing units and the processing units in different groups includes a core, and the cores of the multiple processing units are coupled to each other through a configurable network.
50. A method for operating an image processor, the method comprising: Provides an image processor including an array of processing units; A memory module comprising multiple memory banks; Buffer unit; Collection unit; and multiple microcontrollers; receive a pair of weights through a first set of processing units in the processing unit array, the pair of weights including a first weight and a second weight associated with the target pixel; receive the values of neighboring source pixels associated with the target pixel through a second set of processing units in the array; The second set of processing units calculates the warping result based on the values of the adjacent source pixels and the pair of weights; and provides the warping result to the memory module; The processing unit array is controlled by the plurality of microcontrollers, a portion of the memory module, and the buffer unit; Data is retrieved from the memory module through the buffer unit; the data is sent to the processing unit array through the buffer unit; and multiple requests for retrieving multiple requested data units from the memory module are received through the collection unit. The collection unit sends the plurality of requested data units to the processing unit array. The calculation of the distortion result includes: The first processing unit of the second group calculates the first difference between the first pair of adjacent source pixels and the second difference between the second pair of adjacent source pixels. The first difference is provided to the second processing unit of the second group; The second difference is provided to the third processing unit of the second group; And wherein each processing unit in the first group of processing units and the second group of processing units includes a core, and the cores of the plurality of processing units are coupled to each other through a configurable network.
Citation Information
Patent Citations
Image processor and methods for processing an image
CN108140232A
Quality image warper
US6061477A