Sub-image streaming and processing

By employing an FPGA and DPU to separate payload data from headers and a GPU to process image sections, the system addresses CPU bottlenecks, enhancing image processing efficiency and reducing latency in high-bandwidth applications.

DE102025129653A1Pending Publication Date: 2026-01-29MELLANOX TECHNOLOGIES LTD(IL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025129653
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-29
Filing Date
2025-07-28
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing image processing systems face bottlenecks due to the need for central processing units (CPUs) to handle and copy entire images, leading to latency and inefficiencies in high-speed, high-quality processing scenarios, especially in applications requiring high bandwidth and low latency.

Method used

A system utilizing a field-programmable gate array (FPGA) performs physical-level processing on images, with a data processing unit (DPU) separating payload data from headers, and a graphics processing unit (GPU) processing only image sections, bypassing CPU intervention and enabling direct data transfer and content-level processing.

Benefits of technology

This approach accelerates image processing by reducing CPU intervention, optimizing workflow, and ensuring data integrity and reliability, while supporting high-bandwidth and low-latency applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The systems and methods described herein are used for distributed image processing by at least one data processing unit (DPU) and at least one graphics processing unit (GPU), possibly in association with a field-programmable gate array (FPGA). For example, the FPGA can be used to perform physical-level processing on images acquired by the image sensor or from a simulation, and can provide a media stream to the DPU. The DPU can then provide the GPU with user data only from image segments within the media stream, allowing the GPU to perform content-level processing only on those image segments.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] At least one embodiment relates to image processing for images of a media stream. BACKGROUND

[0002] Video compression can be used to provide reduced media streams while preserving some detail of the underlying video's content. Such media streams can be part of various streaming technologies that extend beyond traditional broadcast markets. For example, Ethernet and other network technologies can contribute to developments in media streaming technologies. However, different applications may have diverse requirements regarding their media streams. For instance, educational online learning platforms can use media streaming for lectures and interactive sessions to make education accessible worldwide. Healthcare applications, such as telemedicine, enable media streaming for consultations and to facilitate remote diagnosis and treatment.Furthermore, gaming applications can leverage streaming platforms to revolutionize how games are played and viewed. In some or all of these applications, Ethernet can play a role in providing stable, high-speed connectivity, which can be crucial for the success of media streaming. For example, there may be requirements for high bandwidth and low latency that can be critical for the aforementioned and other applications. Additionally, with developments in virtual reality (VR) and augmented reality (AR), immersive media streams for entertainment and training are occupying significant bandwidth, along with media streams for smart cities in the form of traffic management and public safety.In one example, efficient video compression and transmission, along with advances in data storage and processing technologies, have made it possible to reliably stream high-quality content over the internet. However, processing is still performed on a large volume of images, which can cause latency in high-speed, high-quality processing situations. SUMMARY

[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described here that may or may not fall within the scope of the claims.

[0004] The systems and methods described herein are used for distributed image processing by at least one data processing unit (DPU) and at least one graphics processing unit (GPU), possibly in association with a field-programmable gate array (FPGA). For example, the FPGA can be used to perform physical-level processing on images acquired by the image sensor or from a simulation, and can provide a media stream to the DPU. The DPU can then provide the GPU with payload data only from image segments within the media stream, allowing the GPU to perform content-level processing only on those image segments.

[0005] Each feature of an aspect or embodiment can be applied to other aspects or embodiments in any suitable combination. In particular, each feature of a process aspect or embodiment can be applied to a device aspect or embodiment, and vice versa. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates a system for separating physical-level processing of images from content-level processing only for image sections in at least one embodiment; Fig. Figure 2 illustrates aspects of a system for providing sequence numbers for user data that represent only image sections of images, in order to enable processing of only the image sections, according to at least one embodiment; Fig. Figure 3 illustrates further aspects of a system for arranging user data representing only image sections in a shared buffer to enable processing of only the image sections, according to at least one embodiment; Fig. Figure 4 illustrates computer and processor aspects of a system for separating image sections from images for processing only image sections in at least one embodiment; Fig. Figure 5 illustrates a process flow for a system for separating physical-level processing for images from content-level processing only for image sections in at least one embodiment; Fig. Figure 6 illustrates a further process flow for a system for providing sequence numbers for user data that only represent image excerpts; and Fig. Figure 7 illustrates a further process flow for a system for arranging user data representing only image sections in a shared buffer in at least one embodiment. DETAILED DESCRIPTION

[0006] Fig. Figure 1 illustrates a system 100 for separating physical-level processing of images from content-level processing of image fragments only, in at least one embodiment. The system 100 provides multi-instance integration for image processing of images for a media stream by encompassing the complexity of data acquisition and processing associated with multiple sensor instances, such as a multi-sensor array 102. As used herein, the media stream, the images, and the image fragments can be part of a video and can be individual frames or segments thereof for the video. Furthermore, as used herein, the images can be generated from an image sensor using the multi-sensor array 102, from stimulation sensors, or from simulated data.

[0007] In one example, the stimulation sensors can be a combination of one or more magnetic and radio sensors capable of providing information to generate an image. For instance, System 100 can be a magnetic resonance imaging (MRI) machine or a computed tomography (CT) machine. System 100 also provides advanced image processing through specialized handling of both static and partial images, as well as dummy sensor compatibility to increase versatility through clever interface design with dummy sensors that act as agile transmitters. In one example, such use of dummy sensors can support a wide range of experimental, simulation, and calibration setups to enable the advanced image processing described here.

[0008] Furthermore, the system 100 comprises a direct transfer of image or sensor data 104 from a multisensor array 102, associated with a field-programmable gate array (FPGA) 106, to a graphics processing unit (GPU) 108 without requiring processing by a central processing unit (CPU) 118 of a host 120, also referred to herein as the host machine. The present system 100 can also support the separation of physical layer processing (PLP) 112 for images 116, which can be performed in the FPGA 106, from content layer processing (CLP) 114. PLP 112 can include processing associated with symbols or a floating-point representation of the images, while CLP 114 can include processing associated with pixels or a metadata representation of the images.

[0009] Furthermore, the CLP 114 in the GPU 108 can only be performed for image sections or sub-images 116A of the images 116. Furthermore, the image sections 116A, although illustrated in one image, can differ in different images 116 and can also comprise different parts of each of the images 116. For example, the image sections 116A can be a region of interest (ROI), an object, or an area of ​​motion relative to other areas in the images.While a media stream 126 may contain packets (“pkt”) furthermore payload (“P”) and an associated header (“H”), which can be provided for the entire images 116, the image sections 116A can be presented as selected payload 116B, which are separated from the packets by a data processing unit (DPU) 122 and provided by the DPU 122 to the GPU 108 for processing.

[0010] This approach avoids bottlenecks that might otherwise require a CPU 118 to intervene in many aspects of image processing. For example, it might otherwise be necessary for the CPU 118 to receive a media stream 126, and it might be necessary for it to copy entire images into the GPU memory 130. The approaches described herein, which may at least include the DPU 122 presenting payload data 116B to the GPU 108 for processing, corresponding to the image sections 116A, accelerate image processing and analysis in the system 100. In at least one embodiment, the present system 100, as part of the present separation of the PLP 112 from the CLP 114, supports header data splitting (HDS) to intelligently delegate the payload data 116B to a GPU 108, with the header sections 116C being delegated to the CPU 118 and stored in a memory 138 of a host 120.This further accelerates and improves image processing, but also reduces the additional organizational effort of the system.

[0011] In at least one embodiment, the system 100 for user data from a DPU 122 of a network interface card (NIC) 124 also includes direct packet placement (DPP) which is directed towards the use of unique identifiers, such as sequence numbers, as with respect to at least Fig. 2 further described. This encompassed approach ensures that the arrival sequence of a media stream 126 is irrelevant and does not depend on it for the GPU. Instead, the sequence numbers can be used to provide flexibility in data handling and optimization of a workflow 128 associated with image processing to be performed by a GPU 108. In at least one embodiment, the system 100 also supports zero-cost duplication by using DPP to provide handling of redundant data streams without additional resource expenditure. The duplication process cleverly writes directly to a buffer or other memory 130 associated with the GPU 108, which may be designated for specific payload data 116B, with respect to at least Fig. Section 3 is described further. This ensures data integrity and reliability for System 100.

[0012] Furthermore, the System 100 can operate in a multi-threaded environment and efficiently manage still images or partial images. For example, the System 100 can receive and process sensor data 104 from simulated dummy sensors instead of, or in addition to, the multi-sensor array 102. In an experimental, simulation, or calibration setup, it is possible to simulate the acquisition of sensor data 104 instead of acquiring it from a multi-sensor array 102. It is also possible to provide a simulated version of the sensor data 104 to which the PLP 112 is applied to correspond to an intended experiment, simulation, or calibration, applying the remaining features for using image sections using at least the configurations with DPU 122 and GPU 108 described herein.In the experimental setup, simulation setup, or calibration setup, an image sensor with the multisensor array 102 may not be required. Instead, an image (or sensor data 104) can be simulated by the FPGA 106 or another DPU. For example, in the experimental setup, simulation setup, or calibration setup, a DPU for generating the simulated sensor data 104 can communicate with the illustrated DPU 122, which, together with the GPU 108 and the host 120, can simulate a compute node for performing all aspects of the live application.

[0013] In at least one embodiment, the system 100 can include an image sensor as part of the multisensor array 102. The image sensor can be associated with the FPGA 106. In one example, the image sensor and the FPGA 106 are communicatively coupled within a camera module. Separately, the FPGA 106 can be coupled to a DPU 122 or a NIC 124. In one example, the FPGA 106 can be coupled to the DPU 122 via an Ethernet connection. However, it is also possible to provide a Peripheral Component Interconnect Express (PCIe) standard intermediary connection or bus between these components. The DPU 122 can, in turn, be coupled to the GPU 108 via a separate PCIe bus. Although illustrated as part of different cards, the DPU 122 and the GPU 108 can be part of a single card and can communicate via the PCIe bus of that single card.

[0014] In at least one embodiment, the single card, comprising a DPU 122 and a GPU 108, can be configured to be self-hosted. For example, the single card provides direct access between the DPU 122 and the GPU 108, enabling the DPU 122 to send payload data from the image sections 116A directly to the GPU 108 without host intervention. The GPU 108 can process the image sections 116A, in part, based on application requirements from at least one application. For example, the application can be associated with a domain-specific algorithm that can be used to perform specific CLP 114 operations on the GPU 108. In one example, the GPU 108 is enabled to perform processes on the image sections 116A that may be based, in part, on different protocols and encapsulation methods.The various protocols can include Hypertext Transfer Protocol (HTTP) Live Streaming (HLS), which can be used to deliver live and on-demand content over the internet. Other protocols can include Real-Time Messaging Protocol (RTMP), which can be used for high-performance transmission of audio, video, and data between Adobe® Flash® platform technologies, and MPEG-DASH®, which offers adaptive streaming by adjusting the quality of video streams in real time based on network conditions.

[0015] The encapsulation methods affect data integrity and transmission efficiency and can include MPEG® Transport Stream (MTS), which can ensure data integrity in a media stream 126, but can also ensure data integrity for error-prone transmission media. Furthermore, RTP, or Real-Time Protocol, can be used to deliver audio and video over networks, and WebRTC® can be used to enable real-time communication directly in web browsers. Additionally, the various streaming aspects can be enabled in the GPU 108 and can serve different use cases, including applications requiring broad compatibility and adaptive streaming capabilities, for which HLS can be used. In low-latency streaming, which can be crucial for live broadcasts, RTMP can be used.MPEG-DASH can be used when applications require flexibility and efficiency in a heterogeneous network environment.

[0016] In one example, if System 100 is part of an MRI machine, the MRI machine can be implemented as a low-field MRI. The low-field MRI can be associated only with magnets and its radio frequency (RF) electronics to acquire sensor data represented by symbols or floating-point values. These symbols or floating-point values ​​can be an analog representation of images to be subsequently rendered by a GPU. System 100 for distributed image processing enables an FPGA 106 to perform PLP on the symbols or floating-point representations intended for use in medical diagnostics and allows a GPU 108 and a DPU 122 to perform CLP on a media stream containing only image segments.Thus, the GPU and DPU combination can be located away from the low-field MRI machine and does not need to be located together with it to enable medical diagnostics.

[0017] Furthermore, the FPGA is configured to perform a PLP 112 on images 116 using the sensor data 104, and is configured to provide a media stream 126, representing a workflow 128, from the image sensor of the multisensor array 102. One result of the PLP 112 is that the FPGA 106 can provide payload data (“P”) with header regions (“H”), which are associated with an arrival sequence of the media stream 126, to a NIC 124 with the DPU 122. The payload data and header regions of the media stream 126 can represent the sensor data 104 of images 116 that have been processed by the PLP 112. The DPU 122 can be configured to perform an arrangement of the image sections 116A from the images 116 by arranging the payload 116B format in the media stream 126.For example, the DPU 122 can access a memory 130 associated with the GPU 108 and can place only image snippets 116A from the payload 116B format into the memory 130. Furthermore, the arrangement performed by the DPU 122 can be partially based on information about image snippets, such as from an application 134 running the CPU 118. For example, the CPU 118 uses input from an application 134 to provide information 216 about one or more image snippets to be provided to a GPU 108. However, the CPU 118 can also use the header regions to provide such information.

[0018] In at least one embodiment, the GPU 108 can perform content-level processing on only the image sections 116A from the images 116 using the payload data 116B, which is located in an associated buffer memory 130, and without further CPU intervention. Furthermore, in one example, the PLP 112 can include one or more analog-to-digital conversion (ADC), noise estimation, or time-stamping operations. In another example, the CLP 114 includes one or more pattern recognition, object recognition, feature extraction, feature characterization, or image segmentation operations. Therefore, it is evident that the physical-level processing in this case can include one or more operations that ignore the content of the images or that are performed only with consideration of raw pixel data associated with the images.In contrast, it is obvious that the content-level processing in this case comprises one or more operations designed to take into account the content of the images, or which are performed on raw pixel data with due consideration of the content within the images.

[0019] Furthermore, the PLP 112 can be free of business logic or can be independent of or agnostic to an application request from a host 120. The application request can originate from an application that wants to use one or more of the image snippets 116A, but the PLP 112 may not need to be aware of this request. Additionally, a GPU kernel 132 can be associated with the GPU 108 and the DPU 122. The GPU kernel 132 can have an interface to the DPU 122 to specify to the DPU 122 only those image snippets 116A from the payload 116B format that it wants to receive and that are to be subjected to the CLP 114. In one example, the GPU kernel can operate using instruction scripts from memory. Therefore, the CPU or the GPU can have application knowledge about one or more image snippets that the GPU should receive from the DPU for image processing according to the application requirements.

[0020] A GPU kernel 132 can function, in part, by having its instruction scripts executed on the GPU 108 to support a range of host kernels associated with the CPU 118 of a host 120. The GPU kernel 132 can be executed many times and can be run in parallel by different threads on the GPU 108. For example, each thread can be assigned a unique identifier or index, which is used for calculating memory addresses and for control decisions. Furthermore, kernel calls associated with a GPU kernel 132 can be executed by various circuits that form multiprocessor cores within the GPU 108. These circuits enable the execution of the different threads. These different threads can be scheduled and can be used to perform image processing for streaming applications.

[0021] In Fig. In this example, system 100 is configured such that GPU 108 can communicate with DPU 122 using the PCIe bus to receive only the image snippets 116A. GPU 108 can perform CLP 114 without intervention from CPU 118. For example, no further directive from CPU 118 to DPU 122 containing the image snippets 116A is required. In another example, host CPU 120's CPU 118 can simply instruct or inform DPU 122 which of the image snippets 116A are relevant to an application's request. This can be based partly on predefined information about the image snippets provided by application 134. However, it can also include the header regions 116C provided by NIC 124. It is possible that the CPU 118 will not provide any further intervention for the GPU 108.The predetermined information can also be provided to a GPU 108 222 to enable the GPU 108 to instruct or inform the DPU 122 which of the image sections 116A are relevant to the request of an application.

[0022] In addition, the DPU 122 can also perform its arrangement of the image segments 116A using the payload 116B in the memory 130 associated with the GPU 108. This can be partially based on a buffer that forms the memory 130 associated with the GPU 108 and is dedicated to the image segments 116A (using the payload 116B). Furthermore, the DPU 122 can also perform the arrangement partially based on providing at least one identifier associated with the image segments 116A (or assigned to each of the payloads in the payload 116B) to, for example, identify each payload 116B as belonging to the same image segments 116A. The GPU 108 can then access the buffer to perform the CLP 114 using the image segments 116A from the buffer and the at least one identifier.

[0023] In at least one embodiment, the media stream 126 can include the header and payload. In one example, one or more PCIe buses can include transactions for the media stream 126 between the FPGA and the DPU, and for payload between the DPU and the GPU. For example, the transactions can include payload representing image segments, which are transferred from the DPU to the GPU's buffer. The transactions can include headers, which are transferred from the DPU to a host's buffer and its associated CPU. In one example, the CPU or the host monitors for the arrival of the complete images (which can include all payload and headers of a complete video frame). In another example, the CPU or the host can trigger the GPU to perform processing on the image segments.

[0024] With regard to experimental, simulation, and calibration setups, the System 100 can be used for various tests by causing the multisensor array 102 to simulate sensor data 104. Thus, the sensor data 104 may not represent an object actually detected by the multisensor array 102, but can be simulated data for testing aspects of image processing using the System 100. For example, a constant bandwidth test can be performed in a multisensor configuration within the System 100. This test can involve simulating multiple sensors from a multisensor array 102. Each of the simulated sensors can maintain a constant bandwidth of 5 Gbit / s. This test can measure power consumption and CPU utilization while ensuring that no packet loss or sender delay occurs for up to 30 sensors.In another example, System 100 can be used to perform a multi-sensor test at full wire speed (FWS). In this test, each sensor of the multi-sensor array 102 can be configured to transmit at its maximum capacity to achieve full wire speed. Furthermore, another test enabled by System 100 can be to determine if a single sensor is unable to reach FWS and to indicate this.

[0025] Another test supported by System 100 is a single-sensor test with an increasing frame rate (FPS - frames per second). In this test, a single sensor of the Multisensor Array 102 can increase the associated FPS. This increase in FPS can, in turn, progressively increase bandwidth usage. To investigate the impact on system resources and data transmission, such a simulation can be performed using the Multisensor Array 102 and the approaches described here for separating physical-level image processing from content-level processing, which can only be performed on image segments. A further test can be an unrestricted single-sensor stability test, which can serve as a long-term stability test using System 100.In this test, a single sensor of the multisensor array 102 can operate at unrestricted speed, while the remaining components perform separation, physical-layer processing, and content-layer processing. System 100 can be monitored for line bandwidth, while the GPU and DPU can be monitored for power consumption, and the CPU can be monitored for core utilization over time. One or more of these tests can define different deployment options for the System 100, according to which one or more aspects of separation, physical-layer processing, and content-layer processing are performed in different environments with different configurations of these aspects. Fig. 1 to Fig. 4 can be carried out.

[0026] In at least one embodiment, at least one circuit of the GPU 108 can perform encoder functions as part of a video encoder. For example, an output of the GPU 108 can be a compressed or encoded media stream 136 for further use in an application 134 by the host 120 or another (and remote) host. In at least one embodiment, at least the CPU aspects of the system 100 can be performed in a data center. The GPU 108 can use standard video compression parameters to perform the video compression or encoding. For example, the GPU 108 can perform such video compression or encoding only on the image sections 116A to provide a compressed or encoded media stream that may be partially based on an H.264 standard, an MPEG2 standard, an AVC standard, an HEVC standard, a VP9 standard, an AV1 standard, or a VVC standard.

[0027] In one example, GPU 108 can be associated with a mode selection module that can be used to perform inter- or intra-mode encoding. Such mode selection can be carried out using this module. The mode selection can allow a selection of parameters that can be associated with available encoding parameters. The result of such mode selection is to provide a specific encoding for the image sections 116A. The mode selection can also allow a determination of how many bits the encoder is willing to sacrifice to hide and / or eliminate distortion that may be relevant for certain parts of the media selection.

[0028] As part of the encoding parameters, a Fourier transform or other related transformation can be performed on blocks within each frame to convert data into a frequency domain and to enable quantization or discarding of information based on selected frequencies. Transformation coefficients at lower frequencies can be quantized less aggressively than those at higher frequencies. Separately, motion estimation can be used to capture and encode motion across video frames. While all such options attempt to improve video compression, they can all serve a similar goal: to enable an encoder to compress video into smaller bitstreams by eliminating noise and artifacts, enabling at least more intensive motion estimation, and exploiting temporal and spatial redundancy.For example, transformation and quantization can be provided as additional parameters by a transformation and quantization module (T and Q module) of the encoder, which influence compression and / or encoding.

[0029] Given all these advantages, encoders can differ, in part, based on the selection of one or more suitable tools designed to enable certain aspects of bit saving. For example, the selection of suitable tools relates to the choice of coding parameters to allow the selection of areas (such as those provided by macroblocks (MBs)) within individual frames of each image section 116A that can be subjected to the compression or encoding described herein. This and other such approaches can be defined within the encoder as different modes that may require more or fewer bits to ensure a desired quality.A rate distortion optimization (RDO) module of the encoder can be associated with a mode selection module within it to address requirements by using RDO metrics, such as sum of squared errors (SSE) or sum of transformed differences (SATD), to determine costs associated with each choice made and to enable selection based on those costs.

[0030] Additional RDO metrics enable further mode selection, which benefits from evaluation using additional quality measures, including VMAF, SSIM, MS-SSIM, or PSNR. Distortion can be determined as a difference from the original image. In at least one embodiment, the GPU 108 supports an enhanced selection of at least the quality measures that can be used to perform video compression for the present image segments 116A. In one example, to provide the video compression or encoding described herein, the encoder can receive transformation coefficients or parameters, such as QPs. The RDO module can be operated to optimize an efficient representation for each point or block of an image segment, which may include segmentation, prediction modes, motion vectors (MV), or the QPs.

[0031] In at least one embodiment, the RDO output is used to select a mode as provided by the RDO module. Furthermore, an RDO can be limited to a single point for each block in each frame 116A and can be represented by a linear equation of the form R + λ * D, where λ (lambda) is a multiplier and an (R, D) pair can be used with the multiplier to minimize a combined R + D value. R can be associated with a bit rate and D can be associated with distortion, as it relates to the quality of the media. The RDO allows, for example, the determination of a ranking of candidate solutions using the linear equation to select one of the candidate solutions. Therefore, the lambda value can be associated with a range from 1 to minimized cost for the set of (R, D).R can be measured in bits and D can be a unit of quality, so the equation provides a measure in distortion units for each bit of a bit rate used in a video compression process.

[0032] To achieve a predetermined bit rate of R, a specific value of Lambda can be used. Furthermore, the selection of encoding parameters, which can include R, D, and Lambda values, allows the RDO to use different quality measures with the image sections 116A. In at least one embodiment, an encoder of the GPU 108 can implement H.264 encoding. The encoder can include modules in hardware or software, such as a prediction module, the T and Q modules, and an entropy encoding module. There can be other modules, such as an inverting module, a filtering module, a motion processing module (to support motion estimation and related aspects), and a previous frame or reference frame module. The video compression or encoding described herein may have no effect on a decoding process for a bitstream provided by the encoder. For example, the decoding process according to the H.264 decoding or another decoding relevant to the encoding format used to provide the output bitstream from the encoder and, in particular, the entropy encoding module.

[0033] A bitstream of individual frames, representing only the image sections 116A of images 116, can be compressed or encoded in the GPU 108 and can comprise different MBs or macroblocks. In at least one embodiment, different MB sizes can be supported in the encoder, in particular 8 × 8, 8 × 16, 16 × 8, 4 × 4, and 16 × 16. The MBs correspond to probabilistically displayed pixel data obtained at the position of the blocks. The prediction module can generate a prediction MB, which can be used to generate residual data reflecting data that undergoes quantization as part of video compression. Several prediction options can be associated with a prediction module, including an intra-prediction associated with previously encoded data derived from a current sequence, such as from each of the image sections 116A.Another option associated with a prediction module includes inter-prediction, which uses encoded data from other previously encoded frames, containing only the 116A image sections, as reference frames, such as those provided by the Previous Frame or Reference Frame module. These reference frames can appear before or after the current frame in the display sequence and can be associated with motion compensation, as with the Motion Process module, which uses previously encoded frames, such as those provided by the Previous Frame or Reference Frame module.

[0034] Another option associated with a prediction module involves using different prediction block sizes, available for both intra-prediction and inter-prediction options. Using different prediction block sizes can alter the accuracy associated with the predictions. A further option associated with a prediction module is using multiple frames during prediction, available in the inter-prediction option, to provide better prediction accuracy. Yet another option is to skip some or all of the remaining data, allowing the encoder to infer the accuracy of the data based partially on the prediction. One or more of these options represent coding parameters that can be applied to compress a section of image 116A.

[0035] In at least one embodiment, the intra-prediction can be based at least partially on spatial data within at least each of the image sections 116A. MBs generated as part of the intra-prediction can differ from the MBs of the individual frames of the image sections 116A. Residual data can be residual MBs generated by subtracting the prediction MB from a current MB. Depending on a mode selected by a mode selection module, which can be associated, for example, with the RDO module to perform the RDO, the residual MB can undergo transformation, quantization, and entropy coding in the provided modules of the GPU 108. Furthermore, in the encoder of the GPU 108, quantized data can be rescaled and inversely transformed in the invert module. An output of the invert module can be filtered in the prediction module and combined with the prediction MB.Motion estimation can be included in the motion processing module. The result can be a reconstructed MB or decoded single frames, which is / are provided to the previous single-frame or reference single-frame module for further predictions. In at least one embodiment, the use of one or more inter-prediction or intra-prediction parameters represents additional coding parameters that can be applied to compress a frame 116A for further communication or processing in a host 120 or a remote host.

[0036] Fig. Figure 2 illustrates aspects of a System 200 for providing sequence numbers for user data that represent only image sections, in order to enable processing of only the image sections, according to at least one embodiment. The aspects of System 200 in Fig. 2 can be all or some of the aspects already mentioned in relation to System 100. Fig. 1 described. For example, the system 200 can include an image sensor, which can be a multisensor array 102 or comprise one such array and can be associated with an FPGA 106. The image sensor can also be associated with a GPU 108 and a DPU 122. The FPGA 106 can provide images 116, captured by the image sensor and located in at least one media stream 126, to the DPU 122. The DPU 122 can separate header regions 204 from payload data associated with the images to provide separate payload data 206 P 11 to P 2N. The DPU 122 can provide sequence numbers 208 for the separate payload data 206, but only for those associated with image sections 116A of the images 116. Therefore, there may be payloads, such as payloads P 11 to P 14, which may not have sequence numbers 218, as they may represent something other than the image sections 116A of images 116.

[0037] In at least one embodiment, the separate header areas 204 are Real-Time Transport Protocol (RTP) header areas. The media stream 126 can be structured in the form of User Datagram Protocol (UDP) ports containing payload data and the RTP header areas. Separating the payload data and directing it to the GPU enables seamless reconstruction of multiple media streams 126 of payload data that can be received simultaneously by the FPGA. This seamless reconstruction allows the payload data from different media streams to be provided as a single stream, at least between the DPU and the GPU. Furthermore, this approach also supports redundancy and packet arrival reordering when required.

[0038] The DPU 122 can provide the separate payload 206 with the sequence numbers 212, which represent only the image sections 116A, for local access by the GPU 108 210. One or more of the DPU and the host can maintain information about a relationship between a sequence number and a header region, partly based on a relationship function 220. For example, the relationship function 220 can be used to determine the sequence numbers. For instance, the relationship function 220 can be a modulo function that extracts a number from a header region and applies a mathematical operation or function, such as the modulo function, to modify the number from the header region to provide a sequence number. Alternatively, the relationship function 220 can be a correlation table that keeps track of sequence numbers obtained from a mathematical operation that relate to a header region.Alternatively, the relationship function 220 is a transformation function that transforms information from the header into a sequence number.

[0039] Therefore, in at least one embodiment, associations between the header region and the sequence numbers can be used to correlate the sequence numbers used with the header region. As used here, local access between a processing unit and a memory can be provided by having such a processing unit and such memory located within the same host machine or card. Furthermore, as used here, local access between a processing unit and a memory can be provided via a PCIe bus instead of a network request (such as Ethernet). With respect to provision by the DPU 122, the separate payload data 206 with the sequence numbers 212 can be provided to a memory 202 associated with the GPU 108.The GPU 108 can access the separate user data P 15 to P 2N, which only contain image sections 116A of images 116 in . Fig. 1 represent, access, and process these, for example, using the sequence numbers S1 to SN. Furthermore, the system 200 can be configured such that at least one media stream 126 can comprise two media streams 126 and 126A. In one example, the two media streams 126 and 126A can be simultaneously acquired by the image sensor and simultaneously provided by the FPGA 106 to the DPU 122.

[0040] With respect to the two media streams 126 and 126A, their respective header areas H11 to H1N and H21 to H2N, which may be associated with them, can be separated from their respective payloads P11 to P1N and P21 to P2N. The respective header areas H11 to H2N can be made available for local access by a host machine 120 with a CPU 118. For example, the respective header areas H11 to H2N can reside in a local memory 104 associated with the CPU 118. Additionally, the header areas H11 to H1N for one of the media streams 126 can be made available in a way that allows separate access by the host machine 120 to these header areas relative to the other header areas H21 to H2N of the other media stream 126A.

[0041] Furthermore, the system 200 can be configured such that a CPU 118 can use information 214, which can be predetermined information provided to and by an application 134. For example, the predetermined information can be associated with different image sections 116A, based in part on the fact that the image sections 116A represent different ROIs. Furthermore, the CPU 118 can also use information from the respective header regions H11 to H1N and H21 to H2N in the local memory 104. All such information can be used to inform the DPU 122 about the separate payloads 206 216, which represent only the image sections 116A of the images 116, to be provided by the DPU 122 for access by the GPU 108.Furthermore, instead of the CPU 118, the GPU 108 can simply inform the DPU 122 about the image sections 216 that are to be received by the GPU 108 for content-level processing, or specify them. In at least one embodiment, the CPU 118 can cause predetermined information 222 to be provided to a GPU 108 to enable the GPU 108 to instruct or inform the DPU 122 which of the image sections 116A are relevant to the requirement of an application. However, in at least one embodiment, the GPU 108 does not need to receive any information about image sections 308, and instead, the GPU 108 can be limited to image processing capabilities and can inform the DPU 112 about the image sections 116A on which the latter must be able to perform its image processing intended for an application 134 and by the GPU 108.

[0042] When two media streams 126 and 126A are simultaneously provided by the FPGA, the payload data P15 to P1N, associated with a frame portion of the first 126 of the two media streams, can be assigned sequence numbers and made available for access by the GPU 108, along with additional payload data P21 to P2N, associated with a second frame portion of the second 126A of the two media streams. Furthermore, the additional payload data P21 to P2N can only represent additional frame portions of the second 126A of the two media streams, but are available together with the payload data P15 to P1N of the first 126 of the two media streams for contiguous access by the GPU.As used here, the contiguous working memory can refer to consecutive blocks of working memory 202, which can be used for the user data P 15 to P 1N and P 21 to P 2N from the different media streams and which can only represent the respective image sections 116A of these different media streams. Contiguous access, as used here, can be achieved such that access to different sequential user data, even if they originate from different media streams, can be maintained for contiguous access by mapping different buffers that store different parts of the user data.

[0043] In at least one embodiment, it is possible to obtain different payloads from different media streams, but to store the different payloads sequentially for contiguous access or in a contiguous buffer. For example, an application may be aware that part of a view may be covered by one camera and another part of the view may be covered by a different camera. Therefore, using the information about the image sections 308, the application can indicate that the payloads from different media streams are related. This indication can cause the GPU to obtain different payloads from the DPU and can cause the GPU to keep the different payloads in a contiguous access or in a contiguous buffer so that they can be combined for use in the application 134.Therefore, an application 134 can be configured to have knowledge of a layout of sensors associated with the multisensor array 102. The sensors can be different cameras. The application 134 can be configured to have knowledge of the field of view of each sensor of the multisensor array 102 and to be aware of the need to acquire a view from the different sensors in order to be able to stitch together image sections for the application 134.

[0044] Therefore, in an example, the DPU 122 can store header areas for initial payload data of the first of the concurrent media streams and additional header areas for additional payload data of the second of the concurrent media streams in several different buffers that may reside in a host machine. These buffers, represented by the host machine's memory 138, differ from a shared buffer, represented by the memory 130 of a GPU 108. For example, the shared buffer is one of: local to the GPU, on a GPU card that includes the GPU, or on an accelerator card or a converged card that includes the GPU and the DPU, whereas the multiple buffers can be local to a CPU 118 of the host machine or the DPU, or reside within the host machine or the DPU.The DPU can arrange the payload and the additional payload belonging to the image sections and additional image sections in contiguous areas of the designated memory locations of the shared buffer. The GPU is then able to use the arrangement of the payload and additional payload to assemble the image sections and additional image sections for use by at least one application or for further processing by the GPU.

[0045] System 200 can include a CPU 118 of a host machine 120, which can be configured to use information 214 from the header regions H11 to H2N to instruct the GPU 108 to process the payload, which represents only the image sections of the images. However, the CPU 118 can use predetermined information from an application 134 to inform a DPU 216 about the image sections to be provided to a GPU 108 to be processed as the payload. For example, the information 214 can cause the DPU 122 to provide only the payload P15 to P1N and P21 to P2N, which relate to image sections 116A.

[0046] Sequence numbers 218 cannot be provided for the remaining payload. System 200 can also allow the header regions H11 to H2N to be received for local access using a CPU 118 of a host machine 120. The CPU 118 can enable the DPU 122 to provide the payload P15 to P1N and P21 to P2N, which represent only the image sections 116A, for local access by the GPU. Furthermore, the CPU 118 can enable the GPU 108 to process the payload P15 to P1N and P21 to P2N, which represent only the image sections, partially based on the sequence numbers 208, by making only this payload available from the DPU 122. Therefore, no intervention by the CPU 118 in the operation of the GPU 108 is necessary in this respect.In addition, the system 200 is configured in such a way that the DPU 122 can control the provision 210 of the separate payload data 206 via a data stream whose bit rate and burst size are associated with predictable workloads at a known consumption rate for the GPU 108.

[0047] Fig. Figure 3 illustrates further aspects of a system 300 for arranging user data representing only image sections in a shared buffer to enable processing of only the image sections, according to at least one embodiment. As with respect to Fig. 2. The aspects of System 300 can be found in Fig. 3. All or some of the same aspects that are already mentioned in relation to one or more of the systems 100 or 200. Fig. 1 or Fig. 2 described. For example, the system 300 can include an image sensor, which can be a multisensor array 102 or comprise one such array and which can be associated with the FPGA 106. The image sensor can also be associated with a GPU 108 and a DPU 122.

[0048] The FPGA 106 can provide images 116 from at least one media stream 126 to the DPU 122 in the format of payload data. The DPU 122 can receive information 308 about only image sections 116A of the images 116. The DPU 122 can arrange payload data 310, representing the image sections 116A, in a shared and contiguous buffer 302 of memory 202. This arrangement can be made accessible to the GPU 108 and can be partially based on designated memory locations 304 in the shared and contiguous buffer 302. In one example, a designated memory location can serve to ensure contiguous memory or contiguous access. The GPU 108 can access the shared and contiguous buffer 302 and can process only the image sections 116A of the images.

[0049] Furthermore, the system 300 can be configured such that the FPGA 106 can also provide simultaneous media streams 126, 126A to the DPU 122, as described in relation to Fig. 2 described. The DPU 122 can also store header areas H15 to H2N for user data P15 to P1N and the first of the concurrent media streams, and can store additional header areas H21 to H2N for additional user data P21 to P2N of the other of the concurrent media streams. However, the different header areas of the different media streams can be stored in different buffers B1 and B2 of the local memory 104 of the host 120. The DPU 122 can arrange the user data P15 to P2N of all the image sections 116A in contiguous of the designated memory locations 304 of the shared and contiguous buffer 302. Furthermore, it is possible to store user data in a way that allows contiguous access instead of using a contiguous buffer 302.

[0050] In at least one embodiment, the system 300 can be configured such that the shared and contiguous buffer 302 can be local to the GPU 108, for example, on a graphics card 110 or another card. The distinct buffers B1, B2 to BN for the head regions H11 to H2N in a host 120 can be local to the host and accessible by a CPU 118 of the host 120. The system 120 can be configured such that the shared and contiguous buffer 302 is located on a GPU or graphics card 110, as illustrated, wherein the GPU or graphics card 110 includes the GPU 108. In at least one example, however, the shared and contiguous buffer 302 can be located on an accelerator or converged card, which can include the GPU 108 and the DPU 122.

[0051] Furthermore, the system 300 can be configured such that an image sensor comprises a multi-sensor array 102 with different sensors to provide different and simultaneous media streams 126, 126A of at least one media stream. The system 300 can also be configured such that the image sensor can communicate simultaneous media streams 126, 126A to the FPGA 106. In addition, the simultaneous media streams 126, 126A can be associated with different User Datagram Protocol (UDP) ports of the FPGA and can use the different UDP ports to identify the different media streams 126, 126A to the DPU. The system can be set up such that, after arranging 310 the payload that represents only the image sections 116A for the GPU 108, the DPU 122 can discard other payload that differs from the image sections 116A 306.

[0052] Fig. Figure 4 illustrates computer and processor aspects 400 of a system for separating image sections from images to support content-level processing of only image sections in at least one embodiment. For example, each of the illustrated processors 402 can comprise one or more processing or execution units 408 that can perform any or all of the aspects of the systems 100-300 for separating image sections from images and enabling content-level processing of only the image sections. Therefore, the processors 402 can be at least one CPU, but can also include aspects of a GPU and a DPU. In addition, the systems 100-300 can include different interfaces between each of the FPGA, the GPU, and the DPU to enable communication, as described throughout.

[0053] The processing or execution units 408 can comprise multiple circuits to support the aspects described herein for separating image sections from images to support content-level processing of only image sections. In at least one embodiment, the processors 402 can comprise CPUs, GPUs, or DPUs, which can be associated with a multi-tenant environment to perform one or more aspects of separating image sections from images to support content-level processing of only image sections. Furthermore, the GPUs can be associated with a DPU (represented by a network control device 434) and a CPU, which is defined by the Fig. The four illustrated processors 402 are represented separately in separate graphics / video cards 412. Therefore, although described in the singular, the graphics / video card 412 can comprise multiple cards and can have multiple GPUs on each card. This can also be the case with multiple DPUs on a network control device 434. In addition, it is also possible for a card to comprise DPUs and GPUs on it to perform the aspects described herein for separating image sections from images to support performing content-level processing of only image sections.

[0054] According to at least one embodiment, the computer and processor aspects 400 can be implemented by one or more processors 402 comprising a system-on-a-chip (SOC) or a combination thereof, which is configured with a processor that may include execution units configured to execute instructions. In at least one embodiment, the computer and processor aspects 400 may in particular include a component, such as a processor 402, configured to use execution units 408 comprising logic to execute algorithms for process data according to the present disclosure, as in an embodiment described herein.In at least one embodiment, the Computer and Processor Aspects 400 may include processors such as PENTIUM® processor family processors, Xeon™, Itanıum®, XScale™ and / or StrongARM™, Intel® Core™ or Intel® Nervana™ microprocessors available from Intel Corporation in Santa Clara, California, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, and the like) may also be used. In at least one embodiment, the Computer and Processor Aspects 400 may run a version of the WINDOWS operating system available from Microsoft Corporation in Redmond, Washington, although other operating systems (for example, UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0055] Embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of executing one or more instructions according to at least one embodiment.

[0056] In at least one embodiment, the computer and processor aspects 400 may in particular comprise a processor 402, which may in particular comprise one or more execution units 408 configured to perform aspects according to techniques relating to at least one or more of the present Fig. 1-3 and Fig. 5-7 are described. In at least one embodiment, the computer and processor aspects 400 are a single-processor desktop or server system, but in another embodiment, the computer and processor aspects 400 may be a multi-processor system.

[0057] In at least one embodiment, the processor 402 may, in particular, comprise a microprocessor of a computer with a complex instruction set (“CISC”), a microprocessor of a computer with a reduced instruction set (“RISC”), a microprocessor with very long instruction words (“VLIW”), a processor implementing a combination of instruction sets, or any other processing device, such as a digital signal processor. In at least one embodiment, a processor 402 may be coupled to a processor bus 410, which can transmit data signals between processors 402 and other components in the computer and processor aspects 400.

[0058] In at least one embodiment, a processor 402 may, in particular, comprise an internal level 1 cache memory (“cache”) 404 (“L1”). In at least one embodiment, a processor 402 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may be located outside of a processor 402. Other embodiments may also include a combination of both internal and external caches, depending on a specific implementation and requirements. In at least one embodiment, a register set 406 may store different data types in different registers, such as, in particular, integer registers, floating-point registers, status registers, and instruction pointer registers.

[0059] In at least one embodiment, a processor 402 also includes an execution unit 408, which in particular comprises logic for performing integer and floating-point operations. In at least one embodiment, a processor 402 can also include a read-only memory ("ROM") containing microcode ("ucode") in which microcode for specific macro instructions is stored. In at least one embodiment, an execution unit 408 can include logic for handling a packed instruction set 409.

[0060] In at least one embodiment, by incorporating a packed instruction set 409 into an instruction set of a general-purpose processor, together with associated circuitry for executing the instructions, operations used by many multimedia applications can be performed using packed data in a processor 402. In at least one embodiment, many multimedia applications can be executed faster and more efficiently by using a full width of the data bus of a processor to perform operations on packed data, which can eliminate the need to transfer smaller data units over the data bus of that processor to perform one or more operations on a data element sequentially.

[0061] In at least one embodiment, an execution unit 408 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer and processor aspects 400 can, in particular, include a main memory 420. In at least one embodiment, a main memory 420 can be a storage device such as dynamic random-access memory (“DRAM”), static random-access memory (“SRAM”), or flash memory, or another storage device. In at least one embodiment, a main memory 420 can store instruction(s) 419 and / or data 421, which are represented by data signals that can be executed by a processor 402.

[0062] In at least one embodiment, a system logic chip can be coupled to a processor bus 410 and a main memory 420. In at least one embodiment, a system logic chip can, in particular, comprise a memory controller hub (MCH) 416, and the processors 402 can communicate with the MCH 416 via the processor bus 410. In at least one embodiment, an MCH 416 can provide a high-bandwidth memory path 418 to a main memory 420 for storing applications and data and for storing graphics instructions, data, and textures. In at least one embodiment, an MCH 416 can route data signals between a processor 402, a main memory 420, and other components in the computer and processor aspects 400, and route data signals between a processor bus 410, a main memory 420, and a system I / O interface 422.In at least one embodiment, a system logic chip can provide a graphics port for coupling with a graphics control unit. In at least one embodiment, an MCH 416 can be coupled to a main memory 420 via a high-bandwidth memory path 418, and a graphics / video card 412 can be coupled to an MCH 416 via an Accelerated Graphics Port (“AGP”) intermediary 414. In at least one embodiment, the graphics / video card 412 can be coupled to one or more of the processors 402 via a PCIe intermediary standard. Likewise, a network control unit 424 can also be coupled to one or more of the processors 402 via a PCIe intermediary standard.

[0063] In at least one embodiment, the computer and processor aspects 400 can use a system I / O interface 422 as a proprietary hub interface bus to couple an MCH 416 with an I / O controller hub (ICH) 430. In at least one embodiment, an ICH 430 can provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, a local I / O bus can, in particular, comprise a high-speed I / O bus for connecting peripheral devices to a memory 420, a chipset, and processors 402.Examples may include, in particular, an audio control device 429, a firmware hub (“Flash BIOS”) 428, a wireless transceiver 426, a data storage device 424, a legacy I / O control device 423 incorporating a user input and keyboard interface(s) 425, a serial expansion port 427, such as a Universal Serial Bus (“USB”) port, and a network control device 434. In at least one embodiment, the data storage device 424 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or another mass storage device.

[0064] Illustrated in at least one embodiment Fig. 4 Computer and processor aspects 400, which include interconnected hardware devices or “chips”, whereas Fig. 4. In other embodiments, an exemplary SoC can be illustrated. In at least one embodiment, in Fig. 4 illustrated devices are interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of the computer and processor aspects 400 are interconnected using Compute Express Link (CXL) interconnects.

[0065] Therefore, the at least one execution unit 408 can be a circuit of at least one processor 402 configured to be associated with a system for separating image segments to support content-level processing of only image segments. The association can be such that the at least one execution unit 408 of at least one processor 402 can implement at least aspects of a GPU, aspects of a DPU, or aspects of a CPU. The association can be such that the at least one execution unit 408 of at least one processor 402 can load and run or execute instructions to implement such aspects. However, the association can also be such that the at least one execution unit 408 of at least one processor 402 can be hard-wired to implement such aspects.

[0066] Furthermore, at least one execution unit 408 can be a circuit of at least one processor 402, which can be a CPU, a DPU or a GPU, as in Fig. 1-3, to implement aspects of separating image sections from images to support performing content-level processing of only image sections. Thus, the computer and processor aspects 400 can include multiple circuits that may include or be part of a GPU and that may include or be part of the FPGA associated with the GPU. The FPGA can be configured to receive images from an image sensor. The FPGA can also perform physical-level processing on the images and can provide a media stream that includes the images after physical-level processing. The FPGA provides the media stream to a data processing unit (DPU). Separately, the GPU can perform content-level processing of only image sections of the images, partly based on the image sections provided by the DPU from the media stream.

[0067] Furthermore, physical-layer processing by the FPGA can include one or more analog-to-digital conversions (ADCs), noise estimation, or timestamping. Separately, content-layer processing by the GPU can include one or more pattern recognition, object detection, feature extraction, feature characterization, or image segmentation. Physical-layer processing can also be an operation that ignores the image content or is performed only considering raw pixel data associated with the images. Finally, physical-layer processing can be an operation free of business logic and independent of, or agnostic to, any application requirement.

[0068] In one example, content-level processing can include one or more operations configured to consider image content or performed on raw pixel data with due regard for the content within the images. The multiple circuits described here can be configured such that the DPU can interface with a GPU kernel. The GPU kernel allows the GPU to specify to the DPU only the image portions to be received for content-level processing. As a result, the GPU receives only those image portions that are to be subjected to content-level processing within the GPU. The GPU can then perform its image processing using command scripts from the GPU kernel.The multiple circuits described here can be configured such that the GPU can also communicate with the DPU using a PCIe bus to specify only the image segments to be received. This ensures that the GPU only receives these image segments from the DPU. The GPU is configured to perform content-level processing without CPU intervention. The multiple circuits described here can also be configured such that the GPU can perform local access to the image segments provided by the DPU, at least because the image segments are stored in a buffer that is local to the GPU.

[0069] Furthermore, the at least one execution unit 408 can be a circuit of at least one processor 402 configured to be associated with a CPU, a DPU or a GPU, as in Fig. 1-3, to implement aspects of separating image sections to support content-level processing of only image sections. Thus, the computer and processor aspects 400 can include multiple circuits that may include or be part of a DPU and that may include or be part of a GPU. For example, the multiple circuits provide a DPU associated with a GPU. The DPU can receive images from at least one media stream from an FPGA, which may be a different additional circuit. The DPU can separate header regions of payload associated with the images and can provide sequence numbers for the payload that represent only image sections. The DPU can make the payload representing only the image sections available for access associated with the GPU.This allows the GPU to access and process user data, which only represents the image sections, using the sequence numbers.

[0070] The multiple circuits can be configured to handle two or more media streams simultaneously. For example, if two media streams are present from an FPGA, the DPU can provide the headers associated with the first of the two media streams and allow access to these headers from a host machine. In one example, the host could include a CPU as another part of the multiple circuits. Furthermore, the DPU could provide additional headers associated with a second of the two media streams. These additional headers could then be accessed separately from the headers associated with the first of the two media streams on the host machine.This may at least be due to the fact that the different header areas of the different media streams can be provided in different buffers, which represent the different access methods.

[0071] The multiple circuits can be configured to allow the DPU to receive information about image sections from the CPU, based in part on predefined information provided to the DPU by an application. The CPU can also provide image section information using the header areas and additional header areas. The DPU can then perform operations on the payload and additional payload, which represent only the image sections that the DPU provides for access by the GPU. The multiple circuits can be configured to provide the payload, associated with multiple media streams representing multiple image sections of those media streams, for contiguous access by the GPU.The multiple circuits can be configured in such a way that the GPU can process the payload data, which only represents the image sections of the images, based partly on input from the CPU part of the multiple circuits, which may be located in a host machine and use the information from the application that requires the image processing of the image sections.

[0072] The multiple circuits can be configured such that the header portions of the media streams can be received for local access using a CPU portion of the multiple circuits in a host machine. The CPU can allow the DPU to provide the payload, representing only the image segments, for local access by the GPU. The CPU can allow the GPU to process the payload, representing only the image segments, partially based on the sequence numbers. The multiple circuits can be configured such that the DPU can control the provision of the payload via a data stream whose bit rate and burst size are associated with predictable workloads at a known GPU consumption rate.

[0073] Furthermore, the at least one execution unit 408 can be a circuit of at least one processor 402 configured to be associated with a CPU, a DPU or a GPU, as in Fig. 1-3, to implement aspects of separating image sections to support content-level processing of only image sections. Thus, the computer and processor aspects 400 can include multiple circuits, which may include or be part of a GPU and which may include or be part of a DPU. For example, the multiple circuits can include the GPU and the DPU, where the DPU can receive images from at least one media stream from an FPGA, which may be another of the multiple circuits. The FPGA may be associated with an image sensor. The DPU can receive information about image sections and can arrange payload data representing the image sections in a shared buffer for the GPU. The arrangement may be based, in part, on designated memory locations in the shared buffer.In one example, the designated memory locations could be contiguous blocks within the shared buffer, assigned to media streams or sequence numbers associated with the payload. The GPU can access the shared buffer to process only the relevant portions of the images.

[0074] The multiple circuits can be configured such that the DPU can also receive simultaneous media streams of at least one media stream. The DPU can store header areas for the payload, along with additional header areas for additional payload from another of the simultaneous media streams, in different buffers associated with a host. For example, the DPU can use a relationship function to maintain header information and sequence numbers, which may be based on a transformation of the header information. Alternatively, the payload and additional payload can be arranged using the contiguous of the designated memory locations of the shared buffer.

[0075] Furthermore, the shared buffer can be local to the GPU by being located on the same card as the GPU, while multiple buffers are local to a CPU by being located within the same host machine that hosts the CPU. Alternatively, the buffer can reside on a GPU card that also houses the GPU. Even further, the shared buffer can reside on an accelerator card or a converged card, which may also house the GPU and the DPU. The multiple circuits can also be configured such that, after arranging the payload data representing only the at least one image segment for the GPU, the DPU can discard other payload data that differs from this at least one image segment.

[0076] Fig. Figure 5 illustrates a process flow or method 500 for a system for separating physical-layer processing of images from content-layer processing of image portions only, in at least one embodiment. The method 500 may include capturing images 502 using an image sensor. The method 500 may include performing physical-layer processing 504 on images using an FPGA to provide a media stream. The method 500 may include verifying or determining 506 that image portions are specified. In one example, this may be based in part on information from a host machine's CPU. The method 500 may include providing, using a DPU, only image portions from the images in the media stream to the GPU.Procedure 500 can also include performing 510 content-level processing on only the image sections of the images using the GPU.

[0077] Method 500 can include a further step or substep in which the GPU is enabled to communicate with the DPU using a PCIe bus. Method 500 can include a further step or substep in which only the image segments are received by the GPU using the PCIe bus. Furthermore, content-level processing can be performed in the GPU without CPU intervention. Method 500 can include a further step or substep in which the image segments are made available by the DPU for local access by the GPU. Method 500 can include a further step or substep in which designated memory locations are determined for local access to payload data representing only the image segments. Then, the payload data representing only the image segments can be accessed locally using the GPU.After local access, the GPU can perform content-level processing for the payload data, which only represents the image sections.

[0078] In Procedure 500, physical-layer processing can be an operation that ignores the content of the images or is performed only with regard to raw pixel data associated with the images. Alternatively, physical-layer processing can be an operation that is free of business logic or that is independent of or agnostic to an application requirement. Content-layer processing can include one or more operations that are configured to consider the content of the images or that are performed on raw pixel data with due consideration of the content within the images.

[0079] Fig. Figure 6 illustrates a further process flow or method 600 for a system for providing sequence numbers for user data that represent only image sections of images, in at least one embodiment. Method 600 from Fig. 6 can be obtained using the procedure 500 from Fig. 5. In an example, the method 600 may include providing 602 an image sensor associated with an FPGA, a GPU, and a DPU. The method 600 may include verifying or determining 604 that images are to be acquired using the image sensor. The method 600 may include, using the FPGA, providing 606 images acquired by the image sensor that are in at least one media stream for the DPU. The method 600 may include, by the DPU, separating 608 header regions of payload data associated with the images.

[0080] Method 600 can include providing sequence numbers for the payload, which represent only image sections. Method 600 can also include providing the payload, which represents only the image sections, by the DPU for access by the GPU. This can be partially based on information from a CPU that has access to the header regions of the payload. Method 600 can also include processing the payload, which represents only the image sections, by the GPU using the sequence numbers.

[0081] Method 600 can be configured such that the at least one media stream comprises two media streams. The header areas can be associated with a first of the two media streams and can be made available for access by a host machine with a CPU. Furthermore, additional header areas can be associated with a second of the two media streams. The additional header areas can be made available by the host machine for separate access relative to the header areas associated with the first of the two media streams. Method 600 can include a further step or substep in which the DPU is informed of the payload data, which represents only the image portions to be made available by the DPU for access by the GPU. This informing can be based, in part, on the CPU providing input using information from an application associated with a DPU.The CPU can also use information from the header areas and the additional header areas.

[0082] Method 600 can be configured such that at least one media stream comprises two media streams, wherein the payload can be associated with a first of the two media streams. The payload can represent only the image portions of the first of the two media streams. The payload can be made available for access by the GPU together with additional payload, which can be associated with a second of the two media streams. For example, the additional payload can represent only additional image portions of the second of the two media streams and can be made available for contiguous access by the GPU together with the payload associated with the first of the two media streams.

[0083] Method 600 can include a further step or substep in which the payload, representing only the image portions, is processed using the GPU. This can be partially based on input from a host machine's CPU. The CPU can use information from an application to provide the input. In one example, the application might be one that requires image processing to be performed on the image or on the image portions. In another example, the CPU might use information from the header regions to provide the input. Method 600 can include a further step or substep in which the DPU controls the provision of the payload via a data stream whose bit rate and burst size are associated with predictable workloads at a known GPU consumption rate.

[0084] Fig. Figure 7 illustrates a further process flow or method 700 for a system for arranging user data representing only image sections in a shared buffer in at least one embodiment. The method 700 from Fig. 7 can be obtained using the 500 method. Fig. 5 or the procedure 600 from Fig.6. Method 700 may include providing 702 an image sensor associated with an FPGA, a GPU, and a DPU. Method 700 may include verifying or determining 704 whether images should be acquired by the image sensor. Method 700 may include providing 706, from the FPGA, images from at least one media stream to the DPU. Method 700 may include receiving 708, by the DPU, information about only image portions. This can be done by a CPU, based partly on access to and information from an application instead of, or together with, the header regions associated with the payload data representing the image portions.Method 700 can include the arrangement 710 of the payload data representing the image sections by the DPU in a shared buffer for the GPU, based partly on designated memory locations in the shared buffer. Method 700 includes the processing 712 of only the image sections by the GPU, based partly on accessing the shared buffer for the image sections.

[0085] Method 700 may include a further step or substep in which the FPGA provides concurrent media streams of at least one media stream to the DPU. Method 700 may include a further step or substep in which the DPU stores header areas for the payload and additional header areas for additional payload of another of the concurrent media streams in different of several buffers. Instead of the header areas, a transformation function performed on aspects of the header area may provide information that can be retained by the DPU. Method 700 may include a further step or substep in which, as part of step 710, the DPU arranges the payload and the additional payload into contiguous of the designated memory locations of the shared buffer.

[0086] Method 700 can be configured such that the shared buffer is local to the GPU. The multiple buffers can be local to a CPU of a host machine. Method 700 can be configured such that the shared buffer resides on a GPU card, which may contain the GPU, or on a NIC, which may contain both the GPU and the DPU. The various buffers can reside within the host machine. Method 700 can be configured such that the image sensor may comprise a multi-array sensor. Method 700 can include a further step or substep in which different sensors of the multi-array sensor provide different and simultaneous media streams of the at least one media stream to the FPGA. The simultaneous media streams can be associated with different UDP ports of the FPGA.Method 700 may include a further step or sub-step in which, after arranging the payload data representing only the image sections, other payload data that differs from the image sections is discarded by the DPU for the GPU.

[0087] The following description sets out numerous specific details to provide a more thorough understanding of at least one embodiment. However, it will be apparent to those skilled in the art that the concepts according to the invention can be implemented without one or more of these specific details.

[0088] Other variations are inherent in the concept of the present disclosure. Disclosed techniques can be modified or alternatively constructed in various ways; however, only certain illustrated embodiments are shown in the drawings and described in detail above. It is understood, however, that the disclosure is not intended to be limited to a specific form or the disclosed forms, but rather, on the contrary, it is intended to cover all modifications, alternative constructions, and equivalents that fall within the concept and scope of the disclosure as defined in the appended claims.

[0089] The use of the terms "a," "an," and "the" and similar references in connection with the description of disclosed embodiments (particularly in connection with the appended claims) is to be interpreted as covering both the singular and the plural, unless otherwise specified or the context clearly contradicts this, and is not to be understood as a definition of terms. The terms "comprise," "have," "exhibit," "contain," "including," and "with" are to be interpreted as non-restrictive, open terms (meaning "including, but not limited to"), unless otherwise specified. "Connected," when not modified and referring to physical connections, is to be interpreted as meaning that the connected elements are partially or completely encompassed, attached to, or joined together, even if something is in between.The mention of ranges of values ​​is intended here merely as an efficient method of referring individually to each separate value falling within the range, unless otherwise specified, and each separate value is deemed to be included in the description as if it had been explicitly mentioned. In at least one embodiment, the use of the term "set" (e.g., "a set of elements") or "subset," unless otherwise specified or the context contradicts, is to be interpreted as a non-empty collection comprising one or more elements. Furthermore, unless otherwise specified or the context contradicts, the term "subset" of a corresponding set does not necessarily denote a proper subset of a corresponding set; rather, the subset and the corresponding set may be identical.

[0090] Language with conjunctions, such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C," is, unless explicitly stated otherwise or the context clearly contradicts it, to be understood in the context that it is generally intended to show that an element, a concept, etc., can be either A, B, C, or any non-empty subset of a set containing A, B, and C. For example, in an illustrative example of a set with three elements, sentences with conjunctions "at least one of A, B, and C" and "at least one of A, B, and C" refer to each of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such a linguistic construction with conjunctions is generally not intended to imply that certain embodiments require, for instance, the presence of at least one of A, at least one of B, and at least one of C.Additionally, unless otherwise stated or the context contradicts, the terms "several" or "a plurality" indicate a plural state (e.g., "a plurality of elements" indicates multiple elements). In at least one embodiment, the number of elements in a plurality (the number of multiple elements) is at least two, but may be greater if either explicitly stated or indicated by context. Furthermore, unless otherwise stated or the context clearly indicates, the phrase "based on" means "based at least in part on" and not "based exclusively on".

[0091] The processes described herein may be carried out in any suitable order, unless otherwise specified or the context clearly contradicts this. In at least one embodiment, a process, such as one of the processes described herein (or variations and / or combinations thereof), is carried out under the control of one or more computer systems configured with executable instructions and is implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors by hardware or combinations thereof. In at least one embodiment, code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising several instructions that can be executed by one or more processors.

[0092] In at least one embodiment, a computer-readable storage medium is a non-volatile computer-readable storage medium that excludes volatile signals (e.g., a propagating transient electrical or electromagnetic transmission) but includes non-volatile data storage circuitry (e.g., buffers, caches, and queues) within transceivers of volatile signals. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-volatile computer-readable storage media on which executable instructions are stored (or on other storage for executable instructions) which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein.In at least one embodiment, the set of non-volatile, computer-readable storage media comprises multiple non-volatile, computer-readable storage media, and one or more individual non-volatile storage media from among multiple non-volatile, computer-readable storage media lack all the code, whereas multiple non-volatile, computer-readable storage media collectively store all the code. In at least one embodiment, executable instructions are executed such that different instructions are executed by different processors—for example, a non-volatile, computer-readable storage medium stores instructions, and a central processing unit (“CPU”) executes some of the instructions, while a graphics processing unit (“GPU”) executes other instructions.In at least one embodiment, different components of a computer system have separate processors, and different processors execute different subsets of instructions.

[0093] In at least one embodiment, an arithmetic logic unit is a set of combinational logic circuits that take one or more inputs to produce a result. In at least one embodiment, an arithmetic logic unit is used by a processor to implement a mathematical operation such as addition, subtraction, or multiplication. In at least one embodiment, an arithmetic logic unit is used to implement logical operations such as logical AND / OR or XOR. In at least one embodiment, an arithmetic logic unit is stateless and consists of physical switching components, such as semiconductor transistors, arranged to form logic gates. In at least one embodiment, an arithmetic logic unit can operate internally as a stateful logic circuit with an associated clock.In at least one embodiment, an arithmetic logic unit can be constructed as an asynchronous logic circuit with an internal state that is not held in an associated register set. In at least one embodiment, an arithmetic logic unit is used by a processor to combine operands stored in one or more registers of the processor and to generate an output that can be stored by the processor in another register or memory location.

[0094] In at least one embodiment, as a result of processing an instruction called by the processor, the processor provides one or more inputs or operands to an arithmetic logic unit. This causes the arithmetic logic unit to generate a result that is based at least partially on instruction code provided to the inputs of the arithmetic logic unit. In at least one embodiment, the instruction codes provided by the processor to the arithmetic logic unit are based at least partially on the instruction being executed by the processor. In at least one embodiment, combinational logic within the arithmetic logic unit processes the inputs and generates an output that is placed on a bus within the processor.In at least one embodiment, the processor selects a destination register, memory location, output device, or output memory location on the output bus, such that the clocking of the processor causes the results generated by the arithmetic logic unit to be sent to the desired location.

[0095] Accordingly, in at least one embodiment, computer systems are configured to implement one or more services that individually or jointly perform the processes described herein, and such computer systems are configured with applicable hardware and / or software that enables the execution of these processes. Furthermore, a computer system implementing at least one embodiment of the present disclosure can be a single unit, or, in another embodiment, it can be a distributed computer system comprising several units that operate differently, such that the distributed computer system performs the processes described herein and a single unit does not perform all of them.

[0096] The use of any and all examples or exemplary language provided herein (e.g., "such as") is intended solely to better illuminate embodiments of the disclosure and does not constitute a limitation of the scope of the disclosure unless claimed otherwise. No wording in the description should be interpreted as indicating that an unclaimed element is essential for the practical implementation of the disclosure.

[0097] In the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It is understood that these terms are not necessarily intended as synonyms. Rather, in certain examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other but nevertheless work together or interact with each other.

[0098] Unless expressly stated otherwise, terms such as "processing", "calculating", "calculating", "determining" or the like are understood to refer to actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or convert data represented as physical or electronic quantities within the registers and / or memory of the computing system into other data represented in a similar manner as physical quantities within the memory, registers or other such information storage, transmission or display devices of the computing system.

[0099] Similarly, the term "processor" can refer to any unit or section of a unit that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As non-restrictive examples, "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used here, "software" processes can, for example, include software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, a given process can also refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently.In at least one embodiment, the terms “system” and “method” are used interchangeably in the present case, insofar as the system can embody one or more methods and insofar as methods can be considered as a system.

[0100] This document may refer to the receiving, acquiring, capturing, receiving, or inputting of analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of receiving, acquiring, capturing, receiving, or inputting analog and digital data can be achieved in various ways, such as receiving data as a parameter of a function call or a call to an application programming interface. In at least one embodiment, processes of receiving, acquiring, capturing, receiving, or inputting analog or digital data can be achieved by transmitting data over a serial or parallel interface.In at least one embodiment, processes of obtaining, acquiring, capturing, receiving, or inputting analog or digital data can be achieved by transmitting data over a computer network from a providing entity to a receiving entity. Reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In at least one embodiment, processes of providing, outputting, transmitting, sending, or presenting analog or digital data can be achieved by transmitting data as an input or output parameter of a function call, a parameter of an application programming interface, or as an interprocess communication mechanism.

[0101] Although the descriptions presented here are exemplary implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to be within the scope of this disclosure. While specific distributions of responsibilities may be defined above for the purpose of description, various functions and responsibilities may also be distributed and subdivided in different ways depending on the circumstances.

[0102] Furthermore, although the subject matter was described in a language specific to structural features and / or methodological actions, it is understood that the subject matter claimed in the attached claims is not necessarily limited to specific features or described actions. Rather, specific features and actions are disclosed as exemplary forms of implementing the claims.

[0103] It is understood that aspects and embodiments described herein are purely exemplary, and that details may be modified within the scope of the claims.

[0104] The individual devices, methods and features disclosed in the description and (where appropriate) the claims and drawings may be provided independently of one another or in any suitable combination.

[0105] Reference numerals appearing in the claims are for illustrative purposes only and do not limit the scope of the claims.

[0106] The disclosure of this application also includes the following numbered clauses: Clause 1. System comprising a data processing unit (DPU) configured to receive image data associated with captured images from at least one media stream, wherein the image data comprises payload and header regions, wherein the DPU is further configured to separate the header regions from the payload, to provide sequence numbers for the payload to represent only portions of the images, and to make only the portions of the images available in a shared buffer for access by a graphics processing unit (GPU) to enable the GPU to process the payload representing only the portions of the images using the sequence numbers. Clause 2. System according to Clause 1, furthermore encompassing: an image sensor for capturing the images; and a field-programmable gate array (FPGA) configured to receive the images and provide the image data as simultaneous media streams of at least one media stream. Clause 3. System according to Clause 2, wherein the DPU is further configured to store the header areas for the payload of a first of the concurrent media streams and additional header areas for additional payload of a second of the concurrent media streams in different of several buffers that are distinct from the shared buffer, wherein the DPU is further configured to arrange the payload and the additional payload belonging to the image sections and additional image sections of the images in contiguous of the designated memory locations of the shared buffer, and wherein the arrangement of the payload and the additional payload enables the GPU to combine the image sections and the additional image sections for use by at least one application or for further processing by the GPU. Clause 4. System according to Clause 3, wherein the shared buffer is local to the GPU and wherein the multiple buffers are local to a central processing unit (CPU) of a host machine or the DPU. Clause 5. System according to Clause 3, wherein the shared buffer is located on a GPU card which has the GPU, and wherein the multiple buffers are located in the host machine or the DPU. Clause 6. System according to Clause 3, wherein the shared buffer is located on an accelerator card or a converged card which includes the GPU and the DPU, and wherein the multiple buffers are located in the host machine or the DPU. Clause 7. System according to one of Clauses 3-6, wherein the DPU is further configured to discard other user data that differs from the image sections after arranging the payload data that represents only the image sections for the GPU. Clause 8. System according to one of Clauses 2-7, wherein the image sensor comprises a multi-array sensor, and wherein different sensors of the multi-array sensor provide different and simultaneous media streams of the at least one media stream. Clause 9. System according to any of Clauses 2-8, wherein the image sensor is further configured to communicate simultaneous media streams of the at least one media stream to the FPGA, and wherein the simultaneous media streams are associated with different User Datagram Protocol (UDP) ports of the FPGA. Clause 10. System according to any of Clauses 1-9, wherein the at least one media stream comprises two media streams, the head areas being associated with a first of the two media streams and being provided in a first buffer for access by a host machine having a central processing unit (CPU), additional head areas being associated with a second of the two media streams, and the additional head areas being provided in a second buffer for separate access by the host machine relative to the head areas associated with the first of the two media streams. Clause 11. System according to Clause 8, wherein the CPU or the GPU is configured to use information from an application to inform the DPU of the payload data representing only the image portions of the images to be provided by the DPU to the GPU. Clause 12. System according to any of the preceding clauses, wherein the at least one media stream comprises two media streams, wherein the payload is associated with a first of the two media streams, representing only the image portions of the first of the two media streams, and is provided in the shared buffer for access by the GPU, and wherein additional payload is associated with a second of the two media streams, representing only additional image portions of the second of the two media streams, and is provided in the shared buffer or a separate buffer for contiguous access with the payload associated with the first of the two media streams by the GPU. Clause 13. System according to any of the preceding clauses, further comprising a central processing unit (CPU) of a host machine, wherein the CPU is configured to use information from an application to induce the GPU to process the payload data representing only the image portions of the images. Clause 14. System according to any of the preceding clauses, wherein the header areas are received for local access using a central processing unit (CPU) of a host machine, wherein the CPU enables the DPU to provide the payload, representing only the image sections, for local access by the GPU, and wherein the CPU enables the GPU to process the payload, representing only the image sections, partly based on the sequence numbers. Clause 15. System according to any of the preceding clauses, wherein the DPU is configured to control the delivery of the payload data via a data stream whose bit rate and burst size are associated with predictable workloads at a known consumption rate for the GPU. Clause 16. Several circuits comprising at least one data processing unit (DPU) configured to receive image data associated with captured images from at least one media stream, wherein the image data comprises payload and header regions, wherein the DPU is further configured to separate the header regions from the payload, to provide sequence numbers for the payload to represent only portions of the images, and to make only the portions of the images available in a shared buffer for access by a graphics processing unit (GPU) to enable the GPU to process the payload representing only the portions of the images using the sequence numbers. Clause 17. Several circuits according to Clause 16, further comprising: an image sensor configured to capture images, wherein the image sensor comprises a multi-array sensor, wherein different sensors of the multi-array sensor provide different and simultaneous media streams of the at least one media stream; and a field-programmable gate array (FPGA) configured to receive the images and provide the image data as the distinct and simultaneous media streams of at least one media stream. Clause 18. Multiple circuits according to Clause 16 or 17, wherein the shared buffer is one of: local to the GPU, on a GPU card which includes the GPU, or on an accelerator card or a converged card which includes the GPU and the DPU. Clause 19. Several circuits according to Clause 16, 17 or 18 configured to store the header areas in a first of several buffers, wherein the several buffers are distinct from the shared buffers and are local to or located in the host machine or the DPU, wherein the header areas are header areas for the payload of a first of the concurrent media streams of the at least one media stream, and wherein additional header areas for additional payload of a second of the concurrent media streams are located in a second of several buffers. Clause 20. Image processing methods, comprehensive: Receiving, in a data processing unit (DPU), image data associated with captured images from at least one media stream, wherein the image data includes user data and header areas, Separation, by the DPU, of the header areas from the user data; Providing sequence numbers for the payload data to represent only image sections of the images; and Providing only the image sections in a shared buffer for access by a graphics processing unit (GPU) to allow the GPU to process the payload, which represents only the image sections, using the sequence numbers. Clause 21. Procedure according to Clause 20, furthermore comprehensively: Capturing images using an image sensor, wherein the image sensor comprises a multi-array sensor, and wherein different sensors of the multi-array sensor provide different and simultaneous media streams of the at least one media stream; Receiving the images in a field-programmable gate array (FPGA); and Providing, from the FPGA, the image data as the different and simultaneous media streams of at least one media stream. Clause 22. Procedure according to Clause 20 or 21, wherein the shared buffer is of: local to the GPU, on a GPU card which includes the GPU, or on an accelerator card or a converged card which includes the GPU and the DPU. Clause 23. Procedure according to Clause 20, 21 or 22, furthermore including: The DPU stores the header areas for the payload of one of the concurrent media streams and additional header areas for additional payload of a second of the concurrent media streams in different of several buffers that are distinct from the shared buffer; Arrange, by the DPU, the payload and the additional payload belonging to the image sections and additional image sections of the images, in contiguous areas of the designated memory locations of the shared buffer; and Enable, using the arrangement of the payload and the additional payload, the GPU to combine the image segments and the additional image segments for use by at least one application or for further processing by the GPU.

[0107] The disclosure of the present application further comprises the following numbered clauses: Clause 1a. Distributed image processing system comprising a graphics processing unit (GPU) and a data processing unit (DPU), wherein the DPU is configured to receive image data associated with captured images and to provide a media stream comprising only image snippets from the images to the GPU, and wherein the GPU is configured to perform content-level processing on the image snippets of the images. Clause 2a. System according to Clause 1a, further comprising: an image sensor for capturing the images; and a field-programmable gate array (FPGA) that is set up to receive the images and perform physical-level processing on the image to provide the image data to the DPU. Clause 3a. System according to Clause 2a, wherein the physical level processing includes processing associated with symbols or a floating-point representation of the images, and wherein the content level processing includes processing associated with pixels or a metadata representation of the images. Clause 4a. System according to Clause 2a or 3a, wherein the physical-level processing includes one or more of channel equalization, analog-to-digital conversion (ADC), noise estimation, error correction or time stamping, and wherein the content-level processing includes one or more of pattern recognition, object recognition, feature extraction, feature characterization or image segmentation. Clause 5a. System according to Clause 2a, 3a or 4a, wherein the processing at the physical level includes one or more operations that ignore the content of the images or that are performed only taking into account raw pixel data associated with the images. Clause 6a. System according to any of Clauses 2a-5a, wherein the content-level processing comprises one or more operations designed to take into account any content of the images or which are performed on raw pixel data in relation to the content within the images. Clause 7a. System according to one of clauses 2a-6a, wherein the processing at the physical level is independent of or agnostic to an application requirement. Clause 8a. System according to any of the preceding clauses, further comprising a GPU core configured to interface with the DPU to specify only the image portions to be received for content-level processing in the GPU. Clause 9a. System according to any of the preceding clauses, wherein the GPU is further configured to communicate with the DPU using a Peripheral Component Interconnect Express (PCIe) bus to receive only the image segments, and wherein the GPU is further configured to perform content-level processing without intervention by a central processing unit (CPU). Clause 10a. System according to any of the preceding clauses, wherein the DPU is further configured to provide only the image sections for local access by the GPU. Clause 11a. Several circuits comprising a graphics processing unit (GPU) and a data processing unit (DPU), wherein the DPU is configured to receive image data associated with captured images and to provide a media stream comprising only image fragments from the images to the GPU, and wherein the GPU is configured to perform content-level processing on the image fragments of the images. Clause 12a. Several circuits according to Clause 11a, further comprising: an image sensor for capturing the images; and a field-programmable gate array (FPGA) that is set up to receive the images and perform physical-level processing on the image to provide the image data to the DPU. Clause 13a. Several circuits according to Clause 12a, wherein the physical-level processing includes processing associated with symbols or a floating-point representation of the images, and wherein the content-level processing includes processing associated with pixels or a metadata representation of the images. Clause 14a. Multiple circuits according to Clause 12a or 13a, wherein the physical-level processing includes one or more of channel equalization, analog-to-digital conversion (ADC), noise estimation, error correction or time stamping, and wherein the content-level processing includes one or more of pattern recognition, object recognition, feature extraction, feature characterization or image segmentation. Clause 15a. Multiple circuits according to Clause 12a, 13a or 14a, wherein the physical-level processing includes one or more operations that ignore the content of the images or that are performed only with regard to raw pixel data associated with the images, and wherein the content-level processing includes one or more operations that are configured to take into account any content of the images or that are performed on raw pixel data with respect to the content within the images. Clause 16a. Distributed image processing methods, comprising: Received, in a data processing unit (DPU), image data associated with images captured using an image sensor; Providing, from the DPU, a media stream of only image excerpts for a graphics processing unit (GPU); and Perform, using the GPU, content-level processing on the image crops. Clause 17a. Procedure according to Clause 16a, furthermore comprehensive: Receiving the images in a field-programmable gate array (FPGA), Performing physical-level processing on the images in the FPGA to provide the image data to the DPU. Clause 18a. Method according to Clause 17a, wherein the physical level processing includes processing associated with symbols or a floating-point representation of the images, and wherein the content level processing includes processing associated with pixels or a metadata representation of the images. Clause 19a. Method according to Clause 17a or 18a, wherein the physical-level processing includes one or more of channel equalization, analog-to-digital conversion (ADC), noise estimation, error correction or time stamping, and wherein the content-level processing includes one or more of pattern recognition, object recognition, feature extraction, feature characterization or image segmentation. Clause 20a. Method according to Clause 17a, 18a or 19a, wherein the physical-level processing comprises one or more operations that ignore the content of the images or that are performed only with regard to raw pixel data associated with the images, and wherein the content-level processing comprises one or more operations that are designed to take into account any content of the images or that are performed on raw pixel data in relation to the content within the images.

Claims

[1] Distributed image processing system comprising a graphics processing unit (GPU) and a data processing unit (DPU), wherein the DPU is configured to receive image data associated with captured images and to provide a media stream comprising only image fragments from the images to the GPU, and wherein the GPU is configured to perform content-level processing on the image fragments of the images. [2] System according to claim 1, further comprising: an image sensor for capturing the images; and a field-programmable gate array (FPGA) that is set up to receive the images and perform physical-level processing on the image to provide the image data to the DPU. [3] System according to claim 2, wherein the processing at the physical level comprises processing associated with symbols or a floating-point representation of the images, and wherein the processing at the content level comprises processing associated with pixels or a metadata representation of the images. [4] System according to claim 2 or 3, wherein the physical level processing comprises one or more of channel equalization, analog-to-digital conversion (ADC), noise estimation, error correction or time stamping, and wherein the content level processing comprises one or more of pattern recognition, object recognition, feature extraction, feature characterization or image segmentation. [5] System according to claim 2, 3 or 4, wherein the processing at the physical level comprises one or more operations that ignore the content of the images or that are performed only taking into account raw pixel data associated with the images. [6] System according to any one of claims 2-5, wherein the content-level processing comprises one or more operations configured to take into account the content of the images or which are performed on raw pixel data with respect to the content within the images. [7] System according to any one of claims 2 to 6, wherein the processing at the physical level is independent of or agnostic to an application requirement. [8] System according to any of the preceding claims, further comprising a GPU kernel configured to form an interface with the DPU to specify only the image sections to be received for content-level processing in the GPU. [9] System according to any of the preceding claims, wherein the GPU is further configured to communicate with the DPU using a Peripheral Component Interconnect Express (PCIe) bus to receive only the image sections, and wherein the GPU is further configured to perform content-level processing without intervention by a central processing unit (CPU). [10] System according to one of the preceding claims, wherein the DPU is further configured to provide only the image sections for local access by the GPU. [11] Several circuits comprising a graphics processing unit (GPU) and a data processing unit (DPU), wherein the DPU is configured to receive image data associated with captured images and to provide a media stream comprising only image fragments from the images to the GPU, and wherein the GPU is configured to perform content-level processing on the image fragments of the images. [12] Several circuits according to claim 11, further comprising: an image sensor for capturing the images; and a field-programmable gate array (FPGA) that is set up to receive the images and perform physical-level processing on the image to provide the image data to the DPU. [13] Multiple circuits according to claim 12, wherein the processing at the physical level comprises processing associated with symbols or a floating-point representation of the images, and wherein the processing at the content level comprises processing associated with pixels or a metadata representation of the images. [14] Multiple circuits according to claim 12 or 13, wherein the physical level processing comprises one or more of channel equalization, analog-to-digital conversion (ADC), noise estimation, error correction or time stamping, and wherein the content level processing comprises one or more of pattern recognition, object recognition, feature extraction, feature characterization or image segmentation. [15] Multiple circuits according to claim 12, 13 or 14, wherein the processing at the physical level comprises one or more operations that ignore the content of the images or that are performed only taking into account raw pixel data associated with the images, and wherein the processing at the content level comprises one or more operations that are configured to take into account any content of the images or that are performed on raw pixel data with respect to the content within the images. [16] Distributed image processing methods, comprising: Received, in a data processing unit (DPU), image data associated with images captured using an image sensor; Providing, from the DPU, a media stream of only image excerpts for a graphics processing unit (GPU); and Perform, using the GPU, content-level processing on the image crops. [17] The method of claim 16, further comprising: Receiving the images in a field-programmable gate array (FPGA); Performing physical-level processing on the images in the FPGA to provide the image data to the DPU. [18] Method according to claim 17, wherein the processing at the physical level comprises processing associated with symbols or a floating-point representation of the images, and wherein the processing at the content level comprises processing associated with pixels or a metadata representation of the images. [19] Method according to claim 17 or 18, wherein the physical level processing comprises one or more of channel equalization, analog-to-digital conversion (ADC), noise estimation, error correction or time stamping, and wherein the content level processing comprises one or more of pattern recognition, object recognition, feature extraction, feature characterization or image segmentation. [20] Method according to claim 17, 18 or 19, wherein the processing at the physical level comprises one or more operations that ignore the content of the images or that are performed only taking into account raw pixel data associated with the images, and wherein the processing at the content level comprises one or more operations that are configured to take into account any content of the images or that are performed on raw pixel data with respect to the content within the images.