Network processing unit, augmented reality display chip and device, and image processing method
By abolishing the image sharding processing and using network processing units to process multiple image data in parallel, the problem of excessive demand for system computing power and cache space in the prior art is solved, and more efficient image processing and delay performance is achieved.
Patent Information
- Application Number
- CN202311582338.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
When processing multiplexed image data, the previous technology needs to slice the entire image, which leads to an increase in the overall computing power and data cache space of the system, thereby limiting the parallel processing capability and delay performance of the network processing unit.
By canceling the need for sharding processing for the entire image, the network processing unit is used to acquire multiple image data row by row in parallel, and each image data is written to the buffer at a preset speed, and the image data is read and processed from the buffer using the full computing resources of the network processing unit.
It improves the accuracy of image processing, reduces the requirements for the overall computing power of the system and data cache space, improves the parallel processing capacity of a single network processing unit, and reduces network inference delay.
Smart Images

Figure CN120047302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of extended reality displays, and in particular, to a network processing unit, an extended reality display chip, an extended reality display device, a parallel processing method for image data, and a computer-readable storage medium. Background Art
[0002] Extended Reality (XR) display technology is an immersive display technology that creates a digital environment combining reality and virtuality through modern high-tech means centered around a computer, bringing an immersive experience of seamless conversion between the virtual world and the real world to the experiencer. It mainly includes various implementation methods such as Virtual Reality (VR) display, Augmented Reality (AR) display, and Mixed Reality (MR) display.
[0003] In the existing technologies in fields such as mobile phones, cameras, and surveillance, usually a separate Network Processing Unit (NPU) is configured for each Image Signal Processor (ISP) for network inference. Therefore, the overall computing power of the system cannot be flexibly and fully utilized, resulting in waste of the overall computing power of the system and increasing network inference latency.
[0004] In addition, after obtaining the entire image data with a resolution of H×W collected by the camera module, the existing technology usually needs to perform tiling processing on the entire image through an Image Signal Processor (ISP). According to the data processing capabilities of the Network Processing Unit (NPU), it is divided into multiple tiled images of N×M (N < H, M < W) and stored in the ISP buffer, and then the Network Processing Unit (NPU) reads the image data of each tiled image from the ISP buffer one by one for network inference. On the one hand, this will introduce additional operations and intermediate data such as tiling, zero-padding, windowing, and cropping overlapping images, which increases the requirements for the overall computing power of the system. On the other hand, it also requires the system to have a larger data cache space, greatly increasing the area, power consumption, and cost of the Image Signal Processor (ISP), thus severely restricting the development and application of existing Network Processing Units (NPU) for parallel processing of multiple paths of image data.
[0005] To overcome the above-mentioned defects existing in the prior art, there is an urgent need in the art for a parallel processing technology for multi-channel image data, which can improve the accuracy of image processing by eliminating the need for tiling the entire image, and reduce the requirements for the overall computing power and data cache space of the system. Thus, under the conditions of the same hardware processing technology and cost, the ability of a single network processing unit (NPU) to parallelly process multi-channel image data can be improved, and the overall computing power of the system can be flexibly and fully utilized to reduce the network inference latency, so as to meet the requirement of parallelly processing binocular image data in XR devices. Summary of the Invention
[0006] The following presents a brief overview of one or more aspects to provide a basic understanding of these aspects. This overview is not an exhaustive survey of all contemplated aspects, and is neither intended to identify key or decisive elements of all aspects nor to attempt to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to a more detailed description to follow.
[0007] To overcome the above-mentioned defects existing in the prior art, the present invention provides a network processing unit, an extended reality display chip, an extended reality display device, a parallel processing method for image data, and a computer-readable storage medium, which can parallelly obtain multi-channel image data row by row, write each channel of image data into a buffer at a preset speed, and then use the full computing resources of the network processing unit via an i model to read and process the i-th channel of image data that has been written into the buffer. By adopting these configurations, the present invention can improve the accuracy of image processing by eliminating the need for tiling the entire image, and reduce the requirements for the overall computing power and data cache space of the system. Thus, under the conditions of the same hardware processing technology and cost, the ability of a single network processing unit (NPU) to parallelly process multi-channel image data can be improved, and the overall computing power of the system can be flexibly and fully utilized to reduce the network inference latency, so as to meet the requirement of parallelly processing binocular image data in XR devices.
[0008] Specifically, the above-mentioned network processing unit according to the first aspect of the present invention is connected to a multi-channel image signal processor and is configured with a buffer and a plurality of pre-trained image processing models. The network processing unit alternately obtains N-channel image data row by row via the N-channel image signal processor and writes each channel of the image data into the buffer at a first speed, where N is an integer greater than 1, and the first speed is not less than N times the second speed at which each channel of the image signal processor outputs image data. In response to reaching at least one preset i-switching node, the image processing model i uses all the computing resources of the network processing unit to read and process the currently written i-th channel image data from the buffer at the second speed, where i is an integer not greater than N. In response to reaching the subsequent i + 1 switching node, the network processing unit caches the i-th channel intermediate data generated by the image processing model i for the next i-switching node to continue reading and processing the subsequently written i-th channel image data.
[0009] Further, in an embodiment of the present invention, the network processing unit obtains the entire i-th channel image data row by row via the image signal processor i. The image processing model i reads the previously cached intermediate data i according to a plurality of corresponding i-switching nodes and processes multiple rows of the currently written i-th channel image data window by window to update the intermediate data i regarding the i-th channel image data and / or generate result data i regarding the i-th channel image data.
[0010] Further, in an embodiment of the present invention, the image processing model i includes a multi-layer neural network structure and is configured to: in response to completing neural network inference of a preset number of layers L according to the currently written multiple rows of the image data, output result data row by row to the corresponding image signal processor i at the second speed and delete multiple rows of the image data that only involve the first L - 1 layer of neural network inference.
[0011] Further, in an embodiment of the present invention, each of the switching nodes is distributed at a preset time interval, or according to the latency requirements of each of the image processing models, or according to the total cached data volume of each of the intermediate data.
[0012] Further, in an embodiment of the present invention, the steps of determining the switching nodes distributed according to the delay requirements and determining the storage space corresponding to each of the image processing models include: dividing the entire storage space of the buffer according to the number of models to be processed in parallel to determine a first storage space distribution for storing the inference data of each of the image processing models; determining the total period of multi-model inference according to the inference period of each of the image processing models, and determining the number of inferences of each of the image processing models within the total period; respectively determining at least one switching node for time-division multiplexing each of the image processing models from multiple processing nodes of each of the image processing models according to the total period and the number of inferences of each of the image processing models; based on the first storage space distribution and the switching nodes of each of the image processing models, using the entire computing resources of the network processing unit to perform multi-model inference on the inference data samples of the multiple image processing models to determine the storage space lacking for each of the image processing models; and optimizing the first storage space distribution according to the entire storage space and the storage space lacking for each of the image processing models to determine a second storage space distribution that meets the multi-model inference requirements.
[0013] Further, in an embodiment of the present invention, the steps of determining the switching nodes distributed according to the total buffer data volume and determining the storage space corresponding to each of the image processing models include: dividing the entire storage space of the buffer according to the number of models to be processed in parallel to determine a third storage space distribution for storing the inference data of each of the image processing models; determining the total period of multi-model inference according to the inference period of each of the image processing models, and determining the number of inferences of each of the image processing models within the total period; respectively performing model inference on the inference data samples of each of the image processing models using the entire computing resources of the buffer based on the third storage space distribution and each of the processing nodes of each of the image processing models to determine the storage space lacking for each of the image processing models at each of the processing nodes; sorting the processing nodes in ascending order according to the lacking storage space, dividing different numbers of subgraphs for each of the image processing models according to the processing nodes, and performing inference tests to determine the required computing duration for each of the image processing models; determining the maximum number of subgraphs whose number of inferences and computing duration meet the total performance requirements of the multi-model inference; and respectively determining the switching nodes of each of the image processing models according to the positions of the processing nodes corresponding to the maximum number of subgraphs.
[0014] Further, in an embodiment of the present invention, the buffer is statically divided into N buffer spaces to buffer the currently to-be-processed image data and the intermediate data generated from previous processing.
[0015] Further, in an embodiment of the present invention, the network processing unit is connected to a binocular camera via two image signal processors and is configured to: alternately acquire left-eye image data and right-eye image data of the binocular camera row by row via the two image signal processors, and write the left-eye image data and the right-eye image data into the buffer respectively; in response to reaching a first switching node, utilize all the computing resources of the network processing unit via a first image processing model to read and process the currently written left-eye image data from the buffer to generate corresponding first intermediate data; in response to reaching a second switching node, cache the first intermediate data, switch all the computing resources of the network processing unit to a second image processing model, and read and process the currently written right-eye image data from the buffer via the second image processing model to generate corresponding second intermediate data; and in response to the next first switching node, cache the second intermediate data, read the first intermediate data, switch all the computing resources of the network processing unit back to the first image processing model, and read and process the further written left-eye image data from the buffer via the first image processing model to update the first intermediate data and / or generate result data regarding the left-eye image data.
[0016] Further, in an embodiment of the present invention, the left-eye image data and the right-eye image data are raw image data with noise, and the result data regarding the left-eye image data and the left-eye image data are denoised image data with noise removed. Alternatively, the left-eye image data and the right-eye image data are raw image data with mosaics, and the result data regarding the left-eye image data and the left-eye image data are restored image data with mosaics removed.
[0017] In addition, the above extended reality display chip provided according to the second aspect of the present invention includes: at least two image signal processors respectively connected to the left-eye camera and the right-eye camera of the extended reality display device; and the above network processing unit provided according to the first aspect of the present invention, which is respectively connected to each of the image signal processors to synchronously acquire the image data output by each of the image signal processors and perform parallel processing thereon.
[0018] In addition, the above extended reality display device provided according to the third aspect of the present invention includes: a binocular camera; and the extended reality display chip provided according to the second aspect of the present invention, wherein the extended reality display chip is connected to the binocular camera to synchronously acquire multiplexed image data output therefrom and perform parallel processing on each of the image data.
[0019] In addition, the parallel processing method of the above image data provided according to the fourth aspect of the present invention includes the following steps: alternately obtaining N-channel image data line by line via an N-channel image signal processor, and writing each channel of the image data into a buffer of a network processing unit at a first speed, where N is an integer greater than 1, and the first speed is not less than N times the second speed at which each channel of the image signal processor outputs image data; in response to reaching at least one preset i switching node, via an image processing model i pre-trained and configured in the network processing unit, using all computing resources of the network processing unit, reading and processing the currently written i-th channel of image data from the buffer at a second speed, where i is an integer not greater than N, and the second speed is not less than N times the first speed; and in response to reaching a subsequent i+1 switching node, caching the i-th channel of intermediate data generated by the image processing model i for the next i switching node to continue reading and processing the subsequently written i-th channel of image data.
[0020] In addition, computer instructions are stored on the computer-readable storage medium provided according to the fifth aspect of the present invention. When the computer instructions are executed by a processor, the parallel processing method of the above image data provided according to the fourth aspect of the present invention is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] After reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings, the above features and advantages of the present invention can be better understood. In the drawings, the components are not necessarily drawn to scale, and components having similar relevant characteristics or features may have the same or similar reference numerals.
[0022] Figure 1 The architecture diagram of an extended reality display chip provided according to some embodiments of the present invention is shown.
[0023] Figure 2 The flowchart of parallel processing of multi-channel image data provided according to some embodiments of the present invention is shown.
[0024] Figure 3 The schematic diagram of obtaining image data line by line provided according to some embodiments of the present invention is shown.
[0025] Figure 4 The structural diagram of an XR display chip provided according to some embodiments of the present invention is shown.
[0026] Figure 5 The schematic diagram of noise reduction processing provided according to some embodiments of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following specific embodiments illustrate the implementation manners of the present invention, and those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention will be introduced in conjunction with the preferred embodiments, this does not mean that the features of this invention are limited to this implementation manner. On the contrary, the purpose of introducing the invention in conjunction with the implementation manner is to cover other alternatives or modifications that may be extended based on the claims of the present invention. In order to provide a deep understanding of the present invention, many specific details will be included in the following description. The present invention can also be implemented without using these details. In addition, in order to avoid confusing or obscuring the key points of the present invention, some specific details will be omitted in the description.
[0028] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0029] In addition, the "upper", "lower", "left", "right", "top", "bottom", "horizontal", and "vertical" used in the following description should be understood as the orientations shown in this paragraph and the related drawings. This relative term is only for convenience of description, and it does not mean that the device described needs to be manufactured or operated in a specific orientation, so it should not be understood as a limitation to the present invention.
[0030] It can be understood that although the terms "first", "second", "third", etc. can be used here to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first component, region, layer, and / or part discussed below can be referred to as the second component, region, layer, and / or part without departing from some embodiments of the present invention.
[0031] As described above, after obtaining the entire image data with a resolution of H×W collected by the camera module, the prior art usually needs to perform tiling processing on the entire image through an Image Signal Processor (ISP), and divide it into multiple tiled images of N×M (N < H, M < W) according to the data processing capacity of the Network Processing Unit (NPU), and then store them in the ISP buffer. Then, the Network Processing Unit (NPU) reads the image data of each tiled image from the ISP buffer one by one for network inference. On the one hand, this will introduce additional operations and intermediate data such as tiling, zero-padding, windowing, and cropping overlapping images, which increases the requirements for the overall computing power of the system. On the other hand, it also requires the system to have a larger data cache space, which greatly increases the area, power consumption, and cost of the Image Signal Processor (ISP), thus severely restricting the development and application of the existing Network Processing Unit (NPU) for parallel processing of multiple-channel image data.
[0032] In order to overcome the above-mentioned defects of the prior art, the present invention provides a network processing unit, an extended reality display chip, an extended reality display device, a parallel processing method for image data, and a computer-readable storage medium, which can parallelly obtain multiple-channel image data row by row, write the image data of each channel into the buffer at a preset speed, and then the i-th model uses all the computing resources of the network processing unit to read and process the i-th channel image data that has been written from the buffer. By adopting these configurations, the present invention can improve the accuracy of image processing by eliminating the need for tiling processing of the entire image, and reduce the requirements for the overall computing power and data cache space of the system. Thus, under the conditions of the same hardware processing technology and cost, the ability of a single Network Processing Unit (NPU) to parallelly process multiple-channel image data can be improved, and the overall computing power of the system can be flexibly and fully utilized to reduce network inference latency, so as to meet the requirements for parallel processing of binocular image data in Extended Reality (XR) display devices.
[0033] In some non-limiting embodiments, the parallel processing method for image data provided in the fourth aspect of the present invention can be implemented via the extended reality (XR) display chip provided in the second aspect of the present invention. Specifically, please refer to Figure 1 , Figure 1 which shows the architecture diagram of the extended reality display chip provided in some embodiments of the present invention.
[0034] In Figure 1In the illustrated embodiment, the above-mentioned extended reality (XR) display chip provided by the second aspect of the present invention can be configured in the above-mentioned extended reality (XR) display device provided by the third aspect of the present invention, which is configured with a memory (not shown), at least two image signal processors 11-12, and the network processing unit 30 provided by the first aspect of the present invention. The memory includes but is not limited to the above-mentioned computer-readable storage medium provided by the fifth aspect of the present invention, on which computer instructions are stored. The at least two image signal processors 11-12 are respectively connected to the left-eye camera and the right-eye camera of the extended reality display device to obtain the real-world scene images collected by them. The network processing unit 30 is respectively connected to the memory and each of the image signal processors 11-12, and is adapted to read and execute the computer instructions stored on the memory to implement the above-mentioned parallel processing method of image data provided by the fourth aspect of the present invention, so as to alternately obtain the image data output by each of the image signal processors 11-12 line by line and perform parallel processing on the obtained image data.
[0035] The working principles of the above-mentioned network processing unit 30, extended reality display chip, and extended reality display device will be described below in conjunction with some embodiments of the parallel processing method of image data. Those skilled in the art can understand that these embodiments of the parallel processing method are only some non-limiting implementation manners provided by the present invention, aiming to clearly show the main concept of the present invention and provide some specific solutions convenient for the public to implement, rather than limiting all functions or all working manners of the network processing unit 30, the extended reality display chip, and the extended reality display device. Similarly, the network processing unit 30, the extended reality display chip, and the extended reality display device are also only a non-limiting implementation manner provided by the present invention, and do not limit the execution subject and execution order of each step in the parallel processing method of image data.
[0036] Please refer to Figure 1 and Figure 2 , Figure 2 which shows a flowchart of parallel processing of image data according to some embodiments of the present invention.
[0037] As shown in Figure 1As shown, the above-mentioned network processing unit 30 provided by the present invention is configured with hardware devices such as buffers 311-312, a multiply accumulation (MAC) array calculation unit 32, a vector processing unit (VPU) 33, and a register 34, and is configured with a software program containing multiple pre-trained image processing models. Its calculation operations and storage operations can be independently performed separately, so as to achieve full calculation resource sharing of the network processing unit 30 through the MAC array calculation unit 32 and the vector processing unit (VPU) 33, so as to flexibly support the network processing unit 30 to perform operations such as data reading, calculation, and caching.
[0038] In some embodiments, those skilled in the art can, in an offline manner, pre-train and compile the input parameters of a neural network model for image processing such as denoising and / or demosaicking according to the size of the entire real scene image, and adapt to the specific usage scenarios of the binocular cameras in the XR display device, and initialize the trained model N (for example: N = 2) times to obtain N neural network models with the same function, so as to respectively correspond to the image signal processors 11-12 of each path.
[0039] Those skilled in the art can understand that the above solution of configuring N neural network models with the same function is only a non-limiting implementation manner provided by the present invention, aiming to clearly show the main idea of the present invention and provide a specific solution convenient for the public to implement, rather than limiting the protection scope of the present invention.
[0040] Optionally, in some other embodiments, those skilled in the art can also adapt to various different image processing requirements such as denoising and demosaicking, and configure multiple neural network models with different parameters, structures, and / or functions to correspondingly meet the requirements of multi-functional parallel processing of binocular image data in the XR device.
[0041] In addition, in some embodiments, the buffers 311-312 may use a static random access memory (SRAM), which is configured at the input end of the network processing unit 30 and is statically divided into N (for example, N=2) independent buffer spaces to adapt to the specific usage scenario of the binocular camera in the XR display device. Further, each buffer space can be preferably divided into an input buffer space, a parameter buffer space, and a feature buffer space to cache the image data currently to be processed by the corresponding image processing model, the model parameter data, and the intermediate data generated by the neural network inference performed by the corresponding image processing model.
[0042] The first buffer 311 and the second buffer 312 are used to refer to the two independent buffer spaces. The first buffer 311 is connected to the image signal processor 11 via the ISP buffer 111, and the second buffer 312 is connected to the image signal processor 12 via the ISP buffer 121. The first buffer 311 and the second buffer 312 are connected to the image signal processor 12 via the ISP buffer 121. 1 , alternately read two channels of image data line by line from the corresponding ISP buffers 111-112, so that each image processing model can read the currently written image data from the corresponding buffers 311-312 according to the corresponding switching nodes, and use the full computing resources of the network processing unit 30 to process the image data, so as to perform parallel processing of the two channels of image data of the binocular camera. Here, the first speed v 1 Not less than the second speed v of the image data output by each image signal processor 11-12 2 N times. The image processing model can include multiple inherent processing nodes respectively, and complete one cycle of neural network reasoning by sequentially executing the relevant operations of each processing node. The switching node can be obtained by screening from each processing node of each image processing model according to the goals of delay priority and / or storage space priority, and correspondingly distributed according to the delay requirements of each neural network model and / or the total cache data volume of intermediate data generated by each neural network model to indicate the opening and closing time of the processing window of each image processing model.
[0043] Specifically, each of the above-mentioned image processing models can respectively include multiple inherent processing nodes, and complete a cycle of neural network reasoning by sequentially executing relevant operations of each processing node.
[0044] For the above embodiments in which the delay requirements of each image processing model are distributed, technicians can, in an offline manner, pre-divide the entire storage space of the network processing unit 30 according to the number N of models to be processed in parallel, so as to determine the first storage space distribution for storing inference data such as the input image data, model parameter data, and intermediate data of each model. In addition, technicians can also determine the total cycle for each of the N models to complete one inference according to the inference cycle for each model to complete one inference, and thereby determine the number of inferences of each model within this total cycle. After that, technicians can respectively select and determine at least one switching node for time-division multiplexing each model from multiple processing nodes of each model according to this total cycle and the number of inferences of each model, and then, based on this first storage space distribution and the switching nodes of each model, use the entire computing resources of the network processing unit 30 to perform multi-model inference on the inference data samples of multiple models, so as to respectively determine the storage space lacking for each model. After that, technicians can optimize the above first storage space distribution according to the total storage space of the network processing unit 30 and the storage space lacking for each model, so as to determine the second storage space distribution that meets the multi-model inference requirements, and thereby determine the static partitioning scheme of the buffers 311 to 312, so as to ensure that the parallel processing of the binocular image data of the XR device can be completed within the specified delay range.
[0045] In addition, for the embodiments of the total cache data volume distribution of the intermediate data generated by each image processing model as described above, those skilled in the art can divide the full storage space of the network processing unit 30 according to the number N of models that need to be processed in parallel, so as to determine the third storage space distribution for storing the inference data of each model. Then, according to the inference cycle of each model as described above, determine the total cycle of multi-model inference, and then determine the number of inferences of each model within the total cycle. After that, those skilled in the art can, based on the third storage space distribution and each processing node of each model, use the full computing resources of the network processing unit 30 to perform model inference on the inference data samples of each model respectively, so as to determine the storage space lacking at each processing node of each model respectively. After that, those skilled in the art can, in the order of the lacking storage space from small to large, divide different numbers of subgraphs for each model according to the processing nodes and conduct inference tests to determine the computing time required for each model respectively. After that, those skilled in the art can determine the maximum number of subgraphs whose number of inferences and computing time meet the total performance requirements of multi-model inference, and according to the positions of the processing nodes corresponding to the maximum number of subgraphs, select and determine the switching nodes that need to cache the minimum amount of intermediate data from the processing nodes of each model respectively. Here, the total performance requirements of the multi-model inference can be characterized by the sum of the number of inferences and computing time of each image processing model. By selecting the positions of the processing nodes corresponding to the maximum number of subgraphs to determine the switching nodes of each model, the present invention can further reduce the amount of intermediate data time-division multiplexed by each image processing model, thereby further reducing the area, power consumption and cost of the network processing unit 30 and the image signal processors 11-12.
[0046] Thus, the present invention can respectively determine at least one switching node from multiple processing nodes of each image processing model according to the specific objectives of latency priority and / or storage space priority, and determine the storage space corresponding to each image processing model, so as to sequentially perform time-division multiplexed inference of each image processing model by using the full computing resources of the network processing unit 30, so as to realize parallel inference with optimized latency of multiple models and / or parallel inference with optimized storage space.
[0047] As Figure 2 shown, in the process of parallel processing of the two-channel image data of the binocular camera, the two image signal processors 11-12 can be respectively connected to the left-eye camera and the right-eye camera of the XR display device to respectively obtain the left-eye image data and the right-eye image data collected by the image sensors (Sensors) of the left-eye camera and the right-eye camera, and perform preprocessing on them. After that, the two image signal processors 11-12 can, within the preset exposure time t 1 inside, in the read-write mode of ping-pong buffer, preprocess the left-eye image data and the right-eye image data according to the above-mentioned second speed v 2It is continuously transmitted line by line to the corresponding ISP buffers 111 to 112. Here, for the image processing function of noise elimination and / or mosaic elimination, the left-eye image data and right-eye image data collected by the binocular camera of the XR display device can be the original image data with noise signals and / or mosaics.
[0048] Specifically, in the above-mentioned ping-pong buffer reading and writing mode, the image signal processors 11 to 12 can synchronously write the left-eye image data collected by the left-eye camera and the right-eye image data collected by the right-eye camera line by line to the corresponding ISP buffers 111 to 121. In response to the preset exposure time t 1 ends, the input buffer spaces of the ISP buffers 111 to 121 will be filled with the 1st to Mth row image data of the left-eye image and the right-eye image at the same time. At this time, the image signal processors 11 to 12 can send an interrupt instruction to the network processing unit 30 to notify it that the above exposure time t 1 is a fixed time slice, and at the above first speed v 1 (v 1 ≥N·v 2 ) writes the 1st to Mth row image data cached on the ISP buffers 111 to 121 line by line to the corresponding buffers 311 to 312 on the network processing unit 30.
[0049] After that, in response to reaching the first switching node that triggers the first model The network processing unit 30 can first determine that the first processing window of the first model is opened, so as to read the 1st to Mth row image data that has been written in the left-eye image from the first buffer 311, and drive the above MAC array calculation unit 32 and vector processing unit (VPU) 33 to support the first model to process the left-eye image data that has been written currently with the full computing resources of the network processing unit 30, and generate the corresponding first intermediate data.
[0050] Then, in response to reaching the second switching node that triggers the second model The network processing unit 30 can determine that the first processing window of the first model is closed, and the second processing window of the second model is opened, so as to first transfer the first model at the first switching node T 11 and the second switching node T 21The first intermediate data generated therebetween is written back to the feature buffer space of the first buffer 311 to free up the computing resources of the network processing unit 30, and then the image data of the first to M rows that has been written to the right-eye image is read from the second buffer 312, and the above-mentioned MAC array computing unit 32 and the vector processing unit (VPU) 33 are driven to switch all the computing resources of the network processing unit 30 to the second model to support the second model to process the right-eye image data that has been written, so as to generate corresponding second intermediate data.
[0051] Further, in the above-mentioned ping-pong buffer reading and writing mode, the image signal processors 11-12 can continue to write the left-eye image data collected by the left-eye camera and the right-eye image data collected by the right-eye camera to the corresponding ISP buffers 111-121 row by row while the network processing unit 30 reads the image data cached in the ISP buffers 111-121, so as to realize the dynamic synchronization of reading and writing data. In response to the preset exposure time 2t 1 ends, the input buffer space of the ISP buffers 111-121 will be filled with the image data of the (M + 1)-th to 2M-th rows of the left-eye image and the right-eye image again. The image signal processors 11-12 can issue an interrupt instruction to the network processing unit 30 again as described above, notifying it to continue with the above exposure time t 1 as a fixed time slice, at the above first speed v 1 (v 1 ≥N·v 2 ) write the image data of the (M + 1)-th to 2M-th rows cached in the ISP buffers 111-121 to the corresponding buffers 311-312 on the network processing unit 30 row by row, so as to prevent the left-eye image data of the (2M + 1)-th to 3M-th rows written to the ISP buffers 111-121 from being backlogged and overflowing during the next exposure time 2t 1 ~3t 1 .
[0052] After that, in response to reaching the first switching node that triggers the above first model again the network processing unit 30 can determine that the second processing window of the above second model is closed, and the first processing window of the first model is opened again, so as to first set the second model at the second switching node T 21 and the first switching node T 12The second intermediate data generated therebetween is written back to the feature buffer space of the second buffer 312 to free up the computing resources of the network processing unit 30. Then, the image data of the M+1 to 2M lines that have been written to the left-eye image is read from the first buffer 311, along with the first intermediate data generated by the first neural network inference in the previous round. The MAC array computing unit 32 and the vector processing unit (VPU) 33 are driven to switch all the computing resources of the network processing unit 30 back to the first model to support the first model to continue processing the left-eye image data that has been written, so as to regenerate the corresponding first intermediate data.
[0053] And so on. The two image processing models in the network processing unit 30 can alternately perform neural network inference on the left-eye image and the right-eye image window by window according to the predetermined switching nodes in the ping-pong buffer reading and writing manner, so as to meet the requirement of parallel processing of binocular image data in the XR device by configuring a relatively small ISP buffer space (for example: 2M lines of image data).
[0054] Please further refer to Figure 3 , Figure 3 which shows a schematic diagram of obtaining image data line by line according to some embodiments of the present invention.
[0055] As Figure 3 shown, the above XR display chip provided by the present invention can perform convolution calculations in the illustrated sliding window in the row-first direction from left to right and from top to bottom. This way of reading and calculating image data is on the one hand consistent with the direction of writing data from the image signal processors 11 to 12 to the ISP buffers 111 to 121. By quantitatively configuring the reading and writing speeds of the image signal processors 11 to 12 and the network processing unit 30, the real-time transfer of image data from the image signal processors 11 to 12 to the network processing unit 30 can be realized, thereby reducing the waiting time of image data in the ISP buffers 111 to 121 and reducing the latency of image processing.
[0056] On the other hand, compared with the traditional column - direction calculation method that requires writing and caching all 22 - line data of the entire image before starting the neural network inference calculation, and thus must divide the entire image into multiple shard images to reduce the cache requirement, the row - first image data reading and calculation method adopted by the present invention only needs to write and cache the image data of a few lines (for example: 3 lines corresponding to the sliding window size), and then can perform the neural network inference of the corresponding image in real - time. Therefore, it can greatly reduce the cache space requirements for the ISP caches 111 - 121, thereby significantly reducing the area, power consumption, and cost of the image signal processor, and eliminating the need for tiling the entire image to reduce the requirements for the overall system computing power and data cache space, so as to reduce the requirements for the overall system computing power.
[0057] For details, please refer to Figure 4 , Figure 4 which shows the structural diagram of an XR display chip provided according to some embodiments of the present invention.
[0058] As Figure 4 shown, by adopting the above - mentioned row - first image data reading and calculation method provided by the present invention, except for the internal cache 41 of the network processing unit 40, the XR display chip only needs to configure a small - area and small - capacity (for example: 0.4 MB) ISP cache 42 outside the network processing unit 30 to meet the requirements for parallel processing of binocular image data in the XR device. Therefore, it is beneficial to the further development and application of the network processing unit 30 for parallel processing of multi - channel image data.
[0059] Furthermore, in some embodiments of the invention, each image - processing model may respectively include a multi - layer neural network structure. In response to reading multiple lines (for example: 3 lines) of image data that meet the above - mentioned sliding window size from the corresponding cache 311 or 312, the image - processing model can perform convolution calculations in the row - first direction from left to right and from top to bottom to synchronously complete the neural network inference of the corresponding layers.
[0060] For details, please refer to Figure 5 , Figure 5 which shows a schematic diagram of noise reduction processing provided according to some embodiments of the present invention. Taking Figure 5Taking the AI noise reduction (AIDenoise) process shown as an example, it is divided into two parts, an encoder and a decoder, at the algorithm module level. AI Denoise inference based on machine learning is a process of estimating a potential clean image from an actually observed noisy image. The image processing model can perform feature mapping on the image through the encoder, and then integrate and restore the image features through the decoder to finally output a clean image with noise removed. Specifically, during the process of neural network inference, the image processing model can use a convolution to form a skip connection part between the encoder and the decoder, so that the features extracted by image downsampling are incorporated into the upsampling part to promote the fusion of feature information. After that, by learning the noise in the training image to obtain the potential mapping of the corresponding reference image, the image processing model can perform neural network inference for noise reduction processing on the obtained noisy image according to this potential mapping to finally obtain a clean image with noise removed. Compared with traditional denoising methods, this AI noise reduction not only better preserves the edge texture details of the image, but also can perform parallel computing using the architecture of the network processing unit (NPU), thereby making full use of the hardware performance to accelerate the operation rate of the calculation.
[0061] After that, in response to completing the neural network inference of the preset number of layers L based on the currently written multi-line image data, the image processing model starts to generate result data of the initial multi-line image for the corresponding left-eye / right-eye image. Here, for the above-mentioned image processing function of noise reduction, the result data generated by the image processing model can be noise-reduced image data with noise removed. Correspondingly, for the above-mentioned image processing function of removing mosaics, the result data generated by the image processing model can also be restored image data with mosaics removed.
[0062] After that, the network processing unit 30 can Figure 1 as shown, in the read-write mode of a ping-pong buffer, output the generated result data row by row to the corresponding ISP buffers 112 to 122 at the above first speed v 1 and the ISP buffers 112 to 122 return the received result data of the preset number of rows to the corresponding image signal processors 11 to 12 at the above second speed v 2 so that the image data in each of the image signal processors 11 to 12 can flow through the network processing unit 30, thereby reducing the size requirements for the ISP buffers 112 to 122 and the internal buffers 311 to 312 of the network processing unit 30 and reducing the latency of image processing.
[0063] Specifically, continuing with Figure 3Taking the original image data shown as an example, assuming that its total resolution is more than 1000 lines, the network processing unit 30 can complete the neural network inference of the preset number of layers L when processing the K-th line of image data, and output the processed result data of the 1st to M-th lines. Along with the opening and closing of the corresponding model processing window, it outputs the subsequent processed result data of the M+1-th to 2M-th lines, the 2M+1-th to 3M-th lines, etc. one window by one window. Here, K can be determined by the receptive field of the image processing model. For the same model processing binocular image data, it can have a same K value of about 60 to 100 lines. The value of M can correspond to the number of lines of image data written to the corresponding ISP buffers 111 to 121 during the above fixed time slice t 1 Write out the number of lines of image data (for example: M = 8) to the corresponding buffers 311 to 312 on the network processing unit 30 to prevent the backlog and overflow of image data in the buffers 311 to 312.
[0064] In this way, by quantitatively and at a fixed speed outputting the generated result data line by line to the corresponding ISP buffers 112 to 122 in the ping-pong buffer reading and writing mode, the network processing unit 30 can, while alternately acquiring multiple paths of original image data provided by the multiple image signal processors 11 to 12, equally return the processed result data to each of the image signal processors 11 to 12, thereby realizing the transfer and dynamic balance of image data.
[0065] Furthermore, in some embodiments, in response to completing the neural network inference of the preset number of layers L and outputting the generated result data to the corresponding ISP buffers 112 to 122, the network processing unit 30 can also preferably delete multiple lines of image data only related to the first L-1 layers of neural network inference to save the feature buffer space of the buffers 311 to 312.
[0066] In summary, the above network processing unit 30, XR display chip, XR display device, parallel processing method of image data, and computer-readable storage medium provided by the present invention can all parallelly acquire multiple paths of image data line by line, write each path of image data into the buffer at a preset speed respectively, and then the i-th model reads and processes the i-th path of image data that has been written from the buffer by using all the computing resources of the network processing unit. By adopting these configurations, the present invention can improve the image processing accuracy by canceling the need for tiling the entire image, and reduce the requirements for the overall system computing power and data buffer space, thereby improving the ability of a single network processing unit (NPU) to parallelly process multiple paths of image data under the conditions of the same hardware processing technology and cost, and thus flexibly and fully utilize the overall computing power of the system to reduce the network inference latency to meet the requirements of parallel processing of binocular image data in XR devices.
[0067] Although the methods described above are illustrated and described as a series of acts for simplicity of explanation, it should be understood and appreciated that the methods are not limited by the order of the acts, since according to one or more embodiments, some acts may occur in different orders and / or concurrently with other acts not illustrated and described herein but understood by those skilled in the art.
[0068] Those skilled in the art will appreciate that information, signals, and data can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described above throughout may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0069] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0070] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A network processing unit is connected to multiple image signal processors and is configured with a buffer and multiple pre-trained image processing models. Characterized in that, The network processing unit alternately obtains N-channel image data line by line via N image signal processors and writes each channel of the image data into the buffer at a first speed respectively, where N is an integer greater than 1, and the first speed is not less than N times the second speed at which each image signal processor outputs image data. In response to reaching at least one preset i-switching node, the image processing model i utilizes all the computing resources of the network processing unit to read and process the currently written i-th channel of image data from the buffer, where i is an integer not greater than N. In response to reaching the subsequent i+1 switching node, the network processing unit caches the i-th channel of intermediate data generated by the image processing model i for the next i-switching node to continue reading and processing the subsequently written i-th channel of image data.
2. The network processing unit according to claim 1, Characterized in that, The network processing unit obtains the entire i-th channel of image data line by line via the image signal processor i. The image processing model i reads the previously cached intermediate data i according to multiple corresponding i-switching nodes and processes multiple currently written lines of the i-th channel of image data window by window to update the intermediate data i regarding the i-th channel of image data and / or generate result data i regarding the i-th channel of image data.
3. The network processing unit according to claim 2, Characterized in that, The image processing model i includes a multi-layer neural network structure and is configured to: In response to completing neural network inference of a preset number of layers L based on the currently written multiple lines of image data, output result data line by line to the corresponding image signal processor i at the first speed and delete multiple lines of the image data that only involve the first L-1 layers of neural network inference.
4. The network processing unit according to any one of claims 1 to 3, Characterized in that, Each of the switching nodes is distributed at preset time intervals, or according to the latency requirements of each image processing model, or according to the total cached data volume of each intermediate data.
5. The network processing unit according to claim 4, Characterized in that, The steps of determining the switching nodes distributed according to latency requirements and determining the storage space corresponding to each image processing model include: Dividing the entire storage space of the buffer according to the number of models to be processed in parallel to determine the first storage space distribution for storing the inference data of each image processing model. Determining the total period of multi-model inference according to the inference period of each image processing model and determining the number of inferences of each image processing model within the total period. Determining at least one switching node for time-division multiplexing each image processing model respectively from multiple processing nodes of each image processing model according to the total period and the number of inferences of each image processing model. Based on the first storage space distribution and the switching nodes of each of the image processing models, using all the computing resources of the network processing unit, perform multi-model inference on the inference data samples of the multiple image processing models to determine the storage space lacking for each of the image processing models; and Optimize the first storage space distribution according to the full storage space and the storage space lacking for each of the image processing models to determine a second storage space distribution that meets the multi-model inference requirements.
6. The network processing unit according to claim 4,[[]]END]] wherein,[[]]END]] The steps of determining the switching nodes distributed according to the total cache data volume and determining the storage space corresponding to each of the image processing models include:[[]]END]] Divide the full storage space of the buffer according to the number of models to be processed in parallel to determine a third storage space distribution for storing the inference data of each of the image processing models; Determine the total period of multi-model inference according to the inference period of each of the image processing models, and determine the number of inferences of each of the image processing models within the total period; Based on the third storage space distribution and each of the processing nodes of each of the image processing models, use all the computing resources of the buffer to perform model inference on the inference data samples of each of the image processing models respectively to determine the storage space lacking for each of the image processing models at each of the processing nodes; According to the order of the lacking storage space from small to large, divide different numbers of subgraphs for each of the image processing models according to the processing nodes and perform inference tests to determine the required computing duration of each of the image processing models; Determine the maximum number of subgraphs whose number of inferences and computing duration meet the total performance requirements of the multi-model inference; and According to the positions of the processing nodes corresponding to the maximum number of subgraphs, determine the switching nodes of each of the image processing models respectively.
7. The network processing unit according to any one of claims 1 to 3,[[]]END]] wherein,[[]]END]] The buffer is statically divided into N buffer spaces to cache the currently to-be-processed image data and the intermediate data generated by previous processing.
8. The network processing unit according to claim 1,[[]]END]] wherein,[[]]END]] The network processing unit is connected to a binocular camera via two image signal processors and is configured to:[[]]END]] Alternately acquire the left-eye image data and the right-eye image data of the binocular camera line by line via the two image signal processors, and write the left-eye image data and the right-eye image data into the buffer respectively; In response to reaching the first switching node, use all the computing resources of the network processing unit via the first image processing model to read and process the currently written left-eye image data from the buffer to generate corresponding first intermediate data; In response to reaching the second switching node, cache the first intermediate data, switch all the computing resources of the network processing unit to the second image processing model, and read and process the currently written right-eye image data from the buffer via the second image processing model to generate corresponding second intermediate data; and In response to the next-mentioned first switching node, cache the second intermediate data, read the first intermediate data, and switch all the computing resources of the network processing unit back to the first image processing model. Read and process the further written left-eye image data from the buffer via the first image processing model to update the first intermediate data and / or generate result data regarding the left-eye image data.
9. The network processing unit according to claim 1 or 8, wherein, the left-eye image data and the right-eye image data are raw image data with noise, and the result data regarding the left-eye image data is noise-reduced image data after noise elimination, or the left-eye image data and the right-eye image data are raw image data with mosaics, and the result data regarding the left-eye image data is restored image data after mosaic elimination.
10. An extended reality display chip, wherein, it includes: at least two image signal processors, respectively connected to the left-eye camera and the right-eye camera of the extended reality display device; and the network processing unit according to any one of claims 1 to 9, respectively connected to each of the image signal processors to synchronously acquire the image data output by each of the image signal processors and perform parallel processing on it.
11. An extended reality display device, wherein, it includes: a binocular camera; and the extended reality display chip according to claim 10, wherein the extended reality display chip is connected to the binocular camera to synchronously acquire multiple paths of image data output by it and perform parallel processing on each path of the image data.
12. A method for parallel processing of image data, wherein, it includes the following steps: alternately acquire N paths of image data row by row via N image signal processors and write each path of the image data into the buffer of the network processing unit at a first speed, where N is an integer greater than 1, and the first speed is not less than N times the second speed of the image data output by each of the image signal processors; in response to reaching at least one preset i switching node, read and process the currently written i-th path of image data from the buffer using all the computing resources of the network processing unit via the image processing model i pre-trained and configured in the network processing unit, where i is an integer not greater than N; and in response to reaching the subsequent i + 1 switching node, cache the i-th path of intermediate data generated by the image processing model i for continued reading and processing of the subsequently written i-th path of image data at the next i switching node.
13. A computer-readable storage medium, on which computer instructions are stored, wherein, when the computer instructions are executed by a processor, the method for parallel processing of image data according to claim 12 is implemented.