Multi-model reasoning method and system, network processing unit, extended reality display chip and apparatus, and image processing method

By implementing a multi-model inference method in the network processing unit, switching nodes and storage space are determined according to the goal of delay priority and/or storage space priority, and time division multiplexing inference is used for full computing resources, the problem of insufficient solidification and flexibility of computing resources in the prior art is solved, and the delay and storage space optimization of multiple models is achieved.

WO2025108230A1PCT designated stage expired Publication Date: 2025-05-30GRAVITYXR ELECTRONICS & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/132671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-11-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems such as high computing resource curing, area, power consumption and cost, and insufficient flexibility in multi-scenario in multi-model inference and extended reality display devices, and cannot adapt to the application needs of multiple scenarios and multi-functions.

Method used

By implementing a multi-model inference method in the network processing unit, switching nodes and storage spaces are determined according to the goal of delay priority and/or storage space priority, time division multiplexing inference is used to optimize the delay and storage space usage of multiple models.

Benefits of technology

Delay optimization inference and/or storage space optimization inference for multiple models are realized, which improves computing efficiency, reduces network inference delay, and reduces the requirements for the overall computing power of the system and data cache space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024132671_30052025_PF_FP_ABST
    Figure CN2024132671_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a multi-model reasoning method and system, a network processing unit, an extended reality display chip and apparatus, and an image processing method. The multi-model reasoning method comprises the following steps: acquiring reasoning data of a plurality of models; determining at least one switching node from a plurality of processing nodes of each model on the basis of a target having a time delay priority and / or a storage space priority, and determining a storage space corresponding to each model; writing the reasoning data of each model into a corresponding storage space of a network processing unit; and on the basis of the switching node, using full computing resources of the network processing unit to sequentially perform time-division multiplexing reasoning of the models so as to implement time delay optimization reasoning and / or storage space optimization reasoning of the plurality of models. By using the described configurations, the multi-model reasoning method can flexibly switch time-division multiplexing among a plurality of neural network models according to actual requirements, thereby implementing time delay optimization reasoning and / or storage space optimization reasoning of the plurality of models.
Need to check novelty before this filing date? Find Prior Art

Description

Multi-model reasoning method and system, network processing unit, extended reality display chip and device, image processing method

[0001] This application claims priority to a patent application filed on November 24, 2023, with Chinese application number 202311587154.0, entitled “A multi-model reasoning method, system and storage medium”, and a patent application filed on November 24, 2023, with Chinese application number 202311582338.8, entitled “Network processing unit, extended reality display chip and device, image processing method”. The entire contents of these two applications are incorporated herein by reference. Technical Field

[0002] The present invention relates to neural network reasoning technology and extended reality display technology, and in particular to a multi-model reasoning method, a multi-model reasoning system, a network processing unit, an extended reality display chip, an extended reality display device, a parallel processing method for image data, and a computer-readable storage medium. Background Art

[0003] A network processing unit (NPU) is a processor that uses circuits to simulate the structure of human neurons and synapses. It is widely used in the field of artificial intelligence (AI) technology to simulate how humans process various complex information such as images, sounds, and language.

[0004] With the continuous development of AI technology, the data dimensions, data volume and model complexity involved have shown an explosive growth trend, which has put a huge test on the computing power of NPU. Existing technologies usually use multiple independently packaged NPU chips or multi-core NPU chips to perform multi-model inference of multiple neural network models. However, independently packaged NPU chips have fixed computing and storage resources. On the one hand, there are problems such as chip area, power consumption and cost that are too high, which makes large-scale integration impossible. On the other hand, there is insufficient flexibility and it is impossible to adapt to the increase or decrease in the number of models to allocate computing and storage resources. Therefore, it cannot adapt to the multi-scenario and multi-functional application requirements of extended reality (XR) display devices. In contrast, although the multiple NPU cores (Cores) of a multi-core NPU chip can share some of the chip's storage resources, the number of its NPU cores and the computing resources of each NPU core are independently fixed. Therefore, it also has the above-mentioned problems of area, power consumption, high cost and insufficient flexibility, which makes it impossible to adapt to the multi-scenario and multi-functional application requirements of XR display devices.

[0005] To overcome the above-mentioned defects existing in the prior art, there is an urgent need in the art for a multi-model inference technology, which is used to flexibly switch the time division multiplexing between multiple neural network models according to actual requirements, so as to achieve latency optimization inference and / or storage space optimization inference of multiple models.

[0006] In addition, Extended Reality (XR) display technology is an immersive display technology that creates a digital environment combining reality and virtuality through modern high-tech means centered around computers, bringing seamless conversion between the virtual world and the real world to the experiencer, mainly including multiple implementation methods such as Virtual Reality (VR) display, Augmented Reality (AR) display, and Mixed Reality (MR) display.

[0007] In the prior art in the fields of mobile phones, cameras, surveillance, etc., usually a separate Network Processing Unit (NPU) is configured for each Image Signal Processor (ISP) to perform network inference, thus unable to flexibly and fully utilize the overall computing power of the system, resulting in waste of the overall computing power of the system and increasing network inference latency.

[0008] In addition, after obtaining the entire image data with a resolution of H×W collected by the camera module, the prior art usually needs to perform tiling processing on the entire image through an Image Signal Processor (ISP), and divide it into multiple tiled images of N×M (N < H, M < W) according to the data processing capabilities of the Network Processing Unit (NPU) and store them in the ISP buffer, and then the Network Processing Unit (NPU) reads the image data of each tiled image from the ISP buffer one by one for network inference. On the one hand, this will introduce additional operations and intermediate data such as tiling, zero-padding, windowing, and cropping overlapping images, which increases the requirements for the overall computing power of the system. On the other hand, it also requires the system to have a larger data cache space, which greatly increases the area, power consumption, and cost of the Image Signal Processor (ISP), thus severely restricting the development and application of the existing Network Processing Unit (NPU) for parallel processing of multiple image data.

[0009] In order to overcome the above-mentioned defects of the existing technology, this field urgently needs a parallel processing technology for multi-channel image data. By eliminating the need for tile processing of the entire image, the image processing accuracy can be improved, and the requirements for the overall system computing power and data cache space can be reduced. Therefore, under the conditions of the same hardware processing technology and cost, the ability of a single network processing unit (NPU) to process multi-channel image data in parallel can be improved, and the overall computing power of the system can be flexibly and fully utilized to reduce network inference latency, so as to meet the needs of parallel processing of binocular image data in XR devices. Summary of the Invention

[0010] The following is a brief summary of one or more aspects to provide a basic understanding of these aspects. This summary is not an exhaustive overview of all conceivable aspects and is neither intended to identify key or critical elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that will be provided later.

[0011] In order to overcome the above-mentioned defects of the prior art, the present invention provides a multi-model reasoning method, a multi-model reasoning system, and a computer-readable storage medium, which can determine at least one switching node from multiple processing nodes of each model according to the goals of latency priority and / or storage space priority, and determine the storage space corresponding to each model, and then use the full computing resources of the network processing unit to perform time-division multiplexing reasoning of each model in turn according to the switching node, so as to achieve latency-optimized reasoning and / or storage space-optimized reasoning of multiple models.

[0012] In addition, the present invention also provides a network processing unit, an extended reality display chip, an extended reality display device, a method for parallel processing of image data, and a computer-readable storage medium, which can acquire multiple channels of image data line by line in parallel, write each channel of image data into a buffer at a preset speed, and then use the full computing resources of the network processing unit through the i-model to read and process the written i-th channel of image data from the buffer. By adopting these configurations, the present invention can improve the accuracy of image processing by eliminating the need for tile processing of the entire image, and reduce the requirements for the overall computing power and data cache space of the system, thereby improving the ability of a single network processing unit (NPU) to process multiple channels of image data in parallel under the same hardware processing technology and cost conditions, and thereby flexibly and fully utilizing the overall computing power of the system to reduce network inference latency, so as to meet the needs of parallel processing of binocular image data in XR devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above features and advantages of the present invention will be better understood after reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings. In the drawings, the components are not necessarily drawn to scale, and components with similar related properties or characteristics may have the same or similar reference numerals.

[0014] FIG1 shows an architecture diagram of a network processing unit provided according to some embodiments of the present invention.

[0015] FIG2 shows a flowchart of a multi-model reasoning method according to some embodiments of the present invention.

[0016] FIG3 shows a flowchart of determining switching nodes and storage space distribution for latency optimization reasoning according to some embodiments of the present invention.

[0017] FIG4 shows a schematic diagram of latency optimization reasoning provided according to some embodiments of the present invention.

[0018] FIG5 shows a flowchart of determining a switching node and storage space distribution for storage space optimization reasoning according to some embodiments of the present invention.

[0019] FIG6 shows a schematic diagram of storage space optimization reasoning provided according to some embodiments of the present invention.

[0020] FIG7 shows an architecture diagram of an extended reality display chip provided according to some embodiments of the present invention.

[0021] FIG8 shows a flowchart of parallel processing of multiple channels of image data according to some embodiments of the present invention.

[0022] FIG9 shows a schematic diagram of acquiring image data line by line according to some embodiments of the present invention.

[0023] FIG10 shows a structural diagram of an XR display chip provided according to some embodiments of the present invention.

[0024] FIG11 shows a schematic diagram of noise reduction processing according to some embodiments of the present invention. DETAILED DESCRIPTION

[0025] The following specific embodiments illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Although the description of the present invention will be introduced in conjunction with the preferred embodiment, this does not mean that the features of this invention are limited to this embodiment. On the contrary, the purpose of introducing the invention in conjunction with the embodiment is to cover other options or modifications that may be extended based on the claims of the present invention. In order to provide a deep understanding of the present invention, the following description will include many specific details. The present invention can also be implemented without using these details. In addition, in order to avoid confusion or blurring the focus of the present invention, some specific details will be omitted in the description.

[0026] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0027] Furthermore, the terms "upper," "lower," "left," "right," "top," "bottom," "horizontal," and "vertical" used in the following description should be understood to refer to the orientations depicted in that section and the accompanying drawings. These relative terms are used solely for convenience of description and do not necessarily imply that the devices described herein must be manufactured or operated in a specific orientation. Therefore, they should not be construed as limiting the present invention.

[0028] It will be understood that although the terms "first," "second," "third," etc. may be used herein to describe various components, regions, layers, and / or portions, these components, regions, layers, and / or portions should not be limited by these terms, and these terms are merely used to distinguish different components, regions, layers, and / or portions. Thus, a first component, region, layer, and / or portion discussed below may be referred to as a second component, region, layer, and / or portion without departing from some embodiments of the present invention.

[0029] As mentioned above, the existing technology usually uses multiple independently packaged network processing unit (NPU) chips or multi-core NPU chips to perform multi-model reasoning of multiple neural network models. However, the independently packaged NPU chip has solidified computing and storage resources. On the one hand, there are problems such as chip area, power consumption, and high cost, which cannot be integrated on a large scale. On the other hand, there is insufficient flexibility and it is impossible to adapt to the increase or decrease in the number of models to allocate computing and storage resources. Therefore, it cannot adapt to the multi-scene and multi-functional application requirements of extended reality (XR) display devices. In contrast, although the multiple NPU cores (Core) of a multi-core NPU chip can share part of the chip's storage resources, the number of its NPU cores and the computing resources of each NPU core are independently solidified. Therefore, there are also the above-mentioned problems of area, power consumption, high cost and insufficient flexibility, which cannot adapt to the multi-scene and multi-functional application requirements of XR display devices.

[0030] In order to overcome the above-mentioned defects of the prior art, the present invention provides a multi-model reasoning method, a multi-model reasoning system and a computer-readable storage medium, which can determine at least one switching node from multiple processing nodes of each model according to the goals of latency priority and / or storage space priority, and determine the storage space corresponding to each model, and then use the full computing resources of the network processing unit to perform time-division multiplexing reasoning of each model in sequence according to the switching node, so as to achieve latency-optimized reasoning and / or storage space-optimized reasoning of multiple models.

[0031] In some non-limiting embodiments, the multi-model reasoning method provided by the first aspect of the present invention can be implemented via the multi-model reasoning system provided by the second aspect of the present invention. Specifically, the multi-model reasoning system can be configured in a network processing unit (NPU) chip, which is configured with a memory and a processor. The memory includes but is not limited to the computer-readable storage medium provided by the third aspect of the present invention, on which computer instructions are stored. The processor is connected to the memory and is configured to execute the computer instructions stored on the memory to implement the multi-model reasoning method provided by the first aspect of the present invention.

[0032] The following will describe the working principles of the above-mentioned NPU chip and multi-model reasoning system in conjunction with some embodiments of the multi-model reasoning method. Those skilled in the art will understand that the embodiments of these multi-model reasoning methods are only some non-restrictive implementation methods provided by the present invention, which are intended to clearly demonstrate the main concept of the present invention and provide some specific solutions that are convenient for the public to implement, rather than to limit all functions or all working modes of the NPU chip and multi-model reasoning system. Similarly, the NPU chip and multi-model reasoning system are only a non-restrictive implementation method provided by the present invention, and do not constitute a limitation on the execution subject and execution order of each step in these multi-model reasoning methods.

[0033] Please refer to Figures 1 and 2. Figure 1 shows an architecture diagram of a network processing unit according to some embodiments of the present invention. Figure 2 shows a flow chart of a multi-model reasoning method according to some embodiments of the present invention.

[0034] As shown in Figure 1, in some embodiments of the present invention, the NPU chip 10 integrates a multiply-accumulate (MAC) array computing unit 11, a vector processing unit (VPU) 12, a time division multiplexing (TDM) unit 13, and a static random access memory (SRAM) for storing neural network inference data. Here, the SRAM can adapt to the needs of neural network inference and is specifically divided into a parameter storage space SRAM_W 141 for storing neural network model parameters and an input and intermediate data storage space SRAM_F 142 for storing input data and intermediate data of neural network inference.

[0035] As shown in Figure 2, during the multi-model inference process, the NPU chip 10 can first obtain inference data of multiple models M1~Mm from a signal source such as the image signal processor (ISP) of the XR display device. Here, the inference data can be image data of multiple complete images synchronously obtained by multiple image signal processors, or it can be multiple sliced ​​image data obtained by one image signal processor after slicing (Tile) a complete image according to the cache space size of SRAM_F. The multiple sliced ​​images correspond to a part of the complete image respectively, and the amount of data is no more than the cache space of SRAM_F. The multiple models M1~Mm can be the same neural network model with the same structure, parameters and functions, or they can be different neural network models with different structures, parameters and functions.

[0036] Afterwards, the TDM unit 13 of the NPU chip 10 can search and determine at least one switching node from the multiple processing nodes of each model M1~Mm according to the goals of delay priority and / or storage space priority, and determine the static storage space corresponding to each model M1~Mm. The NPU chip 10 then writes the inference data of each model M1~Mm into the input and intermediate data storage space SRAM_F in the corresponding static storage space for each model M1~Mm to read. Here, the neural network model can respectively include multiple inherent processing nodes, and complete a cycle of neural network inference by sequentially executing the relevant operations of each processing node. The switching node is obtained by screening from each processing node according to the goals of delay priority and / or storage space priority, and is correspondingly distributed according to the delay requirements of each neural network model and / or the total cache data volume of the intermediate data generated by each neural network model to indicate the opening and closing time of the processing window of each neural network model.

[0037] For details, please refer to Figures 3 and 4. Figure 3 shows a flow chart of determining the switch node and storage space distribution for latency optimization reasoning according to some embodiments of the present invention. Figure 4 shows a schematic diagram of latency optimization reasoning according to some embodiments of the present invention.

[0038] In the embodiment shown in FIG3 , the process of determining the storage space distribution for latency-optimized inference can be implemented offline based on a large number of pre-prepared inference data samples. Specifically, the present invention can first assume that there are m neural network models M1-Mm that need to be processed in parallel, and determine the inference cycles P1-Pm (in milliseconds) for each model M1-Mm to complete an inference. Subsequently, the TDM unit 13 in the NPU chip 10 can statically divide the full memory space of the NPU chip 10 into parameter storage spaces W1-Wm for storing parameters of each model, and input and intermediate data storage spaces F1-Fm for storing input data and intermediate data of each model, based on the number of models m that need to be processed in parallel, to determine the first storage space distribution for storing inference data of each model. Here, the parameter storage spaces W1-Wm for each model M1-Mm can be determined by compiling each model M1-Mm separately and setting the target performance FPS (Frames per second) and total bandwidth. The total size should not exceed SRAM_W 141, that is, SUM(W1:Wm)≤SRAM_W. Correspondingly, the input and intermediate data storage space F1~Fm of each model M1~Mm can also be determined by compiling each model M1~Mm separately and setting the target performance FPS (Frames per second) and total bandwidth. The total size should not exceed SRAM_F 142, that is, SUM(F1:Fm)≤SRAM_F.

[0039] Afterwards, the TDM unit 13 can determine the total cycle N of multi-model reasoning based on the reasoning cycles P1-Pm of each model M1-Mm, and thereby determine the number of inferences performed by each model M1-Mm within the total cycle N. Specifically, the total cycle N can be determined by the least common multiple (LCM) of the reasoning cycles P1-Pm of each model M1-Mm, that is, N=LCM(P1, P2,…, Pm)=LCM(P1, LCM(P2,…, LCM(Pm-1, Pm))). The TDM unit 13 can obtain the number of inferences performed by each model i within a total cycle N by calculating Num(Mi)=N / Pi, and thereby determine the timing distribution of each model M1-Mm within a total cycle N (for example, performing multi-model parallel reasoning of model M1 once, model M2 twice, and model M3 four times within a total cycle N).

[0040] Next, the TDM unit 13 can determine, based on the total cycle N and the number of inferences each model M1-Mm needs to complete within a total cycle N, at least one switching node for each model M1-Mm from the multiple processing nodes inherent to each model M1-Mm to time-division multiplex, so as to prioritize meeting the latency requirements (i.e., the interval between completing two neural network inferences) for each model M1-Mm. In this case, the present invention uses two adjacent switching nodes i and i+1 to form a round-robin time slice T for performing neural network inference on the corresponding model Mi. The number of round-robin time slices T for a model Mi within a total cycle N is Num_T(i) = N / T*Num(Mi).

[0041] Afterwards, the TDM unit 13 can use the full computing resources of the NPU chip 10 to perform multi-model inference on the inference data samples of each model M1~Mm based on the above-mentioned first storage space distribution and the switching node i of each model M1~Mm to determine the storage space missing from each model M1~Mm.

[0042] Specifically, the above-mentioned inference data samples can have the same form and size as the inference data for subsequent multi-model inference, and have similar content (for example, both are of the same target size and have image data with noise signals and / or mosaics). The technician can connect the NPU chip 10 to an external buffer (Buffer), and statically divide the cache space of the external buffer into m blocks according to the compilation requirements of each model M1~Mm, and then load the model parameters of each model M1~Mm into the parameter storage spaces W1~Wm inside the NPU chip 10 to complete the initialization of each model M1~Mm.

[0043] Afterwards, the TDM unit 13 can activate the MAC array computing unit 11 and VPU 12 in the NPU chip 10 to operate the NPU chip 10 as shown in FIG4 . In response to activating the first switching node for the first model M1, the MAC array computing unit 11 and VPU 12 can read the model parameters of the first model M1 from the first storage space W1 corresponding to the first model M1 in the first storage space distribution, and read the inference data samples of the first model M1 from the first storage space F1 corresponding to the first model M1 in the first storage space distribution, thereby utilizing the full computing resources of the NPU chip 10 to perform first model inference. In response to activating the second switching node for the subsequent second model M2, the NPU chip 10 can interrupt the first model inference, first write the first intermediate data generated by the first model inference into the first storage space F1 of the SRAM_F 142, then write the overflowed first intermediate data into the corresponding first external storage space, and count the amount of the overflowed first intermediate data to determine the missing storage space S_M1_j of the first model M1, where j represents the jth cycle.

[0044] Furthermore, the read and write operations and calculation operations of the NPU chip 10 can be performed independently and simultaneously. The NPU chip 10 can also preferably write the inference data samples of the second model M2 into the second storage space F2 corresponding to the second model M2 in the first storage space distribution while performing the first model inference. In this way, in response to starting the subsequent second switching node of the second model M2, the MAC array computing unit 11 and the VPU 12 in the NPU chip 10 can read the inference data samples of the second model M2 from the second storage space F2 in real time via on-chip transmission, and switch the full computing resources of the NPU chip 10 to the second model M2 to perform the second model inference, so as to determine the second external storage space S_M2_j used for the second model inference as described above. By synchronously processing the separated data read and write and calculation operations, the present invention can write the inference data of the subsequent model i+1 into the corresponding storage space Fi in advance while performing the neural network inference calculation of the previous model i, thereby further improving the efficiency and real-time performance of multi-model inference.

[0045] Similarly, the NPU chip 10 can perform round-robin calculations according to the fixed time slice T of each model M1~Mm as described above, so as to complete multiple round-robin calculations of all models M1~Mm when the total cycle N / T time slices are consumed. At this time, each model Mi has completed Num(Mi) inference calculations, and the external storage requirement size S_Mi_j required by the NPU chip 10 for each model i within a total cycle N is obtained. Afterwards, the TDM unit 13 can take the maximum value of the external storage space required for each model i in multiple rounds of round-robin calculations, that is, S_Mi=MAX(S_Mi_j), to respectively determine the storage space S_Mi missing for each model M1~Mm in the first storage space distribution.

[0046] Please continue to refer to Figure 3. After determining the storage space S_Mi missing from each model in the first storage space distribution, the TDM unit 13 can optimize the first storage space distribution accordingly to determine a second storage space distribution that meets the multi-model reasoning requirements.

[0047] Specifically, in the process of determining the second storage space distribution, the TDM unit 13 can preferably expand the corresponding input data and intermediate data storage space Fi (i.e., Fi'=Fi+S_Mi) based on the storage space S_Mi that is missing in the model Mi in the above-mentioned first storage space distribution. In addition, in order to ensure that the total size of the input and intermediate data storage space F1'~Fm' of each model M1~Mm in the second storage space distribution does not exceed SRAM_F 142, that is, SUM(F1':Fm')≤SRAM_F, the TDM unit 13 can also correspondingly reduce the input data and intermediate data storage space Fj of the remaining at least one model Mj (i.e., Fj'=Fj-S_Mi) to determine the second storage space distribution in which each model M1~Mm does not lack storage space. If a distribution scheme in which SUM(F1':Fm')≤SRAM_F cannot be obtained, the TDM unit 13 can determine that the on-chip storage space of the NPU chip 10 cannot support time division multiplexing of m models, and the number of models needs to be reduced.

[0048] Furthermore, after optimizing the first storage space distribution to determine the second storage space distribution, the TDM unit 13 may reinitialize the state of each model M1-Mm based on the second storage space distribution and recalculate whether each model M1-Mm in the second storage space distribution still lacks storage space as described above. If each model M1-Mm in the second storage space distribution still lacks storage space, the TDM unit 13 may further optimize the on-chip storage space of the NPU chip 10 as described above until each model Mi has no lack of storage space, that is, S_Mi = 0.

[0049] Please continue to refer to Figures 2 and 4. After determining the switching node for latency optimization reasoning and the storage space corresponding to each model M1~Mm, the NPU chip 10 can synchronously obtain reasoning data about multiple images through a multi-channel image signal processor (ISP), and / or slice (Tile) at least one acquired reasoning data to obtain multi-channel reasoning data. Then, its MAC array computing unit 11 and VPU 12 use the full computing resources of the NPU chip 10 to perform time-division multiplexing reasoning of each model M1~Mm in sequence according to the switching node determined by the TDM unit 13, so as to achieve latency optimization reasoning of multiple models M1~Mm.

[0050] Specifically, during the latency optimization inference process, in response to starting the first switching node of the first model M1, the MAC array computing unit 11 and the VPU 12 can read the inference data of the first model M1 from the third storage space F1' corresponding to the first model M1 in the predetermined second storage space distribution, so as to utilize the full computing resources of the NPU chip 10 to perform the first model inference. Thereafter, in response to starting the second switching node of the subsequent second model M2, the MAC array computing unit 11 and the VPU 12 can interrupt the first model inference, write the third intermediate data generated by the first model inference into the third storage space F1', and read the inference data of the second model from the fourth storage space F2' corresponding to the second model M2 in the second storage space distribution, so as to utilize the full computing resources of the NPU chip 10 to perform the second model inference. Similarly, in response to starting the mth switching node of the subsequent model Mm, the MAC array computing unit 11 and the VPU 12 can interrupt the m-1th model inference as described above, write the intermediate data generated by the m-1th model inference into the storage space F(m-1)', and read the inference data of the model Mm from the storage space Fm' corresponding to the model Mm in the second storage space distribution, so as to utilize the full computing resources of the NPU chip 10 to perform the mth model inference, thereby realizing the delay optimized inference of multiple models M1~Mm, and effectively controlling the delay of each model M1~Mm (that is, the interval between completing two neural network inferences).

[0051] In addition, please refer to Figures 5 and 6. Figure 5 shows a flow chart of determining a switching node and storage space distribution for storage space optimization reasoning according to some embodiments of the present invention. Figure 6 shows a schematic diagram of storage space optimization reasoning according to some embodiments of the present invention.

[0052] As described above, in an embodiment where each model M1-Mm performs time-division multiplexing inference according to a fixed time slice T0 or a time slice T consisting of delay-priority switching nodes, it is very likely that the on-chip storage space of the NPU chip 10 cannot support the time-division multiplexing of m models, resulting in the failure of multi-model reasoning. To this end, in the embodiment shown in Figure 5, the present invention can also determine the switching nodes for storage space optimization reasoning in an offline manner based on a large number of pre-prepared reasoning data samples.

[0053] Specifically, the present invention can still assume that there are m neural network models M1~Mm that need to be processed in parallel, and determine the inference cycle P1~Pm (in ms) for each model M1~Mm to complete an inference. Afterwards, the TDM unit 13 in the NPU chip 10 can statically divide the full storage space of the NPU chip 10 into parameter storage spaces W1~Wm for storing the parameters of each model, and input and intermediate data storage spaces F1~Fm for storing the input data and intermediate data of each model according to the number of models m that need to be processed in parallel, so as to determine the third storage space distribution for storing the inference data of each model M1~Mm. Here, the parameter storage space W1~Wm of each model M1~Mm can be determined by compiling each model M1~Mm separately and setting the target performance FPS (Frames per second) and total bandwidth, and its total size should not exceed SRAM_W 141, that is, SUM(W1:Wm)≤SRAM_W. Correspondingly, the input and intermediate data storage space F1~Fm of each model M1~Mm can also be determined by compiling each model M1~Mm separately and setting the target performance FPS (Frames per second) and total bandwidth. The total size should not exceed SRAM_F 142, that is, SUM(F1:Fm)≤SRAM_F.

[0054] Thereafter, the TDM unit 13 can determine the total cycle N of multi-model reasoning based on the reasoning cycles P1-Pm of each model M1-Mm, and thereby determine the number of inferences performed by each model M1-Mm within the total cycle N, i.e., N = LCM(P1, P2, …, Pm) = LCM(P1, LCM(P2, …, LCM(Pm-1, Pm))). The TDM unit 13 can calculate Num(Mi) = N / Pi to obtain the number of inferences performed by each model i within the total cycle N, and thereby determine the timing distribution of each model M1-Mm within the total cycle N (for example, performing multi-model parallel reasoning of model M1 once, model M2 twice, and model M3 four times within the total cycle N).

[0055] Afterwards, the TDM unit 13 can use the full computing resources of the NPU chip 10 to perform model reasoning on the inference data samples of each model M1~Mm based on the above-mentioned third storage space distribution and the processing nodes inherent in each model M1~Mm, so as to respectively determine the storage space missing from each processing node of each model M1~Mm.

[0056] Specifically, the inference data sample can have the same form and size as the inference data for subsequent multi-model inference, and have similar content (for example, both are of the same target size and have image data with noise signals and / or mosaics). The technician can connect the NPU chip 10 to an external buffer, and statically divide the cache space of the external buffer into m blocks according to the compilation requirements of each model M1~Mm, and then load the model parameters of each model M1~Mm into the parameter storage spaces W1~Wm inside the NPU chip 10 to complete the initialization of each model M1~Mm.

[0057] Afterwards, the TDM unit 13 can activate the MAC array computing unit 11 and VPU 12 in the NPU chip 10 to utilize the full computing resources of the NPU chip 10 to sequentially perform model inference on the inference data samples of each model M1-Mm. Specifically, during the process of traversing the model inference of any model Mi through its processing nodes, the NPU chip 10 can interrupt the current model inference in response to any processing node in the model Mi, first write the fifth intermediate data generated by the model inference into the fifth storage space Fi corresponding to the model Mi in the third storage space distribution, and then write the overflowed fifth intermediate data into the external storage space. Based on the amount of the overflowed fifth intermediate data, the amount of data S_Mi_j overflowed from the model Mi to the external buffer at each processing node j is counted, and the storage space missing from each model M1-Mm at its processing node is thereby determined. Here, j≤Ni, j is the jth node of the model Mi, and Ni is the total number of nodes of the model Mi.

[0058] Continuing with reference to FIG5 , after determining the amount of storage space missing from each processing node for each model M1-Mm, the TDM unit 13 can sort the amount of storage space missing from each processing node for each model M1-Mm, denoted as S_Mi[N]=sort(S_Mi_j), and divide each model into subgraphs of different numbers n in ascending order of the amount of missing storage space. For example, when dividing into two subgraphs, the TDM unit 13 can determine the location of each subgraph based on the processing node S_Mi[0]. For another example, when dividing into three subgraphs, the TDM unit 13 can determine the location of each subgraph based on the processing nodes S_Mi[0] and S_Mi[1]. Similarly, when divided into n subgraphs, the TDM unit 13 can determine the segmentation position of each subgraph according to the processing nodes S_Mi[0]~S_Mi[n-2], and finally obtain the subgraph model Mi=set{S_i(1),S_i(2),...,S_i(n-1)} of different numbers n divided by each model M1~Mm, where S_i(n-1) indicates that the model Mi is divided into a set of n-1 subgraphs.

[0059] Afterwards, the TDM unit 13 can perform inference tests on the subgraph set of the model Mi to determine the required computation time Mi_Tn, and determine the maximum number of subgraphs n that meet the total performance requirements of multi-model inference based on the number of inferences Num(Mi) and the computation time Mi_Tn. max The corresponding processing node positions determine the switching node positions C_i=set{Node1,Node2,…,Node_j} of the model Mi.

[0060] Specifically, during the inference test, the TDM unit 13 may first activate the MAC array computing unit 11 and the VPU 12 in the NPU chip 10, and sequentially perform model inference on the inference data samples of each model M1 to Mm using the full computing resources of the NPU chip 10. In response to the first processing node of any model Mi, the MAC array computing unit 11 and the VPU 12 may read the model parameters of the model Mi from the sixth storage space Wi corresponding to the model Mi in the third storage space distribution, and read the inference data samples of the model Mi from the sixth storage space Fi corresponding to the model Mi in the third storage space distribution, thereby performing the sixth model inference using the full computing resources of the NPU chip 10. Afterwards, in response to the subsequent second processing node of the model Mi, the NPU chip 10 can interrupt the model reasoning of the model Mi by time-division multiplexing multiple models, first write the sixth intermediate data generated by the model reasoning into the sixth storage space Fi, and then re-read the sixth intermediate data from the sixth storage space to continue the model reasoning between the first processing node and the second processing node, and so on, until the complete reasoning of the model Mi is completed to determine the calculation time Mi_Tn required for the model Mi.

[0061] Afterwards, the TDM unit 13 can determine the total performance requirement of multi-model reasoning, that is, m*N, based on the number of models m and the total cycle N of multi-model reasoning, and then determine the maximum number of subgraphs n that meets the total performance requirement of multi-model reasoning (that is, SUM(Num(Mi)*Mi_Tn)≤m*N) max , so according to the maximum number of subgraphs n max The corresponding processing node positions determine the switching node positions of model Mi, Ci = set{Node1, Node2, …, Node_j}. A larger value of n indicates a greater number of subgraphs, allowing for more frequent switching between models M1-Mm for time-division multiplexing multi-model inference, further reducing the latency of each model M1-Mm (i.e., the interval between completing two neural network inferences).

[0062] Please continue to refer to Figures 2 and 5. After determining the switching node for storage space optimization reasoning and the storage space corresponding to each model M1~Mm, the NPU chip 10 can synchronously obtain reasoning data about multiple images through a multi-channel image signal processor (ISP), and / or slice (Tile) at least one acquired reasoning data to obtain multi-channel reasoning data. Then, its MAC array computing unit 11 and VPU 12 use the full computing resources of the NPU chip 10 to perform time-division multiplexing reasoning of each model M1~Mm in turn according to the switching node determined by the TDM unit 13, so as to realize storage space optimization reasoning of multiple models M1~Mm.

[0063] Specifically, during the process of performing storage space optimization inference, in response to starting the third switching node Ci of the third model Mi, the MAC array computing unit 11 and the VPU 12 can read the inference data of the third model Mi from the seventh storage space F1 corresponding to the third model Mi in the predetermined third storage space distribution, so as to utilize the full computing resources of the NPU chip 10 to perform the third model inference. Thereafter, in response to starting the fourth switching node Ci+1 of the subsequent fourth model Mi+1, the MAC array computing unit 11 and the VPU 12 can interrupt the third model inference, write the seventh intermediate data generated by the third model inference into the seventh storage space F1, and read the inference data of the fourth model Mi+1 from the eighth storage space corresponding to the fourth model Mi+1 in the third storage space distribution, so as to utilize the full computing resources of the NPU chip 10 to perform the fourth model inference. Similarly, in response to starting the mth switching node Cm of the subsequent model Mm, the MAC array computing unit 11 and the VPU 12 can interrupt the m-1th model inference as described above, write the intermediate data generated by the m-1th model inference into the storage space F(m-1), and read the inference data of the model Mm from the storage space Fm corresponding to the model Mm in the third storage space distribution, so as to utilize the full computing resources of the NPU chip 10 to perform the mth model inference, thereby realizing the storage space optimized inference of multiple models M1~Mm, and effectively controlling the intermediate data generated by the time-division multiplexing inference of each model M1~Mm.

[0064] In summary, by adopting a switching node determined according to the goals of latency priority and / or storage space priority, the calculation time of each model M1~Mm is divided into multiple small time slots. The above-mentioned multi-model inference method, multi-model inference system and computer-readable storage medium provided by the present invention can rotate the neural network inference calculation of each model M1~Mm by time-division multiplexing inference, and immediately save the intermediate data such as the state information obtained by the current model Mi calculation after the corresponding time slot ends, so as to rotate the neural network inference calculation of the next model M(i+1), thereby effectively controlling the delay of each model M1~Mm and reducing the cache space required for the intermediate data generated by multi-model inference. In addition, since each model M1~Mm caches intermediate data through on-chip storage and shares the full computing resources of the NPU chip 10 for neural network inference, the present invention can effectively shorten the switching time between each model M1~Mm and avoid idle waste of computing resources in the entire multi-model inference system, thereby improving the computing efficiency of each model M1~Mm as a whole to support concurrent inference of multiple models M1~Mm.

[0065] In addition, as described above, after obtaining the entire image data with a resolution of H×W acquired by the camera module, the prior art typically needs to perform tile processing on the entire image through an Image Signal Processor (ISP), and divide it into multiple tile images of N×M (N < H, M < W) according to the data processing capacity of the Network Processing Unit (NPU), and then store them in the ISP buffer. Then, the Network Processing Unit (NPU) reads the image data of each tile image from the ISP buffer one by one for network inference. On the one hand, this will introduce additional operations and intermediate data such as tiling, zero-padding, windowing, and cropping overlapping images, which increases the requirements for the overall computing power of the system. On the other hand, it also requires the system to have a larger data cache space, which greatly increases the area, power consumption, and cost of the Image Signal Processor (ISP), thus severely restricting the development and application of the existing Network Processing Unit (NPU) for parallel processing of multiple image data.

[0066] To overcome the above-mentioned defects existing in the prior art, the present invention provides a network processing unit, an extended reality display chip, an extended reality display device, a method for parallel processing of image data, and a computer-readable storage medium, which can parallelly obtain multiple image data row by row, write each path of image data into the buffer at a preset speed, and then use the full computing resources of the network processing unit by the i model to read and process the written i-th path of image data from the buffer. By adopting these configurations, the present invention can improve the accuracy of image processing by eliminating the need for tile processing of the entire image, and reduce the requirements for the overall computing power and data cache space of the system, thereby improving the ability of a single Network Processing Unit (NPU) to parallelly process multiple image data under the conditions of the same hardware processing technology and cost, and thus flexibly and fully utilize the overall computing power of the system to reduce network inference latency to meet the requirements for parallel processing of binocular image data in extended reality (XR) display devices.

[0067] In some non-limiting embodiments, the above-mentioned method for parallel processing of image data provided by the seventh aspect of the present invention can be implemented via the above-mentioned extended reality (XR) display chip provided by the fifth aspect of the present invention. Specifically, please refer to FIG. 7, which shows an architecture diagram of an extended reality display chip provided by some embodiments of the present invention.

[0068] In the embodiment shown in FIG7 , the XR display chip provided in the fifth aspect of the present invention can be configured in the XR display device provided in the sixth aspect of the present invention, which is provided with a memory (not shown), at least two image signal processors 71-72, and the network processing unit 80 provided in the fourth aspect of the present invention. The memory includes but is not limited to the computer-readable storage medium provided in the fifth aspect of the present invention, on which computer instructions are stored. The at least two image signal processors 71-72 are respectively connected to the left and right cameras of the XR display device to obtain the real-world scene images captured by them. The network processing unit 80 is respectively connected to the memory and each image signal processor 71-72, and is suitable for reading and executing the computer instructions stored on the memory to implement the parallel processing method of image data provided in the fourth aspect of the present invention, thereby alternately obtaining the image data output by each image signal processor 71-72 line by line, and performing parallel processing on the obtained image data.

[0069] The following will describe the working principles of the above-mentioned network processing unit 80, the extended reality display chip, and the extended reality display device in conjunction with some embodiments of the parallel processing method for image data. Those skilled in the art will understand that these embodiments of the parallel processing method are only some non-limiting implementation methods provided by the present invention, which are intended to clearly demonstrate the main concept of the present invention and provide some specific solutions that are convenient for the public to implement, rather than to limit all functions or all working modes of the network processing unit 80, the extended reality display chip, and the extended reality display device. Similarly, the network processing unit 80, the extended reality display chip, and the extended reality display device are also only a non-limiting implementation method provided by the present invention, and do not constitute a limitation on the execution subject and execution order of each step in the parallel processing method for image data.

[0070] Please refer to FIG. 7 and FIG. 8 . FIG. 8 shows a flowchart of parallel processing of image data according to some embodiments of the present invention.

[0071] As shown in FIG7 , the network processing unit 80 provided by the present invention is configured with hardware devices such as caches 811 to 812, a multiply-accumulate (MAC) array computing unit 82, a vector processing unit (VPU) 83, and a register 84. It is also configured with a software program containing multiple pre-trained image processing models, and its computing operations and storage operations can be performed separately and independently. The full computing resource sharing of the network processing unit 80 is achieved through the MAC array computing unit 82 and the vector processing unit (VPU) 83, so as to flexibly support the network processing unit 80 in performing operations such as data reading, computing, and caching.

[0072] In some embodiments, technicians can train and compile the input parameters of the neural network model for image processing such as denoising and / or demosaicing in advance according to the size of the entire real-scene image in an offline manner, and adapt to the specific usage scenarios of the binocular camera in the XR display device, initialize the trained model N (for example: N = 2) times, and obtain N neural network models with the same functions to correspond to each image signal processor 71~72 respectively.

[0073] Those skilled in the art will understand that the above-mentioned scheme of configuring N neural network models with the same functions is only a non-limiting implementation method provided by the present invention, which is intended to clearly demonstrate the main concept of the present invention and provide a specific scheme that is easy for the public to implement, rather than to limit the scope of protection of the present invention.

[0074] Optionally, in other embodiments, those skilled in the art may also configure a variety of neural network models with different parameters, structures and / or functions to adapt to various image processing requirements such as denoising and demosaicing, so as to correspondingly meet the requirements of multifunctional parallel processing of binocular image data in XR devices.

[0075] In addition, in some embodiments, the buffers 811-812 can be configured as static random access memory (SRAM), which is configured at the input end of the network processing unit 80 and is statically divided into N (for example, N=2) independent buffer spaces to adapt to the specific usage scenario of the binocular camera in the XR display device. Furthermore, each buffer space can be preferably divided into an input buffer space, a parameter buffer space, and a feature buffer space to respectively cache the image data currently to be processed by the corresponding image processing model, the model parameter data, and the intermediate data generated by the neural network inference performed by the corresponding image processing model.

[0076] Hereinafter, the two independent buffer spaces are referred to as first buffer 811 and second buffer 812. The first buffer 811 is connected to the image signal processor 71 via ISP buffer 711, while the second buffer 812 is connected to the image signal processor 72 via ISP buffer 721. Both buffers alternately read two channels of image data line by line from the corresponding ISP buffers 711-712 at a preset first speed v1. Each image processing model, based on the corresponding switching node, alternately reads the currently written image data from the corresponding buffer 811-812 and processes the image data using the full computing resources of the network processing unit 80, thereby performing parallel processing of the two channels of image data for the binocular camera. Here, the first speed v1 is not less than N times the second speed v2 at which each image signal processor 71-72 outputs image data. Each image processing model can include multiple inherent processing nodes, and completes a cycle of neural network inference by sequentially executing the relevant operations of each processing node. The switching node can be obtained by screening from each processing node of each image processing model based on the goals of delay priority and / or storage space priority, and correspondingly distributed according to the delay requirements of each neural network model and / or the total cache data volume of intermediate data generated by each neural network model to indicate the opening and closing time of the processing window of each image processing model.

[0077] Specifically, each of the above-mentioned image processing models may include a plurality of inherent processing nodes, and complete a cycle of neural network reasoning by sequentially executing related operations of each processing node.

[0078] With respect to the above-mentioned embodiment of distribution according to the time delay requirements of each image processing model, the technician can divide the full storage space of the network processing unit 80 in advance according to the number of models N that need to be processed in parallel in an offline manner to determine the first storage space distribution for storing the input image data, model parameter data, intermediate data and other inference data of each model. In addition, the technician can also determine the total cycle of the N model inferences that complete one inference based on the inference cycle of each model, and thereby determine the number of inferences of each model within the total cycle. Afterwards, the technician can select and determine at least one switching node for time-division multiplexing each model from the multiple processing nodes of each model based on the total cycle and the number of inferences of each model, and then, based on the first storage space distribution and the switching nodes of each model, use the full computing resources of the network processing unit 80 to perform multi-model inference on the inference data samples of multiple models to determine the storage space missing for each model. Afterwards, technicians can optimize the above-mentioned first storage space distribution based on the total storage space of the network processing unit 80 and the storage space lacking in each model to determine the second storage space distribution that meets the multi-model reasoning requirements, and thereby determine the static partitioning scheme of the buffers 811~812 to ensure that the parallel processing of the binocular image data of the XR device can be completed within the specified delay range.

[0079] Furthermore, with respect to the above-described embodiment of distributing the total cached data volume of intermediate data generated by each image processing model, a technician can partition the network processing unit 80's full storage space based on the number of models N to be processed in parallel, thereby determining a third storage space distribution for storing the inference data of each model. As described above, the total multi-model inference cycle can be determined based on the inference cycle of each model, and the number of inferences performed by each model within the total cycle can be determined accordingly. Subsequently, based on this third storage space distribution and the processing nodes of each model, the technician can utilize the full computing resources of the network processing unit 80 to perform model inference on the inference data samples of each model to determine the amount of storage space lacking for each model at each processing node. Furthermore, the technician can divide each model into different numbers of subgraphs based on the processing nodes in ascending order of the amount of storage space lacking, and perform inference testing to determine the computation time required for each model. Furthermore, the technician can determine the maximum number of subgraphs whose inference times and computation time meet the overall performance requirements for multi-model inference. Based on the location of the processing node corresponding to this maximum number of subgraphs, the technician can select a switching node from each model's processing nodes that requires caching the minimum amount of intermediate data. Here, the total performance requirement for multi-model inference can be represented by the cumulative sum of the number of inferences and computational time for each image processing model. By selecting the location of the processing node corresponding to the maximum number of subgraphs to determine the switching nodes for each model, the present invention can further reduce the amount of intermediate data generated by time-division multiplexing each image processing model, thereby further reducing the area, power consumption, and cost of the network processing unit 80 and image signal processors 71-72.

[0080] Therefore, the present invention can determine at least one switching node from multiple processing nodes of each image processing model according to the specific goals of latency priority and / or storage space priority, and determine the storage space corresponding to each image processing model, so as to utilize the full computing resources of the network processing unit 80 to perform time-division multiplexing reasoning of each image processing model in sequence, so as to realize parallel reasoning of multiple models with optimized latency and / or parallel reasoning of optimized storage space.

[0081] As shown in Figure 8, during the parallel processing of two channels of image data from the binocular cameras, the two image signal processors 71-72 can be connected to the left and right cameras of the XR display device, respectively, to obtain the left and right image data captured by the image sensors of the left and right cameras, respectively, and pre-process them. Subsequently, the two image signal processors 71-72 can continuously transfer the pre-processed left and right image data line by line to the corresponding ISP buffers 711-712 at the aforementioned second speed v2 within a preset exposure time t1, using a ping-pong buffer read-write method. Here, for the image processing functions of noise removal and / or mosaic removal, the left and right image data captured by the binocular cameras of the XR display device can be raw image data with noise signals and / or mosaics.

[0082] Specifically, using the aforementioned ping-pong buffer read / write method, the image signal processors 71-72 can synchronously write the left-eye image data captured by the left camera and the right-eye image data captured by the right camera to the corresponding ISP buffers 711-721, line by line. Upon completion of the preset exposure time t1, the input buffers of the ISP buffers 711-721 will be filled with the 1st to Mth rows of image data for both the left and right images. At this point, the image signal processors 71-72 can issue an interrupt command to the network processing unit 80, instructing it to write the 1st to Mth rows of image data cached in the ISP buffers 711-721, line by line, to the corresponding buffers 811-812 on the network processing unit 80, using the aforementioned exposure time t1 as a fixed time slice and at the aforementioned first speed v1 (v1 ≥ N·v2).

[0083] Afterwards, in response to reaching the predetermined first switching node T that triggers the first model 11 The network processing unit 80 may first determine that the first processing window of the first model is open, thereby reading the 1st to Mth rows of image data currently written into the left-eye image from the first buffer 811, and driving the MAC array computing unit 82 and the vector processing unit (VPU) 83 to support the first model in processing the currently written left-eye image data with the full computing resources of the network processing unit 80, and generating corresponding first intermediate data.

[0084] Then, in response to reaching the predetermined second switching node T that triggers the second model 21 The network processing unit 80 may determine that the first processing window of the first model is closed and the second processing window of the second model is opened, thereby first switching the first model at the first switching node T 11 and the second switching node T 21 The first intermediate data generated between the two modes is written back to the feature buffer space of the first buffer 811 to free up the computing resources of the network processing unit 80, and then the 1st to Mth rows of image data currently written in the right-eye image are read from the second buffer 812, and the above-mentioned MAC array computing unit 82 and vector processing unit (VPU) 83 are driven to switch the full computing resources of the network processing unit 80 to the second model to support the second model to process the currently written right-eye image data to generate corresponding second intermediate data.

[0085] Furthermore, using the aforementioned ping-pong buffering read / write method, the image signal processors 71-72 can simultaneously read image data from the ISP buffers 711-721 while simultaneously writing line-by-line the left image data captured by the left camera and the right image data captured by the right camera to the corresponding ISP buffers 711-721, thereby achieving dynamic synchronization of data reading and writing. Upon expiration of the preset exposure time 2t1, the input buffers of the ISP buffers 711-721 will once again be filled with image data from lines M+1 to 2M of the left and right images. As described above, the image signal processors 71~72 can again send an interrupt instruction to the network processing unit 30, notifying it to continue to use the above-mentioned exposure time t1 as a fixed time slice, and write out the M+1~2Mth lines of image data cached in the ISP buffers 711~721 to the corresponding buffers 811~812 on the network processing unit 80 line by line at the above-mentioned first speed v1 (v1≥N·v2), thereby preventing the 2M+1~3Mth lines of left-eye image data written into the ISP buffers 711~721 within the next exposure time 2t1~3t1 from being accumulated and overflowing.

[0086] Afterwards, in response to reaching the first switching node T that triggers the first model again 12 The network processing unit 80 may determine that the second processing window of the second model is closed and the first processing window of the first model is opened again, thereby first switching the second model to the second switching node T 21 and the first switching node T 12The second intermediate data generated between the two processes is written back to the feature buffer space of the second buffer 812 to free up the computing resources of the network processing unit 80, and then the image data of the left-eye image currently written in the M+1 to 2M rows and the first intermediate data generated by the previous round of the first neural network inference are read from the first buffer 811, and the MAC array computing unit 82 and the vector processing unit (VPU) 83 are driven to switch all the computing resources of the network processing unit 80 back to the first model to support the first model to continue processing the currently written left-eye image data to regenerate the corresponding first intermediate data.

[0087] Similarly, the two image processing models in the network processing unit 80 can perform neural network inference on the left-eye image and the right-eye image alternately on a window-by-window basis according to the predetermined switching nodes and the reading and writing method of the ping-pong cache, thereby meeting the demand for parallel processing of binocular image data in the XR device by configuring less ISP cache space (for example: 2M lines of image data).

[0088] Please further refer to FIG. 9 , which shows a schematic diagram of acquiring image data line by line according to some embodiments of the present invention.

[0089] As shown in FIG9 , the XR display chip provided by the present invention can use the sliding window shown in the figure to perform convolution calculations from left to right and from top to bottom in a row-first direction. This method of reading and calculating image data is consistent with the direction in which the image signal processors 71-72 write data to the ISP buffers 711-721. By quantitatively configuring the read and write speeds of the image signal processors 71-72 and the network processing unit 80, real-time flow of image data from the image signal processors 71-72 to the network processing unit 80 can be achieved, thereby reducing the waiting time for image data to be cached in the ISP buffers 711-721, thereby reducing the latency of image processing.

[0090] On the other hand, compared with the traditional column-wise calculation method, which requires writing and caching all 22 rows of data of the entire image before starting the neural network inference calculation, the entire image must be divided into multiple tiled images to reduce the cache requirements. The row-priority image data reading and calculation method adopted by the present invention only needs to write and cache a few rows (for example: 3 rows corresponding to the sliding window size) of image data to carry out the neural network inference of the corresponding image in real time. Therefore, the cache space requirements of the ISP buffers 711~721 can be greatly reduced, so as to significantly reduce the area, power consumption and cost of the image signal processor, and eliminate the need for tile processing of the entire image to reduce the requirements for the overall system computing power and data cache space, so as to reduce the requirements for the overall system computing power.

[0091] Please refer to FIG. 10 for details, which shows a structural diagram of an XR display chip provided according to some embodiments of the present invention.

[0092] As shown in FIG10 , by adopting the above-mentioned row-priority image data reading and calculation method provided by the present invention, in addition to the internal buffer 91 of the network processing unit 90, the XR display chip only needs to configure a small-area, small-capacity (for example, 0.4 MB) ISP buffer 92 outside the network processing unit 80 to meet the requirements of parallel processing of binocular image data in the XR device, thereby facilitating the further development and application of the network processing unit 80 to parallel process multi-channel image data.

[0093] Furthermore, in some embodiments of the present invention, each image processing model may include a multi-layer neural network structure. In response to reading multiple lines (e.g., three lines) of image data that meet the above sliding window size from the corresponding buffer 811 or 812, the image processing model may perform convolution calculations from left to right and from top to bottom in a row-first direction to simultaneously complete the neural network inference of the corresponding number of layers.

[0094] Please refer to Figure 11 for details, which shows a schematic diagram of the noise reduction processing provided according to some embodiments of the present invention. Taking the AI ​​noise reduction (AIDenoise) processing shown in Figure 11 as an example, it is divided into two parts, an encoder and a decoder, at the algorithm module level. AIDenoise reasoning based on machine learning is a process of estimating a potential clean image from an actually observed noisy image. The image processing model can perform feature mapping on the image through the encoder, and then integrate and restore the image features through the decoder to finally output a clean image with noise eliminated. Specifically, in the process of performing neural network reasoning, the image processing model can use convolution to form a jump connection part between the encoder and the decoder, so that the features extracted by image downsampling are integrated into the upsampling part to promote the fusion of feature information. Afterwards, by learning the noise in the training image, the potential mapping of the corresponding reference image is obtained, and the image processing model can perform neural network reasoning on the acquired noisy image according to the potential mapping to finally obtain a clean image with noise eliminated. Compared with traditional denoising methods, this AI denoising not only better preserves the edge texture details of the image, but also can use the architecture of the network processing unit (NPU) for parallel computing, thereby making full use of hardware performance to speed up the computing operation rate.

[0095] After completing the neural network inference for the preset number of layers L based on the currently written multiple lines of image data, the image processing model begins generating result data for the initial multiple lines of the corresponding left / right image. For the aforementioned image processing function for noise removal, the result data generated by the image processing model can be denoised image data. Similarly, for the aforementioned image processing function for de-mosaicing, the result data generated by the image processing model can also be restored image data after de-mosaicing.

[0096] Afterwards, the network processing unit 80 can output the generated result data line by line to the corresponding ISP buffers 712~722 at the first speed v1 according to the reading and writing method of the Ping-Pong buffer as shown in Figure 7, and the ISP buffers 712~722 return the received preset number of lines of result data to the corresponding image signal processors 71~72 at the second speed v2, so that the image data in each image signal processor 71~72 can flow through the network processing unit 80, thereby reducing the size requirements of the ISP buffers 712~722 and the internal buffers 811~812 of the network processing unit 80, and reducing the latency of image processing.

[0097] Specifically, continuing with the example of the original image data shown in FIG9 , assuming its total resolution is above 1000 lines, the network processing unit 80 can complete the neural network inference of the preset number of layers L when processing the Kth line of image data, and output the processing result data of lines 1 to M. As the corresponding model processing window is opened and closed, the subsequent processing result data of lines M+1 to 2M, 2M+1 to 3M, and so on is output window by window. Here, K can be determined by the receptive field of the image processing model. For the same model processing binocular image data, it can have the same K value of approximately 60 to 100 lines. The value of M can correspond to the number of image data lines written by the ISP buffers 711 to 721 to the corresponding buffers 811 to 812 on the network processing unit 80 at the fixed time slice t1 (for example, M = 8), to prevent backlog and overflow of image data in buffers 811 to 812.

[0098] In this way, by outputting the generated result data line by line to the corresponding ISP buffers 712~722 in a quantitative and constant speed manner according to the reading and writing method of the ping-pong cache, the network processing unit 80 can alternately obtain the multiple original image data provided by the multiple image signal processors 71~72 while returning the processing result data to each image signal processor 71~72 in equal amounts, thereby realizing the flow and dynamic balance of image data.

[0099] Furthermore, in some embodiments, in response to completing the neural network inference of a preset number of layers L and outputting the generated result data to the corresponding ISP buffers 712 to 722, the network processing unit 80 may also preferably delete multiple rows of image data that only involve the first L-1 layers of neural network inference to save feature buffer space of the buffers 811 to 812.

[0100] In summary, the network processing unit 80, XR display chip, XR display device, parallel processing method of image data, and computer-readable storage medium provided by the present invention can all acquire multiple channels of image data line by line in parallel, write each channel of image data into a buffer at a preset speed, and then use the full computing resources of the network processing unit through the i-model to read and process the written i-th channel of image data from the buffer. By adopting these configurations, the present invention can improve the accuracy of image processing by eliminating the need for tile processing of the entire image, and reduce the requirements for the overall system computing power and data cache space. Therefore, under the conditions of the same hardware processing technology and cost, the ability of a single network processing unit (NPU) to process multiple channels of image data in parallel is improved, and the overall computing power of the system is flexibly and fully utilized to reduce network inference delay, so as to meet the needs of parallel processing of binocular image data in XR devices.

[0101] Although the above methods are illustrated and described as a series of acts for simplicity of explanation, it is to be understood and appreciated that these methods are not limited by the order of the acts, as some acts may occur in a different order and / or concurrently with other acts from those illustrated and described herein or not illustrated and described herein but understandable to those skilled in the art according to one or more embodiments.

[0102] Those skilled in the art will appreciate that information, signals, and data may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips cited throughout the foregoing description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0103] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.

[0104] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read and write information from / to the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside in a user terminal as discrete components.

[0105] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any media that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0106] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-model reasoning method, characterized in that: The following steps are involved: Get inference data for multiple models; According to the goals of delay priority and / or storage space priority, at least one switching node is determined from a plurality of processing nodes of each of the models, and a storage space corresponding to each of the models is determined; Writing the inference data of each model into the corresponding storage space of the network processing unit respectively; as well as According to the switching node, the full computing resources of the network processing unit are utilized to perform time-division multiplexing reasoning of each of the models in turn, so as to realize delay optimization reasoning and / or storage space optimization reasoning of the multiple models.

2. The multi-model reasoning method according to claim 1, characterized in that: According to the delay priority target, the steps of determining at least one switching node from a plurality of processing nodes of each of the models, and determining a storage space corresponding to each of the models include: Dividing the full storage space of the network processing unit according to the number of models that need to be processed in parallel to determine a first storage space distribution for storing the inference data of each of the models; Determine the total cycle of multi-model reasoning according to the reasoning cycle of each model, and determine the number of reasoning times of each model in the total cycle; According to the total cycle and the number of inferences of each model, respectively determine at least one switching node for time-division multiplexing each model from a plurality of processing nodes of each model; Based on the first storage space distribution and the switching nodes of each of the models, using the full computing resources of the network processing unit, multi-model reasoning is performed on the reasoning data samples of the multiple models to determine the storage space missing from each of the models; and According to the total storage space of the network processing unit and the storage space lacking in each of the models, the first storage space distribution is optimized to determine a second storage space distribution that meets the multi-model reasoning requirements.

3. The multi-model reasoning method according to claim 2, characterized in that: The step of performing multi-model reasoning on the reasoning data samples of the multiple models based on the first storage space distribution and the switching nodes between the models using the full computing resources of the network processing unit to determine the storage space missing from each model includes: connecting the network processing unit to an external buffer; In response to starting a first switching node of the first model, reading an inference data sample of the first model from a first storage space corresponding to the first model in the first storage space distribution, performing first model inference using full computing resources of the network processing unit, and counting a first external storage space used for the first model inference; and A storage space lacking in the first model is determined according to the first external storage space.

4. The multi-model reasoning method according to claim 3, characterized in that: The step of, in response to starting the first switching node of the first model, reading the inference data sample of the first model from the first storage space corresponding to the first model in the first storage space distribution, performing the first model inference using the full computing resources of the network processing unit, and counting the first external storage space used for the first model inference comprises: In response to starting a second switch node of a subsequent second model, interrupting the first model reasoning, first writing the first intermediate data generated by the first model reasoning into the first storage space, and then writing the overflowed first intermediate data into the external storage space; and The first external storage space is determined according to the data amount of the overflowed first intermediate data.

5. The multi-model reasoning method according to claim 4, characterized in that: The step of performing multi-model reasoning on the reasoning data samples of the multiple models based on the first storage space distribution and the switching nodes between the models using the full computing resources of the network processing unit to determine the storage space missing from each model also includes: While performing the first model inference, writing the inference data samples of the second model into a second storage space corresponding to the second model in the first storage space distribution; and In response to the second switching node, the inference data sample of the second model is read from the second storage space, the full computing resources of the network processing unit are switched to the second model, the second model inference is performed using the full computing resources, and the second external storage space used for the second model inference is determined.

6. The multi-model reasoning method according to claim 3, characterized in that: The inference data includes model parameter data, input data and intermediate data, the storage space distribution involves parameter storage space, input data storage space and intermediate data storage space of each model, and the step of optimizing the first storage space distribution according to the total storage space of the network processing unit and the storage space lacking of each model to determine the second storage space distribution that meets the multi-model inference requirement includes: According to the storage space lacking in each of the models in the first storage space distribution, the corresponding input data storage space and intermediate data storage space are expanded, and / or the input data storage space and intermediate data storage space of the remaining models are reduced to determine a second storage space distribution in which each of the models has sufficient storage space.

7. The multi-model reasoning method according to claim 1, characterized in that: The step of performing time-division multiplexing reasoning of each of the models in sequence according to the switching node using the full computing resources of the network processing unit to achieve delay optimization reasoning and / or storage space optimization reasoning of the multiple models includes: In response to starting the first switch node of the first model, reading the inference data of the first model from a third storage space corresponding to the first model in a predetermined second storage space distribution, and performing the first model inference using the full computing resources of the network processing unit; and In response to starting a subsequent second switching node of a second model, the first model reasoning is interrupted, the third intermediate data generated by the first model reasoning is written into the third storage space, and the reasoning data of the second model is read from the fourth storage space corresponding to the second model in the second storage space distribution, and the second model reasoning is performed using the full computing resources of the network processing unit to achieve delay optimized reasoning of the multiple models.

8. The multi-model reasoning method according to claim 1, characterized in that: According to the storage space priority target, the steps of determining at least one switching node from a plurality of processing nodes of each of the models, and determining the storage space corresponding to each of the models include: Dividing the full storage space of the network processing unit according to the number of models that need to be processed in parallel to determine the distribution of the third storage space for storing the inference data of each of the models; Determine the total cycle of multi-model reasoning according to the reasoning cycle of each model, and determine the number of reasoning times of each model in the total cycle; Based on the third storage space distribution and the processing nodes of the models, using the full computing resources of the network processing unit, respectively performing model reasoning on the reasoning data samples of the models to determine the storage space lacking in the processing nodes of the models; In the order of the lack of storage space from small to large, dividing each model into different numbers of subgraphs according to the processing nodes, and performing reasoning tests to determine the calculation time required for each model; Determine the maximum number of subgraphs whose inference times and computation time meet the total performance requirements of the multi-model inference; and According to the positions of the processing nodes corresponding to the maximum number of subgraphs, the switching nodes of each of the models are determined respectively.

9. The multi-model reasoning method according to claim 8, characterized in that: The step of performing model reasoning on the reasoning data samples of each model based on the third storage space distribution and each processing node of each model by using the full computing resources of the network processing unit to determine the storage space lacking of each model at each processing node comprises: connecting the network processing unit to an external buffer; Using the full computing resources of the network processing unit, perform model reasoning on the reasoning data samples of each model in turn; In response to any of the processing nodes of any of the models, interrupt the current model reasoning of the corresponding model, first write the fifth intermediate data generated by the model reasoning into the fifth storage space corresponding to the model in the third storage space distribution, and then write the overflowed fifth intermediate data into the external storage space; and According to the data amount of the overflowed fifth intermediate data, a fifth external storage space lacking of the model at the processing node is determined.

10. The multi-model reasoning method according to claim 8, characterized in that: The step of performing an inference test to determine the computation time required for each model includes: Using the full computing resources of the network processing unit, perform model reasoning on the reasoning data samples of each model in turn; In response to a first processing node of any of the models, reading inference data samples of the model from a sixth storage space corresponding to the model in the third storage space distribution, and performing sixth model inference using full computing resources of the network processing unit; In response to a subsequent second processing node of the model, interrupting the model reasoning of the model, first writing the sixth intermediate data generated by the model reasoning into the sixth storage space, and then reading the sixth intermediate data from the sixth storage space, so as to continue the model reasoning between the first processing node and the second processing node, and so on, until the complete reasoning of the model is completed; and The computation time required for the model is determined based on the time required to complete the complete inference of the model.

11. The multi-model reasoning method according to claim 8, characterized in that: Before determining the maximum number of subgraphs whose inference times and calculation duration meet the total performance requirements of the multi-model inference, the multi-model inference method further includes the following steps: The total performance requirement of the multi-model reasoning is determined according to the number of the models and the total cycle of the multi-model reasoning.

12. The multi-model reasoning method according to claim 1, characterized in that: The step of performing time-division multiplexing reasoning of each of the models in sequence according to the switching node using the full computing resources of the network processing unit to achieve delay optimization reasoning and / or storage space optimization reasoning of the multiple models includes: In response to starting a third switch node of a third model, reading inference data of the third model from a seventh storage space corresponding to the third model in a predetermined third storage space distribution, and performing inference of the third model using full computing resources of the network processing unit; and In response to starting the fourth switching node of the subsequent fourth model, the reasoning of the third model is interrupted, the seventh intermediate data generated by the reasoning of the third model is written into the seventh storage space, and the reasoning data of the fourth model is read from the eighth storage space corresponding to the fourth model in the third storage space distribution, and the fourth model reasoning is performed using the full computing resources of the network processing unit to achieve storage space optimized reasoning of the multiple models.

13. The multi-model reasoning method according to claim 1, characterized in that: The step of obtaining the inference data of multiple models includes: Synchronously acquiring inference data about multiple channels of images via multiple channel image signal processors; and / or Data slicing is performed on the at least one path of acquired reasoning data to obtain multiple paths of reasoning data.

14. A multi-model reasoning system, characterized in that: include: a memory having computer instructions stored thereon; as well as A processor is connected to the memory and configured to execute computer instructions stored in the memory to implement the multi-model reasoning method according to any one of claims 1 to 13.

15. The multi-model reasoning system according to claim 14, characterized in that: The multi-model reasoning system is configured in a network processing unit.

16. A network processing unit connected to multiple image signal processors and equipped with a buffer and multiple pre-trained image processing models, characterized in that: The network processing unit alternately obtains N channels of image data line by line through N channels of image signal processors, and writes each channel of the image data into the buffer at a first speed, wherein N is an integer greater than 1, and the first speed is not less than N times a second speed at which each channel of the image signal processor outputs the image data. In response to reaching at least one preset i-switching node, the image processing model i uses the full computing resources of the network processing unit to read and process the i-th channel of image data currently written from the buffer, where i is an integer not greater than N, In response to reaching the subsequent i+1 switching node, the network processing unit caches the i-th path of intermediate data generated by the image processing model i, waiting for the next i switching nodes to continue reading and processing the i-th path of image data that is subsequently written.

17. The network processing unit according to claim 16, characterized in that: The network processing unit acquires the entire i-th channel of image data line by line via the image signal processor i, The image processing model i reads the previously cached intermediate data i according to multiple corresponding i switching nodes, and processes the multiple rows of the i-th image data that have been written currently, window by window, to update the intermediate data i about the i-th image data, and / or generate result data i about the i-th image data.

18. The network processing unit according to claim 17, characterized in that: The image processing model i includes a multi-layer neural network structure and is configured as follows: In response to completing the preset number L of neural network inferences based on the multiple lines of image data currently written, the result data is output line by line to the corresponding image signal processor i at the first speed, and the multiple lines of image data that only involve the first L-1 layers of neural network inference are deleted.

19. The network processing unit according to any one of claims 16 to 18, characterized in that: The switching nodes are distributed according to preset time intervals, or according to the delay requirements of the image processing models, or according to the total cached data volume of the intermediate data.

20. The network processing unit according to claim 19, characterized in that: The steps of determining the switching nodes distributed according to the delay requirements and determining the storage space corresponding to each of the image processing models include: Dividing the entire storage space of the buffer according to the number of models that need to be processed in parallel to determine a first storage space distribution for storing the inference data of each of the image processing models; Determine the total cycle of multi-model reasoning according to the reasoning cycle of each of the image processing models, and determine the number of reasoning times of each of the image processing models within the total cycle; Determine at least one switching node for time-division multiplexing each of the image processing models from a plurality of processing nodes of each of the image processing models according to the total cycle and the number of inferences of each of the image processing models; Based on the first storage space distribution and the switching nodes of each of the image processing models, using the full computing resources of the network processing unit, multi-model reasoning is performed on the inference data samples of the multiple image processing models to determine the storage space missing from each of the image processing models; and Based on the full storage space and the storage space lacking in each of the image processing models, the first storage space distribution is optimized to determine a second storage space distribution that meets the multi-model reasoning requirements.

21. The network processing unit according to claim 19, characterized in that: The steps of determining the switching nodes distributed according to the total cache data volume and determining the storage space corresponding to each of the image processing models include: Dividing the entire storage space of the buffer according to the number of models that need to be processed in parallel to determine the distribution of the third storage space for storing the inference data of each of the image processing models; Determine the total cycle of multi-model reasoning according to the reasoning cycle of each of the image processing models, and determine the number of reasoning times of each of the image processing models within the total cycle; Based on the third storage space distribution and the processing nodes of the image processing models, the full computing resources of the buffer are used to perform model reasoning on the inference data samples of the image processing models to determine the storage space lacking in the processing nodes of the image processing models; In the order of the lack of storage space from small to large, dividing each of the image processing models into different numbers of subgraphs according to the processing nodes, and performing reasoning tests to determine the calculation time required for each of the image processing models; Determine the maximum number of subgraphs whose inference times and computation time meet the total performance requirements of the multi-model inference; and According to the positions of the processing nodes corresponding to the maximum number of sub-graphs, the switching nodes of the image processing models are determined respectively.

22. The network processing unit according to any one of claims 16 to 18, characterized in that: The buffer is statically divided into N buffer spaces to buffer the image data currently to be processed and the intermediate data generated by previous processing.

23. The network processing unit according to claim 16, characterized in that: The network processing unit is connected to the binocular camera via two image signal processors and is configured as follows: Alternately acquiring left-eye image data and right-eye image data of the binocular camera line by line through the two image signal processors, and writing the left-eye image data and the right-eye image data into the buffer respectively; In response to reaching the first switching node, using the full computing resources of the network processing unit via the first image processing model, reading and processing the currently written left-eye image data from the buffer to generate corresponding first intermediate data; In response to reaching a second switching node, the first intermediate data is cached, and the full computing resources of the network processing unit are switched to a second image processing model, and the right-eye image data currently written is read and processed from the cache via the second image processing model to generate corresponding second intermediate data; as well as In response to the next first switching node, the second intermediate data is cached, the first intermediate data is read, and the full computing resources of the network processing unit are switched back to the first image processing model, and the further written left-eye image data is read and processed from the cache via the first image processing model to update the first intermediate data, and / or generate result data about the left-eye image data.

24. The network processing unit according to claim 16 or 23, characterized in that: The left-eye image data and the right-eye image data are original image data with noise, and the result data about the left-eye image data and the right-eye image data are noise-reduced image data after the noise is eliminated, or The left-eye image data and the right-eye image data are original image data with mosaics, and the result data regarding the left-eye image data and the right-eye image data are restored image data with the mosaics removed.

25. An extended reality display chip, characterized in that: include: At least two image signal processors, connected to a left camera and a right camera of the augmented reality display device respectively; as well as The network processing unit according to any one of claims 16 to 24 is connected to each of the image signal processors, respectively, to synchronously acquire image data output by each of the image signal processors, and process the image data in parallel.

26. An extended reality display device, characterized in that: include: Binocular camera; as well as The extended reality display chip as described in claim 25, wherein the extended reality display chip is connected to the binocular camera to synchronously acquire multiple channels of image data outputted by the binocular camera and perform parallel processing on each channel of the image data.

27. A method for parallel processing of image data, characterized in that: The following steps are involved: Alternately acquiring N channels of image data line by line through N channels of image signal processors, and writing each channel of the image data into a buffer of a network processing unit at a first speed, wherein N is an integer greater than 1, and the first speed is not less than N times a second speed at which each channel of the image signal processor outputs the image data; In response to reaching at least one preset i-switching node, reading and processing the currently written i-th channel of image data from the buffer by using the image processing model i pre-trained and configured in the network processing unit and utilizing the full computing resources of the network processing unit, wherein i is an integer not greater than N; and In response to reaching the subsequent i+1 switching node, the i-th channel of intermediate data generated by the image processing model i is cached, waiting for the next i switching nodes to continue reading and processing the i-th channel of image data that is subsequently written.

28. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the multi-model reasoning method according to any one of claims 1 to 13, or the parallel processing method for image data according to claim 27 is implemented.

Citation Information

Patent Citations

  • Image acquisition card, image acquisition method and image acquisition system

    CN114286035A

  • Neural network model processing method and device

    CN115511693A

  • Neural network algorithm task compiling and executing method, chip and electronic equipment

    CN116010049A

  • Concurrent optimization of machine learning model performance

    US20210019652A1