Multi-processor system and image processing method for multi-lens camera

By using multiprocessor systems and image processing methods, the problem of insufficient computing resources and performance of traditional centralized processor systems in multi-lens cameras has been solved, memory bandwidth and computing speed have been improved, and efficient image processing has been achieved.

CN115942103BActive Publication Date: 2026-02-27COOL BOLE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111106724.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-12
Filing Date
2021-09-22
Publication Date
2026-02-27
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

Traditional centralized processor systems are insufficient in computing resources, memory bandwidth, computing speed, and processing performance when processing image data from multi-lens cameras, making it difficult to meet the demands of high resolution and multiple lenses.

Method used

A multiprocessor system is adopted, which connects multiple processor elements and links to achieve distributed processing of image data. Multiple processor elements are used to transmit data in a single direction, and image stitching and mixing are performed through image processing methods, thereby improving memory bandwidth and computing speed.

Benefits of technology

It improves the memory bandwidth and computing speed of multi-lens cameras, enhances processing performance, overcomes the shortcomings of traditional centralized processor systems, and achieves efficient image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115942103B_ABST
    Figure CN115942103B_ABST
Patent Text Reader

Abstract

A multi-processor system and image processing method for a multi-lens camera are disclosed. The system includes a plurality of processor elements and a plurality of links. Each processor element includes a plurality of input / output (I / O) ports and a processing unit. The multi-lens camera captures a field of view having an X degree horizontal field of view and a Y degree vertical field of view, where X<=360 and Y<180. Each link connects one of the plurality of I / O ports of one of the plurality of processor elements to one of the plurality of I / O ports of another of the plurality of processor elements, such that each processor element is connected to one or two adjacent processor elements with two or more links, each link being configured to transmit data in a single direction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to image processing, and in particular, to a multi-processor system and an image processing method for a multi-lens camera. BACKGROUND

[0002] Conventionally, a centralized processor system is used to process image data generated by a multi-lens camera, which has the advantages of low hardware cost and low power consumption. However, as the resolution of images or videos captured by a camera is getting higher and the number of lenses is getting larger, the centralized processor system has the disadvantage of high cost in terms of computing resources, memory bandwidth, computing speed and processing performance. Therefore, there is an urgent need for a new architecture and method to solve the above problems, and thus the present application is proposed. SUMMARY

[0003] In view of the above problems, one of the objectives of the present application is to provide a multi-processor system for a multi-lens camera, so as to increase memory bandwidth and computing speed and improve processing performance.

[0004] According to an embodiment of the present application, a multi-processor system is provided, which includes a plurality of processor elements and a plurality of links. The plurality of processor elements is coupled to a multi-lens camera, each processor element including a plurality of input / output (I / O) ports and a processing unit coupled to the plurality of I / O ports. The multi-lens camera captures a field of view having a horizontal field of view of X degrees and a vertical field of view of Y degrees, where X <= 360 and Y <= 180. Each link connects one of the plurality of I / O ports of one of the plurality of processor elements to one of the plurality of I / O ports of another of the plurality of processor elements, such that each processor element is connected to one or two adjacent processor elements with two or more links, each link being configured to transmit data in a single direction.

[0005] According to another embodiment of the present application, an image processing method is provided for a multi-processor system coupled to a multi-lens camera, the multi-lens camera capturing a field of view having a horizontal field of view of X degrees and a vertical field of view of Y degrees, the multi-processor system including a plurality of processor elements and a plurality of links, each processor element being connected to one or two adjacent processor elements with two or more links, each link being configured to transmit data in a single direction, the method including, at a processor element j: obtaining n j lens images captured from the multi-lens camera; in a first transmission phase, selectively receiving and transmitting data with respect to the n jEach lens image and zero or more input and output first edge image data related to the overlapping area are passed to one or two neighboring processor elements; according to a first vertex list, the n j Based on the lens image and the input edge image data, the optimal joining coefficients for multiple control regions responsible for the overlapping area are determined; in a second transmission stage, input and output joining coefficients are selectively transmitted and received from one or two neighboring processor elements; and, based on the first vertex list, the multiple optimal joining coefficients, the input joining coefficients, the input edge image data, and the n... j Each shot produces n images. j n working face images, where n j >=1, X<=360, and Y<180. Wherein, based on the plurality of responsible control regions, the plurality of output joining coefficients are selected from the plurality of optimal joining coefficients; wherein, the first vertex sublist contains a plurality of first vertices having a first data structure, the plurality of first data structures defining n j A first vertex mapping between a lens image and a projected image, wherein the projected image is a working surface image from all processor elements.

[0006] The above and other objects and advantages of the present invention will be described in detail below with reference to the following illustrations, detailed descriptions of the embodiments, and the scope of the claims. Attached Figure Description

[0007] FIG. 1 According to the present invention, a block architecture diagram of a multiprocessor system suitable for multi-lens cameras is shown.

[0008] FIG. 2A This shows two different side views of a quad-lens camera.

[0009] FIG. 2B This shows two different side views of a triple-lens camera.

[0010] FIG. 2C This shows two different side views of a twin-lens camera.

[0011] FIG. 3A This illustrates the relationship between a cube structure and a sphere.

[0012] FIG. 3B This shows an example of a triangular mesh used to model the surface of a sphere.

[0013] FIG. 3C This shows an example of a polygonal grid used to compose / model the multiple rangeside panoramic images.

[0014] FIG. 3DAn example of a long rectangular panoramic image with four overlapping regions A(0)-A(3) and twenty control regions R(1)-R(20).

[0015] FIG. 4A According to an embodiment of the present application, a block diagram of a four-processor system suitable for a four-lens camera is shown.

[0016] FIG. 4B-FIG. 4C According to the present application, a flowchart of an image processing method suitable for the multi-processor system 100 / 400 / 800 / 900 is shown.

[0017] FIG. 5A An example of how the mismatched image defects of an object are improved after modifying the texture coordinates of all vertices in each lens image according to the best blending coefficients is shown.

[0018] FIG. 5B An example of the relationship between the target vertex P and ten control regions R(1)-R(10) in the lens image iK1 is shown.

[0019] FIG. 6 According to an embodiment of the present application, a schematic diagram of a graphics processing unit (GPU) is shown.

[0020] FIG. 7A According to an embodiment of the present application, a flowchart of a method for determining the best blending coefficients of all control regions in a measurement mode is shown.

[0021] FIG. 7B According to an embodiment of the present application, a flowchart of a method for performing the coefficient determination operation of step S712 by the GPU 132 is shown.

[0022] FIG. 7C An example of a link metric is established.

[0023] FIG. 8 According to an embodiment of the present application, a block diagram of a two-processor system suitable for a four-lens camera is shown.

[0024] FIG. 9A According to an embodiment of the present application, a block diagram of a three-processor system suitable for a three-lens camera is shown.

[0025] FIG. 9B An example of a wide-angle image with two overlapping regions A(0)-A(1) and ten control regions R(1)-R(10) is shown.

[0026] REFERENCE NUMERALS:

[0027] 11A cube architecture

[0028] 11B, 11C architecture of the camera

[0029] 12 sphere

[0030] 50 ideal imaging position

[0031] 54 image center

[0032] 55 object

[0033] 56, 57 lens center

[0034] 58 actual imaging position

[0035] 100 multi-processor system

[0036] 110 multi-lens camera

[0037] 110A multi-lens camera

[0038] 110B three-lens camera

[0039] 110C two-lens camera

[0040] 120 main processor element

[0041] 121-12m auxiliary processor elements

[0042] 120-1-12m-1 processing units

[0043] 131 image signal processor

[0044] 132 image processing unit

[0045] 133 image quality enhancement unit

[0046] 134 encoding and transmission unit

[0047] 140-14m lens groups

[0048] 151-157, 15t0 / t m I / O port

[0049] 160-16m local non-volatile memory

[0050] 170-17m local volatile memory

[0051] 180 receiver

[0052] 400 four-processor system

[0053] 481, 482, 901 chain

[0054] 610 rasterization engine

[0055] 620 Texture Mapping Circuit

[0056] Texture mapping engine 621-622

[0057] 630 Hybrid Unit

[0058] 650 measurement units

[0059] 800 dual-processor system

[0060] 900 Triple Processor System Detailed Implementation

[0061] Throughout this specification and in subsequent claims, the singular forms of terms such as "a" and "the" have both singular and plural meanings, unless otherwise specified herein. The definitions of related terms used throughout this specification and in subsequent claims are as follows, unless otherwise specified herein. Throughout this specification, circuit elements with the same function use the same reference numerals.

[0062] One of the features of this invention is the use of a multiprocessor architecture to process image data from a multi-lens camera, thereby making full use of computing resources, increasing memory bandwidth and computing speed, and improving processing performance.

[0063] FIG. 1 According to the present invention, a block architecture diagram of a multiprocessor system suitable for multi-lens cameras is shown. (Reference) FIG. 1 A multiprocessor system 100, used to process image data from a multi-lens camera 110, includes a main processor component (PC) 120, m auxiliary processor components 121-12m, and multiple links (shown in...). FIG. 4A / FIG. 8 / FIG. 9A The camera 110 has (m+1) lens groups 140 to 14m, which are coupled to the main processor element 120 and the m auxiliary processor elements 121 to 12m via input / output ports 151, where n0, ..., n m>=1. According to the present invention, the implementation of each processor element 12j can use an integrated circuit device, such as a programmable processor, an application-specific integrated circuit (ASIC), and a dedicated processor element, where 0 <= j <= m. In one embodiment, the multiprocessor system 100 / 400 / 800 / 900 is a system-on-a-chip (SOC) that can be integrated into a computing device (such as a mobile phone, tablet computer, or wearable computer) to perform image processing on images or videos captured by the camera 110.

[0064] The multi-lens camera 110 can capture still or moving images. The multi-lens camera 110 can be a panoramic camera (e.g.,...). FIG. 2A A quad-lens camera (110A) or a wide-angle camera (such as...) FIG. 2B-FIG. 2C The system uses a triple-lens camera 110B and a dual-lens camera 110C. Correspondingly, a receiver 180 receives an encoded video stream (en or en0~enm) from the system 100 to form a projected image, which can be a panoramic image or a wide-angle image. FIG. 2B-FIG. 2C In this design, the two edges or working surfaces of the camera 110B / C architecture 11B / C (lenses K0 and K1 are respectively mounted on these two edges) form a 120-degree angle. Note that this 120-degree angle is merely an example and not a limitation of the invention; in actual implementation, the two edges of the architecture 11B / C may form other angles. This multi-lens camera 110 can simultaneously capture a field of view encompassing an X-degree horizontal field of view (FOV) and a Y-degree vertical FOV to generate multiple lens images, where X <= 360 and Y < 180, for example, 360 × 160 or 180 × 90. For example, FIG. 2A The camera 110A comprises four lenses (not shown) mounted on the four working surfaces of a cubic structure 11A to simultaneously capture a field of view with a horizontal FOV of 360 degrees and a vertical FOV of 90 degrees, thereby generating four-lens images. Note that the invention does not limit the number of lenses in the camera 110 as long as a field of view with an X-degree horizontal FOV and a Y-degree vertical FOV can be captured, where X <= 360 and Y < 180. A necessary condition is that there should be sufficient overlap between the fields of view of any two adjacent lenses to facilitate image stitching.

[0065] Each processor element 12j includes a processing unit 12j-1, a local nonvolatile memory (NVM) 16j, a local volatile memory (VM) 17j, and multiple I / O ports 151 to 15t. j Where 0 <= j <= m and tj >=3. Each processor element 12j operates together with its own local non-volatile memory 16j and local volatile memory 17j. Note the number t of the I / O ports of each processor element 12j. j The changes are influenced by factors such as whether the primary processor element 120 incorporates working surface images / enhanced images from other auxiliary processor elements, the size of the number m, the camera type (panoramic or wide-angle camera), the type of processor element 12j (primary or auxiliary), and its position relative to the primary processor element 120. The aforementioned I / O ports 151 to 15t... j It can be a conventional design, or it can include circuitry to modify data to conform to a high-speed serial interface standard, such as, but not limited to, the Mobile Industry Processor Interface (MIPI). In the following embodiments, each I / O port 151-15t j Taking MIPI ports as an example, it should be understood that I / O ports 151 to 15t are... j This is not a limitation; existing or future high-speed serial interface standards are applicable to the concepts of this invention. I / O ports 151~15t j Each can be configured as an input MIPI port or an output MIPI port. Each link connects one of the plurality of I / O ports of one of the plurality of processor elements to one of the plurality of I / O ports of the other of the plurality of processor elements.

[0066] Each processing unit 12j-1 includes an image signal processor (ISP) 131, an image processing unit (GPU), an image quality enhancement (IQE) unit 133, and an encoding and transmission unit 134. Note that the IQE unit 133 is not an essential element of this invention, therefore... FIG. 1 The text is represented by dashed lines. Local volatile memory 170-17m stores various data used by the aforementioned processing units 120-1-12m-1, such as programs or image data acquired from camera 110. Local non-volatile memory 160-16m contains multiple programs or instructions, which are executed by the aforementioned processing units 120-1-12m-1 respectively, for the purpose of execution... FIG. 4B-FIG. 4C and FIG. 7A-FIG. 7Ball the steps of the methods to be described later. In addition, the processing unit 120-1 executes the programs stored in the local non-volatile memory 160 to perform various data processing or operations on the image data acquired from the n0 lenses 140 of the camera 110 and stored in the local volatile memory 170, and to control all the operations of the multi-processor system 100, including the camera 110 and the m auxiliary processor elements 121-12m. The processing unit 12j-1 of each auxiliary processor element 12j operates independently of the main processor element 120, executes the programs stored in the local non-volatile memory 16j to perform various data processing or operations on the image data acquired from the n j j lenses 14j of the camera 110 and stored in the local volatile memory 17j, where 1<=j<=m. Specifically, the plurality of ISPs 131 respectively receive electrical signals from the image sensors (not shown) of the corresponding lens group 14j of the camera 110 through their own input ports 151, and convert the plurality of electrical signals into a plurality of digital lens images; based on the plurality of lens images, the original and corrected main vertex secondary lists, the m original and m corrected auxiliary vertex secondary lists, the plurality of GPUs 132 respectively execute the programs stored in the local non-volatile memories 160-16m to determine the optimal blending coefficients and perform rasterization, texture mapping and blending operations to form a main working face image F0 and m auxiliary working face images F1-Fm (to be described later). The plurality of IQE units 133 perform contrast enhancement, low-pass filtering and sharpness processing on the main working face image F0 and the m auxiliary working face images F1-Fm to generate a main enhanced image F0' and m auxiliary enhanced images F1'-Fm'. Finally, the encoding and transmission units 134 of each processor element 120-12m respectively encode the main enhanced image F0' and the m auxiliary enhanced images F1'-Fm' into (m+1) encoded video streams en0-enm, and then transmit the (m+1) encoded video streams en0-enm to a receiver 180 to generate a panoramic image or a wide-angle image.

[0067] The following terms are defined throughout the specification and the following claims as follows, unless otherwise expressly provided herein. The term "texture coordinates" refers to coordinates in a texture space (e.g., a texture image or a lens image). The term "rasterization operation" refers to a computation process of mapping scene geometry (or a projected image) to texture coordinates of each lens image. The term "tranceive" refers to: transmit and / or receive. The term "projection" refers to: flattening a surface of a sphere into a two-dimensional (2D) plane, e.g., a projected image.

[0068] The multi-processor system 100 of the present application is applicable to various projection methods. The various projection methods include, but are not limited to, equirectangular projection, cylindrical projection, and modified cylindrical projection. The modified cylindrical projection includes, but is not limited to, Miller projection, Mercator projection, Lambert cylindrical equal area projection, Pannini projection, etc. Accordingly, the projected image includes, but is not limited to, an equirectangular projected image, a cylindrical projected image, and a modified cylindrical projected image. FIG. 3B-FIG. 3C Regarding the equirectangular projection. The embodiments of the cylindrical projection and the modified cylindrical projection are well known to those skilled in the art and are not described herein.

[0069] For the sake of clarity and convenience, the following examples and embodiments are described with respect to the equirectangular projection and the equirectangular panoramic image, and assume that the panoramic camera 110A includes four lenses K0-K3 and is mounted on four working surfaces (right, left, front, back) of the cubic structure 11A. It should be noted that the operation of the multi-processor system 100 of the present application is also applicable to a wide-angle camera, the cylindrical projection, and the modified cylindrical projection.

[0070] FIG. 3A The relationship between a cubic structure 11A and a sphere 12 is shown. As shown in FIG. 2A and FIG. 3A The four lens cameras 110A include four lenses K0-K3 and are mounted on four working surfaces of a cubic structure 11A, any two adjacent surfaces of the four working surfaces are substantially orthogonal, e.g., respectively facing the longitude 0 degrees, 90 degrees, 180 degrees, and 270 degrees of the virtual sphere 12, to simultaneously capture a field of view with a 360-degree horizontal FOV and a 90-degree vertical FOV, and generate four lens images. Please refer to FIG. 3Dpixels in regions A(0) - A(3) are overlapped by two lens / texture images, while pixels in other regions b0 - b3 are from a single lens / texture image. Therefore, the multiple overlapping regions 13 can be stitched or blended to form an equirectangular panorama image. Generally, the size of the overlapping regions (e.g. A(0) - A(3) of the equirectangular panorama image) varies with the lens FOA, lens sensor resolution and lens mounting angle of the camera 110A. FIG. 3D

[0071] The processing pipeline of the multi-processor system 100 is divided into an offline phase and an online phase. In the offline phase, once the lens FOA, lens sensor resolution and lens mounting angle of the camera 110A are fixed, the size of the overlapping regions A(0) - A(3) is also fixed. Then, the four lenses of the camera 110A are calibrated respectively, and then a suitable image registration technique is used to generate a raw vertex list, where each vertex in the raw vertex list provides a mapping between the multiple equirectangular panorama images and the multiple lens images (or between the multiple equirectangular coordinates and the multiple texture coordinates). For example, a sphere 12 with a radius of 2 meters (r = 2) is divided into many circles, which are used as longitude and latitude, and their intersection points are considered as calibration points. The four lenses K0 - K3 capture the calibration points, and the positions of the calibration points on the multiple lens images are known. Then, because the view angle of the calibration points and the multiple texture coordinates are linked, the mapping between the multiple equirectangular panorama images and the multiple lens images can be established. In this specification, a calibration point with the above mapping is defined as a "vertex". In short, in the offline phase, the relationship between each vertex in the multiple equirectangular panorama images and the multiple lens images is calibrated to obtain the raw vertex list.

[0072] FIG. 3B A triangular mesh is shown to model a sphere surface. Referring to FIG. 3B A triangular mesh is used to model the surface of a sphere 12. FIG. 3C A polygon mesh is shown to compose / model the multiple equirectangular panorama images. By performing an equirectangular projection on FIG. 3B the triangular mesh, the polygon mesh is generated, where FIG. 3C the polygon mesh is a collection of quadrilaterals and / or triangles. FIG. 3C

[0073] ​​During the offline phase, based on the geometry of the multiple rectangular panoramic images and the multiple lens images, the polygonal mesh ( FIG. 3C For each vertex of the polygon mesh, its equidistant rectangular coordinates and texture coordinates are calculated to generate an initial vertex list. During the offline phase, after the lens FOA, lens sensor resolution, and lens mounting angle of the camera 110A are fixed, this initial vertex list only needs to be calculated / generated once. This initial vertex list is a list of multiple vertices that form the polygon mesh. FIG. 3C The original list of vertices contains multiple quadrilaterals and / or triangles, each vertex defined by a corresponding data structure. This data structure defines the vertex mapping between a destination space and a texture space (or between the multiple rectangular coordinates and the texture coordinates). Table 1 shows an example of the data structure for each vertex in the original vertex list.

[0074]

[0075]

[0076] FIG. 4A According to one embodiment of the present invention, a block architecture diagram of a four-processor system suitable for a four-lens camera is shown. (Refer to...) FIG. 4A The four-processor system 400 processes image data from a four-lens camera 110A. It includes a main processor element 120, three auxiliary processor elements 121-123, and nine links, each link connecting two processor elements. Link 481 is not mandatory. Through input ports 151, the four processor elements 120-123 are connected to the four lenses K0-K3 of the camera 110A, respectively. (For clarity and convenience,...) FIG. 4A Only four processor elements 120-123, their included I / O ports, and the nine links are shown, and will be described in detail later. In this embodiment, the main processor element 120 includes six I / O ports 151-155 and 157, the auxiliary processor elements 121 / 123 include five I / O ports 151-155, and the auxiliary processor element 122 includes six I / O ports 151-156. I / O ports 151, 153-154, and 157 are designated as input ports, while I / O ports 152 and 155-156 are designated as output ports. Each link connects an input port of one processor element to an output port of another processor element. Note that although... FIG. 4A , FIG. 8 and FIG. 9A The display shows multiple links between the input port of one processor element and the output port of another processor element. However, in reality, these multiple links refer to the same links. These multiple links simply represent data transmission between the two processor elements at multiple different points in time via the same links; for example...FIG. 4A The two links 482 between the I / O port 152 of the processor element 120 and the I / O port 154 of the processor element 121 represent data transmissions through the links 482 at two different points in time (transmission phases two and three); FIG. 8 The four links 801 between the I / O port 152 of the processor element 121 and the I / O port 153 of the processor element 120 represent data transmissions through the links 801 at four different points in time (at different transmission phases one to four).

[0077] It is also noted that, FIG. 4A , FIG. 8 and FIG. 9A The display of the connection topology of the processor elements does not necessarily represent the physical arrangement of the processor elements. Similarly, in the entire specification and the subsequent claims, it should be understood that the use of the terms “adjacent” or “nearby” in relation to the processor elements refers to the connection topology, not a specific physical arrangement.

[0078] In the offline phase, because the four-processor system 400 comprises four processor elements 120-123, the original vertex list (as in Table 1) is divided into four original vertex sublists, i.e., one original primary vertex sublist or0and three original auxiliary vertex sublists or1-or3, according to the equidistant rectangular coordinates, and the four original vertex sublists or0-or3are respectively stored in the four local non-volatile memories 160-163 for subsequent image processing.

[0079] FIG. 4B-FIG. 4C To show the flowchart of the image processing method applicable to the multi-processor system 100 / 400 / 800 / 900 according to the present application, the following describes the operation of the four-processor system 400 according to the flowchart of FIG. 4B-FIG. 4C In step S402, the ISP 131 of each processor element 12j receives and parses the MIPI packet through the MIPI input port 151, converts the electrical signals into a digital lens image iKj(including the electrical signals related to the image sensor of a corresponding lens of the camera 110A), and stores the digital lens image iKjin the local volatile memory 17j according to the data type (such as 0x2A) of the packet header, where 0<=j<=3.

[0080] In the example of FIG. 4A each processor element is responsible for a single overlap region, such as FIG. 3DIn one embodiment, the four processor elements 120-123 respectively acquire the lens images iK0-iK3, and are respectively responsible for the four overlapping areas A(3), A(0), A(1), and A(2). For the sake of clarity and convenience of description, the following examples and embodiments are described assuming that the four processor elements 120-123 respectively acquire the lens images iK0-iK3, and are respectively responsible for the overlapping areas A(0), A(1), A(2), and A(3).

[0081] In step S404 (i.e., the transmission phase one), in order to form the aforementioned four overlapping areas, each processor element needs to transmit the left edge data of its own lens image to a neighboring processor element through the output port 155, and receive the left edge data of a neighboring right lens image from another neighboring processor element through the input port 153. For each processor element, the output left edge data of its own lens image is located at the opposite edge of the overlapping area it is responsible for; meanwhile, the right edge data of its own lens image and the left edge data of the neighboring right lens image it receives form the overlapping area it is responsible for, and the size of the aforementioned right edge data of its own lens image and the left edge data of the neighboring right lens image is related to the size of the overlapping area it is responsible for; for example, the edge data rK0' and iK1' form the overlapping area A(0), and the size of the edge data rK0' and iK1' is related to the size of A(0). As described above, once the lens FOA of the camera 110A, the lens sensor resolution, and the lens mounting angle are fixed, the sizes of the overlapping areas A(0)-A(3) are determined. Assuming that the left edge data and the right edge data of a lens image respectively refer to the leftmost quarter (i.e., HxW / 4; H and W respectively represent the height and width of the lens image) and the rightmost quarter of the lens image, for the sake of convenience of description, the "quarter" is simply referred to as "quarter" hereinafter. Since the processor element 120 acquires the lens image iK0 and is responsible for the overlapping area A(0), the ISP 131 of the processor element 120 needs to transmit the leftmost quarter iK0' of its own lens image iK0 to the processor element 123 through the output port 155, and the GPU 132 of the processor element 120 receives and parses the MIPI packet (containing the leftmost quarter iK1' of the neighboring right lens image iK1) from the ISP 131 of the processor element 121 through the input port 153, and stores the leftmost quarter iK1' in its own local volatile memory 170 according to the data type (such as 0x30, representing the input edge data) of the packet header, so that the leftmost quarter iK1' and the rightmost quarter rK0' of its own lens image iK0 form the overlapping area A(0). In step S404, the processor elements 121-123 operate in a similar manner to the processor element 120.

[0082] In an ideal situation, the four lenses K0~K3 are located at the center 53 of the camera system of the cubic architecture 11A, thus a single ideal imaging point 50 of an object 55 is located on an image plane 12 with a radius of 2 meters (r=2), as shown on the left side of FIG. 5A . Taking lenses K1 and K2 as an example, because the ideal imaging point 50 of the lens image iK1 coincides with the ideal imaging point 50 of the lens image iK2, the multiple equirectangular panoramic images will show a perfect stitching / mixing result after the image stitching / mixing operation is completed. However, in an actual situation, the lens centers 56 and 57 of the lens image iK1 and the lens image iK2 have an offset ofs relative to the system center 53, and as a result, the multiple equirectangular panoramic images will clearly show a mismatched image defect after the image stitching / mixing operation is completed.

[0083] FIG. 3D An example of an equirectangular panoramic image with four overlapping regions A(1)~A(4) and twenty control regions R(1)~R(20) is shown. Please refer to FIG. 3D , each overlapping region A(1)~A(4) includes P1 control regions arranged in a column, where P1>=3. The following examples and embodiments are described by taking an example of each overlapping region of an equirectangular panoramic image including five (P1=5) control regions. In the example of FIG. 3D , the multiple equirectangular panoramic images have twenty control regions R(1)~R(20), and the multiple control regions R(1)~R(20) respectively have twenty warping coefficients C(1)~C(20), which respectively represent different warping degrees of the twenty control regions R(1)~R(20).

[0084] In the measurement mode, each GPU modifies the texture coordinates of each vertex in the aforementioned original vertex sub-lists or0~or3 in each lens image according to the blending weights of the "test" warping coefficients of the two control regions closest to a target vertex and a corresponding warping coefficient of the target vertex, to generate the area error of each control region (steps S705 & 706); and in the display mode, each GPU modifies the texture coordinates of each vertex in the aforementioned original vertex sub-lists or0~or3 in each lens image according to the blending weights of the "best" warping coefficients of the two control regions closest to a target vertex and a corresponding warping coefficient of the target vertex, to minimize the mismatched image defect (step S409). FIG. 5B An example of the positional relationship between the target vertex P and the ten control regions R(1)~R(10) in the lens image iK1 is shown. In the example of FIG. 5BIn this example, the angle θ is clockwise and forms between a first vector V1 and a second vector V2; the first vector V1 is centered at image center 51 (with texture coordinates (u center ,v center Starting from the image center 51 and ending at the target vertex P(u)... P ,v P The endpoint is θ = 119.5°. Since there are five control areas on the left and right sides of the lens image iK1, 90° / 4 = 22.5°, idx = θ / 22.5° = 5 and θmod 22.5° = θ - idx × 22.5° = 7°. During the offline phase, it can be determined which two control areas (such as R(4) and R(5)) are closest to the target vertex P, and their index values ​​(4 and 5) are written / stored in the "Joining Coefficient Index" field of the lens image iK1 in the data structure of vertex P in the original vertex sublist or1 (as shown in Table 1); in addition, during the offline phase, the mixed weight (= 7 / 22.5) of the joining coefficients (C(4) and C(5)) is also calculated and stored in the "Joining Coefficient Mixing Weight (Alpha)" field of the lens image iK1 in the data structure of vertex P in the original vertex sublist or1. Note that a set of twenty test joining coefficients (C) in the measurement mode t (1)~C t (20)) and a set of twenty optimal binding coefficients (C(1) to C(20)) in the imaging mode are respectively arranged as a one-dimensional (1D) binding coefficient array or a one-dimensional data stream. Furthermore, in the measurement mode (step S702), according to FIG. 5A The offset ofs is used to specify the set of twenty test bonding coefficients (C). t (1)~C t The value of (20) is determined at the end of the measurement mode (steps S406 and S772). The values ​​of the group of twenty optimal bonding coefficients (C(1) to C(20)) are used in the imaging mode (step S409).

[0085] One of the features of this invention is that, in measurement mode, within a preset number of loops ( FIG. 7A The optimal engagement factor for twenty control zones is determined within the range of (max). This preset number of loops relates to an offset ofs, which is the distance (refer to reference) between the lens center 56 of the aforementioned camera 110A and its camera system center 53. FIG. 5A In measurement mode, according to FIG. 5A The offset ofs will be used to determine the twenty test bonding coefficients C. t (1)~C t(20) Set to different numerical ranges to measure the error amounts E(1) to E(20) of the multiple regions, and set the twenty test bonding coefficients to the same value in each (or each loop) cycle. For example, assuming ofs = 3 cm, the twenty test bonding coefficients C t (1)~C t (20) is set to a value range of 0.96 to 1.04, and if each increment is 0.01, a total of nine measurements will be taken. FIG. 7A (max = 9); Assuming ofs = 1 cm, the twenty test bonding coefficients C t (1)~C t (20) Set to a value range of 0.99 to 1.00, with each increment being 0.001, a total of ten measurements will be taken. FIG. 7A (max = 10). Note that the offset ofs is detected or determined during the offline phase, hence the twenty test bonding coefficients C. t (1)~C t The value of (20) is also predetermined and stored in the local non-volatile memory 16j, where 0 <= j <= 3.

[0086] In step S406, in measurement mode, execute FIG. 7A The method for determining the optimal bonding coefficient of the control area is described below for clarity and convenience. FIG. 7A The method for determining the optimal bonding coefficients C(6) to C(10) of the control areas R(6) to R(10) and FIG. 7B The method for coefficient decision-making is illustrated using GPU 132 of auxiliary processor element 121 as an example, assuming ofs = 3 cm. It should be understood that: FIG. 7A Methods for determining the optimal bonding coefficient of the control area and FIG. 7B The method of coefficient decision operation is also applicable to GPU 132 of processor elements 120 and 122-123 to generate optimal bonding coefficients C(1)-C(5) and C(11)-C(20), respectively.

[0087] Step S702: Set the loop number Q1 and the test engagement coefficient to new values. In one embodiment, Q1 is set to 1 in the first loop, and Q1 is increased by 1 in each subsequent loop; if ofs = 3 cm, the plurality of test engagement coefficients C are set to new values ​​in the first loop. t (1)~C t (20) All are set to 0.96 (i.e., C) t (1) = ... = C t (20) = 0.96), and in subsequent loops, the plurality of bonding coefficients C are sequentially set. t (1)~C t(20) is set to 0.97,…,1.04.

[0088] Step S704: Clear all regional error quantities E(i) to 0, where i = 6, ..., 10.

[0089] Step S705: Based on the test bonding coefficient C t (1)~C t The value of (10) and the original auxiliary vertex sublist or1 are used to generate a modified vertex sublist m1. Below, again using... FIG. 5B For example, after receiving the original auxiliary vertex sublist or1 from the local non-volatile memory 161, the GPU 132 of the auxiliary processor element 121 extracts two test joining coefficients (C) from the one-dimensional test joining coefficient array based on the "joining coefficient index" field (i.e., 4 and 5) in the lens image iK1 of the target vertex P in the data structure of the target vertex P. t (4) and C t (5) Then, based on the "blending weight (Alpha)" column (i.e., 7 / 22.5) of the target vertex P in the image iK1 (see Table 1), the interpolation blending coefficient C' is calculated according to the following equation: C' = C t (4)×(7 / 22.5)+C t (5)×(1-7 / 22.5). Then, the GPU 132 of processor element 121 calculates the corrected texture coordinates (u) of the target vertex P in the lens image iK1 according to the following equation. P ',v P '):u P '=(u P -u center )×C'+u center ;v P '=(v P -v center )×C'+v center In this manner, the GPU 132 of the processor element 121 is configured according to the ten test bonding coefficients C. t (1)~C t (10) The texture coordinates of the lens image iK1 of each vertex from the original auxiliary vertex sublist or1 are sequentially corrected to generate a corrected auxiliary vertex sublist m1. Similarly, the GPU 132 of processor elements 120 and 122-123 also adjusts the texture coordinates of the lens image iK1 according to the twenty test bonding coefficients C. t (1)~C t(20), sequentially correct all the texture coordinates of the lens images iK0and iK2~iK3of each vertex from the three original helper vertex sublists or0and or2~or3to generate three corrected helper vertex sublists m0and m2~m3. Table II shows an example of the data structure of each vertex in the corrected helper vertex sublists.

[0090]

[0091]

[0092] Step S706: Measure the area error amounts E(6)~E(10) of the five control regions R(6)~R(10) of the rectangular panoramic image according to the corrected helper vertex sublist ml, the lens image iK1and the input leftmost quad iK2' by the GPU 132 of the processor element 121 (to be described in detail). FIG. 6 For convenience of description, use E(i) = f(C t (i)) to represent this step S706, where i = 6,..., 10, and f() represents the measurement of the area error amounts E(6)~E(10) according to the corrected helper vertex sublist ml, the lens image iK1and the input leftmost quad iK2' by the GPU 132 of the processor element 121.

[0093] Step S708: Store all the area error amounts E(6)~E(10) and the values of all the test seam coefficients in a two-dimensional (2D) error table. Table III shows an example of the 2D error table when ofs= 3 cm (the value range of the test seam coefficient is 0.96~1.04). In Table III, there are five area error amounts E(6)~E(10) and nine values of the test seam coefficient.

[0094]

[0095] Step S710: Determine whether the loop number Q1 has reached the upper limit max (= 9). If yes, go to step S712, otherwise, return to step S702.

[0096] Step S712: Perform the coefficient decision operation according to the above-mentioned 2D error table.

[0097] Step S714: Output the optimal seam coefficient C(i), where i = 6,..., 10.

[0098] FIG. 7B According to an embodiment of the present application, a flow chart of the method of performing the coefficient decision operation of step S712 is shown.

[0099] Step S761: Set Q2 to 0 for initialization.

[0100] Step S762: From the 2D error table, a selected decision group is extracted. Returning to FIG. 3D , generally each control region is adjacent to two control regions, a selected control region and its adjacent two control regions form a selected decision group to determine the optimal joining coefficient of the selected control region. For example, a selected control region R(9) and its adjacent two control regions R(8) and R(10) form a selected decision group. However, if a selected control region (e.g. R(6)) is located at the top or bottom of the overlapping region A(l), the selected control region R(6) will only form a selected decision group with its only adjacent control region R(7) to determine its optimal joining coefficient C(6). The subsequent steps are described assuming that a control region R(7) is selected and R(7) and its adjacent two control regions R(6) and R(8) form a selected decision group to determine its optimal joining coefficient C(7).

[0101] Step S764: In the area error amounts of the control regions of the selected decision group, local minima are determined. Table IV shows an example of the area error amounts of R(6) ~ R(8) and the test joining coefficients C t (6) ~ C t (8).

[0102]

[0103] As shown in Table IV, there is only one local minimum in the nine area error amounts of R(6), and there are two local minima in the nine area error amounts of R(7) and R(8) respectively, in which each local minimum is marked with an asterisk (*) in Table IV.

[0104] Step S766: From the local minima, candidates are selected. Table V shows the candidates selected from the local minima of Table IV, in which ID represents index, WC represents joining coefficient, and RE represents area error amount. The number of candidates is equal to the number of local minima in Table IV.

[0105]

[0106] Step S768: From the candidates of Table V, a link metric is established. As shown in FIG. 7C , a link metric is established from the candidates of Table V.

[0107] Step S770: In all paths of the link metric, the minimum sum of link metric values is determined. The minimum value between two link metric values and is the minimum sum of link metric values between two link metric values and the minimum value between them After that, the sum of the link metric values of path 0-0-0 and path 0-1-1 is calculated as follows: and Because Therefore, it can be determined that (path 0-1-1) is the minimum sum of the link metric values among all paths, as shown by the solid line path in FIG. 7C

[0108] Step S772: Determine the optimal blending factor of the selected control region. In the example shown in step S770, because (path 0-1-1) is the minimum sum of the link metric values among all paths, it is determined that 1.02 is the optimal blending factor of control region R(7). However, if the sum of the link metric values of two or more paths are the same at the end of the calculation, the blending factor of the node with the smallest region error is selected as the optimal blending factor of the selected control region. At this point, the value of the loop number Q2 is incremented by 1.

[0109] Step S774: Determine whether the loop number Q2 has reached the upper limit 5. If yes, end the flowchart, otherwise, return to step S762 to process the next control region. In this way, the GPU 132 of each processor element 120-123 forms its own 2D error table, and determines five optimal blending factors for the five control regions in the overlap region respectively.

[0110] FIG. 6 According to an embodiment of the present application, a schematic diagram of a GPU is shown. Please refer to FIG. 6 The GPU 132 of each processor element includes a rasterization engine 610, a texture mapping circuit 620, a blending unit 630 (controlled by a control signal CS2), and a measurement unit 650 (controlled by a control signal CS1). Note that in the measurement mode, if the isometric rectangular coordinates of a point fall within the five control regions responsible for it, the blending unit 630 will be disabled and the measurement unit 650 will be enabled by the two control signals CS1 and CS2; in the display mode, the blending unit 630 will be enabled and the measurement unit 650 will be disabled by the two control signals CS1 and CS2. The texture mapping circuit 620 includes two texture mapping engines 621-622. FIG. 3C The polygon mesh of FIG. 3C ​pixels within the quadrilateral are subjected to a quadrilateral rasterization operation, or a triangle is formed from each set of three vertices from the modified vertex sub-list (as shown in FIG. 6) and pixels within the triangle are subjected to a triangle rasterization operation. FIG. 3C pixels within the triangle are subjected to a triangle rasterization operation.

[0111] In the case of a quadrilateral rasterization operation, assume that a set of four vertices (A, B, C, D) from the modified primary vertex sub-list m0 (forming a quadrilateral of the polygon mesh) is located within the range of one of the five control regions of overlap region A (0) and is overlapped by two camera images (iK0 and iK1; N = 2), the four vertices (A, B, C, D) respectively include the following data structures: Vertex A: {(x A ,y A ), 2, ID iK0 , (u 1A ,v 1A ), w 1A , ID iK1 , (u 2A ,v 2A ), w 2A}; Vertex B: {(x B ,y B ), 2, ID iK0 , (u 1B ,v 1B ), w 1B , ID iK1 , (u 2B ,v 2B ), w 2B}; Vertex C: {(x C ,y C ), 2, ID iK0 , (u 1C ,v 1C ), w 1C , ID iK1 , (u 2C ,v 2C ), w 2C}; Vertex D: {(x D ,y D ), 2, ID iK0 , (u 1D ,v 1D ), w 1D , ID iK1 , (u 2D ,v 2D ), w 2D} The rasterization engine 610 of the processor element 120 (responsible for A(0)) performs a quad rasterization operation directly on the points / pixels within the quad ABCD. Specifically, the rasterization engine 610 of the processor element 120 computes the texture coordinates for a point Q (having isometric coordinates (x,y) and located within the quad ABCD of the polygon grid) using the following steps: (1) using a bi-linear interpolation method, compute four spatial weights (a,b,c,d) based on the isometric coordinates (x A ,y A ,x B ,y B ,x C ,y C ,x D ,y D ,x,y); (2) compute the facet blending weight for a sample point Q iK0 (corresponding to the point Q) in the lens image iK0: fw1 = a x w 1A + b x w 1B + c x w 1C + d x w 1D ; compute the facet blending weight for a sample point Q iK1 (corresponding to the point Q) in the lens image iK1: fw2 = a x w 2A + b x w 2B + c x w 2C + d x w 2D ; (3) compute the texture coordinates for the sample point Q iK0 (corresponding to the point Q) in the lens image iK0: (u1,v1) = (a x u 1A + b x u 1B + c x u 1C + d x u 1D , a x v 1A + b x v 1B + c x v 1C + d x v 1D ; compute the texture coordinates for the sample point Q iK1 (corresponding to the point Q) in the lens image iK1: (u2,v2) = (a x u 2A + b x u 2B + c x u 2C + d x u 2D , a x v 2A + b x v 2B + c x v 2C + d x v 2D). Finally, the rasterization engine 610 of the processor element 120 parallel transfers the two texture coordinates (u1, v1) and (u2, v2) to the two texture mapping engines 621-622. Where a + b + c + d = 1 and fw1 + fw2 = 1. Based on the two texture coordinates (u1, v1) and (u2, v2), the two texture mapping engines 621-622 texture map the texture data of the two lens images iK0and iK1using any suitable method (e.g., nearest-neighbour interpolation, bilinear interpolation, or trilinear interpolation) to generate two sample values s1, s2. Each of the sample values can be a luma value, a chroma value, an edge value, a pixel color value (RGB), or a motion vector.

[0112] In the case of triangle rasterization operation, the rasterization engine 610 and the two texture mapping engines 621-622 of the processor element 120 perform similar operations (similar to the case of quad rasterization operation described above) on each pixel inside a triangle formed by any three vertices from the modified primary vertex sub-list m0 FIG. 3C ) to generate two corresponding sample values s1, s2, except that the rasterization engine 610 uses a barycentric weighting method in step (1) instead of the bilinear interpolation method described above, based on the barycentric coordinates (x A ,y A ,x B ,y B ,x C ,y C ,x,y) of the pixel, to calculate three spatial weights (a, b, c) of the three vertices (A, B, C).

[0113] Next, the raster engine 610 of the processor element 120 determines whether the point Q falls into one of the five responsible control regions R(l)-R(5) according to the equidistant rectangular coordinates (x, y) of the point Q. If the point Q is determined to fall into one of the five responsible control regions, the control signal CS1 is asserted to cause the measurement unit 650 to start measuring the area error of the control region. The measurement unit 650 of the processor element 120 can use any known algorithm, such as sum of absolute differences (SAD), sum of squared differences (SSD), median absolute deviation (MAD), etc., to estimate / measure the area error of the control regions. For example, if the point Q is determined to fall into control region R(l), the measurement unit 650 uses the following equation: E = \s1-s2\; E(l)+=E, to accumulate the absolute value of the difference between the sample values of each point in control region R(l) in the lens image iK0 and the corresponding point in control region R(l) in the lens image iKl to obtain a SAD value as the area error E(l) of control region R(l). In this way, the measurement unit 650 measures the area errors E(l)-E(5) of the five control regions R(l)-R(5). In the same way, the measurement unit 650 of the processor element 121 measures the area errors E(6)-E(10) of the five control regions R(6)-R(10) according to the modified auxiliary vertex sub-list ml, the lens image iKl, and the leftmost quarter iK2' of the adjacent right lens image iK2; the measurement unit 650 of the processor element 122 measures the area errors E(l l)-E(15) of the five control regions R(l l)-R(15) according to the modified auxiliary vertex sub-list m2, the lens image iK2, and the leftmost quarter iK3' of the adjacent right lens image iK3; and the measurement unit 650 of the processor element 123 measures the area errors E(16)-E(20) of the five control regions R(16)-R(20) according to the modified auxiliary vertex sub-list m3, the lens image iK3, and the leftmost quarter iK0' of the adjacent right lens image iK0 (step S706).

[0114] In step S408 (i.e., transmission stage two), the GPU 132 of each processor element transmits the optimal bonding coefficients of the five control regions in the overlapping area it is responsible for to the GPU 132 of a neighboring processor element through the output port 152, and receives the optimal bonding coefficients of the five control regions in the adjacent left overlapping area from another neighboring processor element through the input port 154. For example, GPU 132 of processor element 122 transmits the optimal binding coefficients C(11) to C(15) of the five control regions R(11) to R(15) in the overlapping region A(2) it is responsible for to GPU 132 of processor element 123 through output port 152, and receives and parses MIPI packets (containing the optimal binding coefficients C(6) to C(10) of the five control regions R(6) to R(10)) from GPU 132 of processor element 121 through input port 154, and stores the input optimal binding coefficients C(6) to C(10) in its own local volatile memory 172 according to the data type of the packet header (such as 0x31, representing the input optimal binding coefficient). The operation of GPU 132 of other processor elements 120 to 121 and 123 is similar to that of GPU 132 of processor element 122.

[0115] In step S409, similar to step S705, based on the aforementioned twenty optimal bonding coefficients C(1) to C(20), the GPU 132 of processor elements 120 to 123 respectively corrects the texture coordinates of each vertex in the aforementioned original vertex sublists or0 to or3 in the four-lens image iK0 to iK3, generating a corrected primary vertex sublist m0' and three corrected auxiliary vertex sublists m1' to m3'. Hereinafter, again... FIG. 5B For example, after receiving five optimal bonding coefficients C(1) to C(5) from processor element 120, GPU 132 of auxiliary processor element 121 extracts two test bonding coefficients (C(4) and C(5)) from the one-dimensional test bonding coefficient array based on the "bonding coefficient index" field (i.e., 4 and 5) in the data structure of the target vertex P in the lens image iK1. Then, based on the "bonding coefficient mixing weight (Alpha)" field (i.e., 7 / 22.5) in the data structure of the target vertex P in the lens image iK1, it calculates the interpolation bonding coefficient C' according to the following equation: C' = C(4) × (7 / 22.5) + C(5) × (1 - 7 / 22.5). Afterwards, GPU 132 of processor element 121 calculates the corrected texture coordinates (u) of the target vertex P in the lens image iK1 according to the following equation. P ',v P '):u P '=(u P -u center )×C'+u center ;vP ' = (v P - v center ) x C' + v center In this way, the GPU 132 of the processor element 121 sequentially corrects the texture coordinates of the lens image iK1 of each vertex from the original auxiliary vertex sub-list or1 according to the ten best blending coefficients C(1)~C(10) to generate a corrected auxiliary vertex sub-list m1'. After the texture coordinates of each vertex in the four original vertex sub-lists or0~or3 in the four lens images iK0~iK3 are corrected according to the twenty best blending coefficients C(1)~C(20), the mismatch image defect problem caused by the lens center offset (i.e., a lens center 56 is offset from its system center 53 by an offset distance ofs) of the camera 110A can be greatly improved (i.e., the actual imaging position 58 is pushed toward the ideal imaging position 50), as shown in FIG. 5A Please note that since the sphere 12 is virtual, the object 55 can be located outside, inside, or on the surface of the sphere 12.

[0116] In step S410, the rasterization engine 610, the texture mapping circuit 620, and the blending unit 630 of each processor element operate together to generate a working face image according to the self lens image, the leftmost quarter of the neighboring right lens image, and the corrected vertex sub-list of the self. For example, the rasterization engine 610, the texture mapping circuit 620, and the blending unit 630 of the processor element 123 operate together to generate a working face image F3 according to the self lens image iK3, the leftmost quarter iK0' of the neighboring right lens image iK0, and the corrected vertex sub-list m3' of the self. The term "working face image" refers to an image generated from the projection of a corresponding lens image from the camera 110; the projection is, for example, an equirectangular projection, a cylindrical projection, a Miller projection, a Mercator projection, a Lambert cylindrical equal-area projection, or a Pannini projection. In the present disclosure, each working face image includes a non-overlapping region and an overlapping region. For example, as shown in FIG. 3D Since the processor element 123 is responsible for the overlapping region A(3), the processor element 123 generates a working face image F3 that includes a non-overlapping region b3 and an overlapping region A(3).

[0117] Returning to FIG. 6, the grid engine 610 and the texture mapping circuit 620 operate in the same manner as in the measurement mode. The following again uses the above example (a point Q has an equidistant rectangular coordinate (x, y) and is located in the quadrangle ABCD of the polygon grid, and the quadrangle ABCD is overlapped by two lens images (iK0and iK1; N = 2). After the two texture mapping engines 621-622 of the processor element 120 texture map the texture data of the lens images iK0and iK1to generate two sample values s1, s2, the blending unit 630 of the processor element 120 blends the two sample values s1, s2in the following equation to generate a blended value Vbfor the point Q: Vb= fw1x s1+ fw2x s2. Finally, the blending unit 630 of the processor element 120 stores the blended value Vbfor the point Q in the local volatile memory 170. In this way, the blending unit 630 of the processor element 120 stores all the blended values Vbin the local volatile memory 170 until all the points in the quadrangle ABCD are processed, and once all the quadrangles and triangles are processed, a working surface image F0is stored in the local volatile memory 170. Similarly, the GPU 132 of the processor element 121 generates a working surface image F1from the lens image iK1, the leftmost quarter iK2' of the adjacent right lens image iK2, and the modified vertex sub-list m1'; the GPU 132 of the processor element 122 generates a working surface image F2from the lens image iK2, the leftmost quarter iK3' of the adjacent right lens image iK3, and the modified vertex sub-list m2'; and the GPU 132 of the processor element 123 generates a working surface image F3from the lens image iK3, the leftmost quarter iK0' of the adjacent right lens image iK0, and the modified vertex sub-list m3'.

[0118] In step S412 (i.e., the transmission phase three), the GPU 132 of each of the processor elements 120-123 cuts the working surface image thereof into a plurality of tiles having a predetermined size, calculates histograms H1and Hr of the leftmost column of tiles and the rightmost column of tiles of the lens image thereof, and transmits a predetermined section of the working surface image and the histograms H1and Hr to two adjacent processor elements. In an embodiment, the predetermined size of the tiles is equal to 64x64, and the predetermined section of the working surface image is the leftmost eight columns of pixels and the rightmost eight columns of pixels of the working surface image. It should be noted that the predetermined size of the tiles and the predetermined section of the working surface image are merely examples and are not limiting of the present application, and other sizes of tiles and other numbers of columns of pixels can be used in actual implementation. In FIG. 4AIn the embodiment, the GPU 132 of the processor element 123 transmits the histogram Hr3 of the rightmost column of tiles of the working surface image F3 and the pixels Fr3 of the rightmost eight columns of tiles to the IQE unit 133 of the processor element 120 through the output port 152, and transmits the histogram Hl of the leftmost column of tiles of the working surface image F3 and the pixels Fl3 of the leftmost eight columns of tiles to the IQE unit 133 of the processor element 122 through the output port 155. The IQE unit 133 of the processor element 123 receives and parses the MIPI packets (including the histograms Hl0 and Hr2 and the segments Fl0 and Fr2) from the processor elements 120 and 122 through the input ports 153 and 154, and stores the histograms Hl0 and Hr2 and the segments Fl0 and Fr2 in the local volatile memory 173 thereof according to the data type (e.g. 0x32 for input histogram; 0x33 for input segment) of the packet header. The GPUs 132 and the IQE units 133 of the other processor elements 120-122 operate in a similar manner as the GPU 132 and the IQE unit 133 of the processor element 123.

[0119] In step S414, after receiving the tile histograms and segments of the two adjacent working surface images from the two adjacent processor elements, the IQE unit 133 of each processor element performs image quality enhancement processing on the working surface image thereof. The image quality enhancement processing includes, but is not limited to, contrast enhancement, low-pass filtering and image sharpening processing. The contrast enhancement can be implemented using any known algorithm, such as contrast limited adaptive histogram equalization (CLAHE). For example, the IQE unit 133 of the processor element 123 performs image quality enhancement processing on the working surface image F3 according to the histograms Hl0 and Hr2 and the segments Fl0 and Fr2 to generate an enhanced image F3'. The IQE units 133 of the other processor elements 120-122 operate in a similar manner as the IQE unit 133 of the processor element 123.

[0120] After step S414 is completed, FIG. 4C the flow can directly proceed to step S416 (hereinafter referred to as "Method I", without the link 481 and step S415); in step S416, the four encoding and transmission units 134 respectively encode the four enhanced images F0'-F3' into four encoded video streams en0-en3, and transmit the four encoded video streams en0-en3 to the receiver 180, so as to generate a panoramic image. Alternatively, FIG. 4CThe flow of the method can re-enter step S416 (hereinafter referred to as "Method Two"; link 481) through step S415, and the operation is as follows: in step S415 (transmission stage four), the IQE units 133 of the three auxiliary processor elements 121-123 respectively transmit the three enhanced images F1'-F3' to the encoding and transmission unit 134 of the processor element 120 through the output ports 155, 156 and 152. The encoding and transmission unit 134 of the processor element 120 receives and parses the MIPI package (including the three enhanced images F1'-F3') through the input ports 153, 157 and 154, and stores the three enhanced images F1'-F3' in the local volatile memory 170 according to the data type of the package header (such as 0x34, representing input enhanced image). In step S416, the encoding and transmission unit 134 of the processor element 120 merges the three enhanced images F1'-F3' into the enhanced image F0' to form a single bit stream, encodes the single bit stream into a single encoded video stream en, and transmits the single encoded video stream en to the receiver 180. Please note that in Method Two, only the encoding and transmission unit 134 of the processor element 120 is required, and the encoding and transmission units 134 of the other processor elements 121-123 can be discarded.

[0121] Please note that, as mentioned above, in each processor element and Method Two, the IQE unit 133 is not required, so steps S412, S414 and S415 are also not required, and thus are shown in dashed lines in FIG. 4B-FIG. 4C . Assuming that the circuit discards all IQE units 133, the GPU 132 of each processor element transmits the respective working surface image F0-F3 to the respective encoding and transmission unit 134 for subsequent encoding and transmission operation after generating the working surface image F0-F3 (Method One; steps S412, S414 and S415 are discarded); or the GPUs 132 of the three auxiliary processor elements 121-123 respectively transmit the three working surface images F1-F3 to the encoding and transmission unit 134 of the processor element 120 for subsequent encoding and transmission operation (Method Two; steps S415 and steps S412 and S414 are discarded).

[0122] FIG. 8 According to another embodiment of the present application, a block architecture diagram of a dual-processor system suitable for a four-lens camera is shown. Referring to FIG. 8 , the dual-processor system 800 is used to process four-lens image data from a four-lens camera 110A, and includes a main processor element 120, an auxiliary processor element 121, and four links. Please also refer to FIG. 1The processor element 120 is connected to the two lenses K0-K1 of the camera 110A through the input port 151, and the processor element 121 is connected to the two lenses K2-K3 of the camera 110A. For the sake of clarity and convenience of description, FIG. 8 Only the two processor elements 120-121, the I / O ports included therein, and the four links are shown, and the operation mode will be described later in detail. In the present embodiment, each processor element 120 / 121 includes three I / O ports 151-153. In the offline stage, because the dual-processor system 800 includes two processor elements 120-121, the original vertex list (Table 1) is divided into two original vertex sublists, i.e., an original primary vertex sublist or01 (used by the processor element 120) and an original secondary vertex sublist or23 (used by the processor element 121), according to the equidistant rectangular coordinates, and the two original vertex sublists or01-or23 are respectively stored in two local non-volatile memories 160-161 for subsequent image processing.

[0123] The operation mode of the dual-processor system 800 will be described below according to the flow of FIG. 4B-FIG. 4C In step S402, the ISP 131 of the processor element 120 receives and parses MIPI packets containing electrical signals related to the image sensors of the lenses K0 and K1 of the camera 110A through the MIPI input port 151, converts the electrical signals into two lens images iK0 and iK1, and stores the two lens images iK0 and iK1 in its own local volatile memory 170 according to the data type (such as 0x2A) of the packet header. The operation mode of the ISP 131 of the processor element 121 is similar to that of the ISP 131 of the processor element 120. Please note that each processor element is responsible for two overlapping areas, as shown in FIG. 3D In one embodiment, the processor element 120 obtains the lens images iK0-iK1 and is responsible for the two overlapping areas A(3) and A(0); the processor element 121 obtains the lens images iK2-iK3 and is responsible for the two overlapping areas A(1) and A(2). For the sake of clarity and convenience of description, the following examples and embodiments assume that the processor element 120 obtains the lens images iK0-iK1 and is responsible for the two overlapping areas A(0) and A(1), and the processor element 121 obtains the lens images iK2-iK3 and is responsible for the two overlapping areas A(2) and A(3).

[0124] At step S404 (transmission phase one), for forming the four overlapping areas, each processor element needs to transmit its left edge data of the two lens images to another processor element through the output port 152, and to receive the left edge data of the adjacent two lens images from another processor element through the input port 153. For each processor element, the left edge data of its two lens images to be transmitted are located at the opposite edge of the two overlapping areas it is responsible for, while the right edge data of its two lens images and the left edge data of the adjacent two right lens images it receives form a corresponding overlapping area, and the right edge data of its two lens images and the left edge data of the adjacent two right lens images are related to the size of the corresponding overlapping area; for example, the edge data rK1’ and K2’ form the overlapping area A(l) and are related to the size of A(l). As mentioned above, once the lens FOA, the lens sensor resolution and the lens mounting angle of the camera 110A are fixed, the sizes of the overlapping areas A(0)-A(3) are determined. Assuming that the left edge data and the right edge data of the two lens images refer to the leftmost quarter (i.e. HxW / 4) and the rightmost quarter of the left lens image and the right lens image of the two lens images respectively, because the processor element 120 obtains the lens images iK0-iK1 and is responsible for the overlapping areas A(0)-A(l), the ISP 131 of the processor element 120 needs to transmit the leftmost quarter iK0’ of the lens image iK0 to the processor element 121 through the output port 152, and the GPU 132 of the processor element 120 receives and parses the MIPI packet (including the leftmost quarter iK2’ of the adjacent right lens image iK2) from the ISP 131 of the processor element 121 through the input port 153, and stores the leftmost quarter iK2’ in its own local volatile memory 170 according to the data type (such as 0x30, representing the input edge data) of the packet header, so that the leftmost quarter iK2’ and the rightmost quarter rK1’ of the two lens images iK0-iK1 form the overlapping area A(l). At step S404, the ISP 131 and the GPU 132 of the processor element 121 operate in a similar manner to the ISP 131 and the GPU 132 of the processor element 120.

[0125] At step S406, the GPU 132 of the processor element 120 forms a 2D error table (such as Table 3) including different values of the twenty test joint coefficients (related to FIG. 7A and FIG. 7B . At step S408, the GPU 132 of the processor element 120 determines the best joint coefficient value (e.g. the value of the joint coefficient J) from the 2D error table, and transmits the best joint coefficient value to the ISP 131 of the processor element 120 through the local bus 160. At step S410, the ISP 131 of the processor element 120 performs the image stitching according to the best joint coefficient value. FIG. 5Aa 2D error table (e.g. Table 3) including the twenty test joint coefficients and the area error amounts E(11)~E(20) of the ten control areas R(11)~R(20) within the overlap area A(2)~A(3) responsible for the GPU 132 of the processor element 121, the 2D error table used to determine the optimal joint coefficients C(11)~C(20) of the ten control areas R(11)~R(20).

[0126] At step S408 (i.e. transmission phase two), the GPU 132 of the processor element 120 transmits the optimal joint coefficients C(6)~C(10) of the five control areas R(6)~R(10) within the overlap area A(1) responsible for the GPU 132 of the processor element 121 through the output port 152, and receives and parses the MIPI packet (including the optimal joint coefficients C(16)~C(20) of the five control areas R(16)~R(20)) from the GPU 132 of the processor element 121 through the input port 153, and stores the optimal joint coefficients C(16)~C(20) in the local volatile memory 170 according to the data type (e.g. 0x31, representing the input optimal joint coefficients) of the packet header. The GPU 132 of the processor element 121 operates in a similar manner as the GPU 132 of the processor element 120.

[0127] In step S409, according to the best joint coefficients C(l)-C(10) and C(16)-C(20), the GPU 132 of the primary processor element 120 corrects the texture coordinates of each vertex in the aforementioned original primary vertex sub-list or01 in the two lens images iK0-iK1 to generate a corrected primary vertex sub-list m01'; according to the best joint coefficients C(6)-C(20), the GPU 132 of the auxiliary processor element 121 corrects the texture coordinates of each vertex in the aforementioned original auxiliary vertex sub-list or23 in the two lens images iK2-iK3 to generate a corrected auxiliary vertex sub-list m23'. In step S410, the rasterization engine 610, the texture mapping circuit 620 and the blending unit 630 of the primary processor element 120 operate together to generate two working surface images F0-F1 according to the two lens images iK0-iK1 of itself, the leftmost quad iK2' and the corrected primary vertex sub-list m01' of itself; the rasterization engine 610, the texture mapping circuit 620 and the blending unit 630 of the auxiliary processor element 121 operate together to generate two working surface images F2-F3 according to the two lens images iK2-iK3 of itself, the leftmost quad iK0' and the corrected auxiliary vertex sub-list m23' of itself.

[0128] In step S412 (i.e., transmission phase three), the GPU 132 of each processor element 120-121 cuts the working surface images into a plurality of tiles (e.g., 64x64) of a predetermined size, and calculates the histogram Hl of the leftmost column of tiles and the histogram Hr of the rightmost column of tiles of the two lens images of itself. The GPU 132 of the processor element 120 transmits the histogram Hr1 of the rightmost column of tiles of the working surface image F1, a predetermined section (e.g., the rightmost eight columns of pixels) Fr1, and the histogram Hl0 of the leftmost column of tiles of the working surface image F0, a predetermined section (the leftmost eight columns of pixels) Fl0 to the IQE unit 133 of the processor element 121 through the output port 152; the IQE unit 133 of the processor element 120 receives and parses the MIPI packet (containing the histograms Hl2 and Hr3 and the sections Fl2 and Fr3) from the processor element 121 through the input port 153, and stores the histograms Hl2 and Hr3 and the sections Fl2 and Fr3 in the local volatile memory 170 of itself according to the data type (e.g., 0x32, representing input histogram; 0x33, representing input section) of the packet header. The GPU 132 and the IQE unit 133 of the other processor element 121 operate in a similar manner to the GPU 132 and the IQE unit 133 of the processor element 120.

[0129] At step S414, the IQE unit 133 of the processor element 120 performs image quality enhancement on the two working surface images F0 and F1 according to the received histograms Hl2 and Hr3 and the segments Fl2 and Fr3 to generate two enhanced images F0' and F1'. The IQE unit 133 of the processor element 121 performs image quality enhancement on the two working surface images F2 and F3 according to the received histograms HlO and Hr1 and the segments FlO and Fr1 to generate two enhanced images F2' and F3'.

[0130] If method one is performed, at step S416, the two encoding and transmitting units 134 encode the four enhanced images F0' - F3' into two encoded video streams en01 - en23, respectively, and transmit the two encoded video streams en01 - en23 to the receiver 180 to generate a panoramic image. If method two is performed, at step S415 (transmission stage four), the IQE unit 133 of the auxiliary processor element 121 transmits the two enhanced images F2' - F3' to the encoding and transmitting unit 134 of the processor element 120 through the output port 152. Then, the encoding and transmitting unit 134 of the processor element 120 receives and parses the MIPI packet (containing the two enhanced images F2' - F3') through the input port 153, and stores the two enhanced images F2' - F3' in its local volatile memory 170 according to the data type of the packet header (e.g. 0x34, representing input enhanced images). Then, at step S416, the encoding and transmitting unit 134 of the processor element 120 merges the two enhanced images F2' - F3' with the enhanced images F0' - F1' to form a single bitstream, encodes the single bitstream into a single encoded video stream en, and transmits the single encoded video stream en to the receiver 180.

[0131] FIG. 9A According to another embodiment of the present application, a block architecture diagram of a three-processor system suitable for a three-lens camera is shown. Note that the three-processor system 900 is used to generate three working surface images to form a wide-angle image as shown in FIG. 9B , while the four-processor system 400 and the two-processor system 800 are used to generate four working surface images to form a panoramic image as shown in FIG. 3D . Referring to FIG. 9A, three-processor system 900 for processing image data from three lenses K0-K2 of three-lens camera 110B, includes one main processor element 120, two auxiliary processor elements 121-122, and five links, where link 901 is optional. In this embodiment, processor elements 120 / 122 include four I / O ports, and processor element 121 includes five I / O ports. In the offline phase, because three-processor system 900 includes three processor elements 120-122, the original vertex list (e.g., Table 1) is divided into three original vertex sublists, namely one original main vertex sublist or0 (for use by processor element 120) and two original auxiliary vertex sublists or1-or2 (for use by processor elements 121 and 122, respectively), according to the equidistant rectangular coordinates, and the three original vertex sublists or0-or2 are stored in three local non-volatile memories 160-162, respectively, for subsequent image processing.

[0132] The following describes the operation of three-processor system 900 according to the flow of FIG. 4B-FIG. 4C Step S402, three ISPs 131 of three-processor system 900 obtain three lens images iK0-iK2 in a manner similar to ISPs 131 of four-processor system 400. In one embodiment, processor element 121 is responsible for overlap region A(0), processor element 122 is responsible for overlap region A(l), but processor element 120 is not responsible for any overlap region. For clarity and ease of description, the following examples and embodiments assume that processor element 120 is responsible for overlap region A(0), processor element 121 is responsible for overlap region A(l), but processor element 122 is not responsible for any overlap region.

[0133] At step S404 (transmission phase one), for forming two overlapping areas, the ISP 131 of the processor element 121 transmits the left edge data (e.g. the leftmost quarter iK1') of its own lens image iK1 through the output port 155 to the processor element 120; and the GPU 132 of the processor element 121 receives and parses the MIPI packet (containing the leftmost quarter iK2' of the adjacent lens image iK2) from the processor element 122 through the input port 153, and stores the leftmost quarter iK2' in its own local volatile memory 171 according to the data type (e.g. 0x30, representing the input edge data) of the packet header, so that the leftmost quarter iK2' and the rightmost quarter rK1' of its own lens image iK1 form an overlapping area A(l). Because the processor element 120 acquires the lens image iK0 and is responsible for the overlapping area A(0), the GPU 132 of the processor element 120 receives and parses the MIPI packet (containing the leftmost quarter iK1' of the adjacent lens image iK1) from the processor element 121 through the input port 153, and stores the leftmost quarter iK1' in its own local volatile memory 170 according to the data type (e.g. 0x30) of the packet header, so that the leftmost quarter iK1' and the rightmost quarter rK0' of its own lens image iK0 form the overlapping area A(0). Because the processor element 122 acquires the lens image iK2 and is not responsible for any overlapping area, the ISP 131 of the processor element 122 transmits the leftmost quarter iK2' of its own lens image iK2 through the output port 155 to the GPU 132 of the processor element 121.

[0134] At step S406, according to the method of FIG. 7A and FIG. 7B , the GPU 132 of the processor element 120 forms a 2D error table (e.g. Table III) containing different values of ten test blending coefficients (related to the offset ofs of FIG. 5A ) and the area error amounts E(l)-E(5) of five control areas R(l)-R(5) (located in the overlapping area A(0) which it is responsible for), which is used to determine the optimal blending coefficients C(l)-C(5) of the five control areas R(l)-R(5); and the GPU 132 of the processor element 121 forms a 2D error table (e.g. Table III) containing different values of the aforementioned ten test blending coefficients and the area error amounts E(6)-E(10) of five control areas R(6)-R(10) (located in the overlapping area A(l) which it is responsible for), which is used to determine the optimal blending coefficients C(6)-C(10) of the five control areas R(6)-R(10).

[0135] At step S408 (i.e., the second transmission stage), the GPU 132 of the processor element 120 transmits the optimal seam coefficients C(1)-C(5) of the five control regions R(1)-R(5) in the overlap region A(0) responsible for the processor element 120 to the processor element 121 through the output port 152; the GPU 132 of the processor element 121 transmits the optimal seam coefficients C(6)-C(10) of the five control regions R(6)-R(10) in the overlap region A(1) responsible for the processor element 121 to the processor element 122 through the output port 152, and receives and parses the MIPI packet (containing the five optimal seam coefficients C(1)-C(5)) from the GPU 132 of the processor element 120 through the input port 154, and stores the optimal seam coefficients C(1)-C(5) in the local volatile memory 171 thereof according to the data type (e.g., 0x31) of the packet header. The GPU 132 of the processor element 122 receives and parses the MIPI packet (containing the five optimal seam coefficients C(6)-C(10)) from the GPU 132 of the processor element 121 through the input port 154, and stores the optimal seam coefficients C(6)-C(10) in the local volatile memory 172 thereof according to the data type of the packet header.

[0136] At step S409, the GPU 132 of the primary processor element 120 corrects the texture coordinates of each vertex in the aforementioned original primary vertex sub-list or0 in the lens image iK0 according to the optimal seam coefficients C(1)-C(5), to generate a corrected primary vertex sub-list m0'; the GPU 132 of the secondary processor element 121 corrects the texture coordinates of each vertex in the aforementioned original secondary vertex sub-list or1 in the lens image iK1 according to the optimal seam coefficients C(1)-C(10), to generate a corrected secondary vertex sub-list m1'; and the GPU 132 of the secondary processor element 122 corrects the texture coordinates of each vertex in the aforementioned original secondary vertex sub-list or2 in the lens image iK2 according to the optimal seam coefficients C(6)-C(10), to generate a corrected secondary vertex sub-list m2'. At step S410, the rasterization engine 610, the texture mapping circuit 620 and the blending unit 630 of the primary processor element 120 operate together to generate a working surface image F0 (e.g., Fig. 6B) according to the lens image iK0 thereof, the leftmost quad iK1' inputted and the corrected primary vertex sub-list m0' thereof; the rasterization engine 610, the texture mapping circuit 620 and the blending unit 630 of the secondary processor element 121 operate together to generate a working surface image F1 (e.g., Fig. 6C) according to the lens image iK1 thereof, the leftmost quad iK2' inputted and the corrected secondary vertex sub-list m1' thereof; and the rasterization engine 610, the texture mapping circuit 620 and the blending unit 630 of the secondary processor element 122 operate together to generate a working surface image F2 (e.g., Fig. 6D) according to the lens image iK2 thereof, the leftmost quad iK1' inputted and the corrected secondary vertex sub-list m2' thereof. FIG. 9B ) according to the lens image iK0 thereof, the leftmost quad iK1' inputted and the corrected primary vertex sub-list m0' thereof; the rasterization engine 610, the texture mapping circuit 620 and the blending unit 630 of the secondary processor element 121 operate together to generate a working surface image F1 (e.g., Fig. 6C) according to the lens image iK1 thereof, the leftmost quad iK2' inputted and the corrected secondary vertex sub-list m1' thereof; and the rasterization engine 610, the texture mapping circuit 620 and the blending unit 630 of the secondary processor element 122 operate together to generate a working surface image F2 (e.g., Fig. 6D) according to the lens image iK2 thereof, the leftmost quad iK1' inputted and the corrected secondary vertex sub-list m2' thereof. FIG. 9BThe rasterization engine 610, texture mapping circuit 620, and mixing unit 630 of the auxiliary processor element 122 work together to generate a working surface image F2 (e.g., based on its own lens image iK2 and its own modified auxiliary vertex sublist m2') FIG. 9B ).

[0137] In step S412 (i.e., transmission stage three), the GPU 132 of each processor element 120-122 cuts each working surface image into multiple tiles of a preset size (e.g., 64x64), calculates the histogram Hl of the leftmost column of tiles and / or the histogram Hr of the rightmost column of tiles in its own working surface image, and transmits a preset segment of its own working surface image to one or two adjacent processor elements. In one embodiment, the preset segment of the aforementioned working surface image is the leftmost eight columns of pixels and / or the rightmost eight columns of pixels in the working surface image. FIG. 9A As shown, the GPU 132 of the processor element 120 transmits the histogram Hr0 of the rightmost column of the working surface image F0 and the rightmost eight columns of pixels Fr0 to the IQE unit 133 of the processor element 121 through the output port 152. Then, it receives and parses the MIPI packet (including the histogram Hl1 of the leftmost column of the working surface image F1 and the leftmost eight columns of pixels Fl1) from the processor element 121 through the input port 153. According to the data type of the packet header, the histogram Hl1 and the leftmost eight columns of pixels Fl1 are stored in its own local volatile memory 170. The GPU 132 of processor element 121 transmits the histogram Hr1 of the rightmost column of tiles and the rightmost eight columns of pixels Fr1 of the working surface image F1 to the IQE unit 133 of processor element 122 through output port 152, transmits the histogram Hl1 of the leftmost column of tiles and the leftmost eight columns of pixels Fl1 of the working surface image F1 to the IQE unit 133 of processor element 120 through output port 155, receives and parses MIPI packets (including the histogram Hl2 of the leftmost column of tiles and the leftmost eight columns of pixels Fl2 of the working surface image F2) through input port 153, receives and parses MIPI packets (including the histogram Hr0 and the segment Fr0 of the working surface image F0) through input port 154, and stores the histograms Hr0 and Hl2 and the segments Fr0 and Fl2 in its own local volatile memory 171 according to the data type of the packet header. The GPU 132 of the processor element 122 transmits the histogram Hl2 of the leftmost column of the working surface image F2 and the leftmost eight columns of pixels Fl2 to the IQE unit 133 of the processor element 121 through the output port 155. It receives and parses MIPI packets (containing the histogram Hr1 and segment Fr1 of the working surface image F1) through the input port 154, and stores the histogram Hr1 and segment Fr1 in its own local volatile memory 172 according to the data type of the packet header.

[0138] At step S414, the IQE unit 133 of the processor element 120 performs image quality enhancement on the working face image F0 according to the received histogram Hl1 and the section Fl1 to generate an enhanced image F0'; the IQE unit 133 of the processor element 121 performs image quality enhancement on the working face image F1 according to the received histograms Hl2 and Hr0 and the sections Fl2 and Fr0 to generate an enhanced image F1'; the IQE unit 133 of the processor element 122 performs image quality enhancement on the working face image F2 according to the received histogram Hr1 and the section Fr1 to generate an enhanced image F2'. If method one (without the link 901) is performed: at step S416, the encoding and transmitting unit 134 of the processor elements 120-122 respectively encodes the three enhanced images F0'-F3' into three encoded video streams en0-en2, and transmits the three encoded video streams en0-en2 to the receiver 180 to generate a wide-angle image. If method two (with the link 901) is performed: at step S415 (transmission stage four), the IQE units 133 of the auxiliary processor elements 121-122 transmit the two enhanced images F1'-F2' to the encoding and transmitting unit 134 of the processor element 120 through the output ports 155 and 152, then the encoding and transmitting unit 134 of the processor element 120 receives and parses the MIPI packets (containing the two enhanced images F1'-F2') through the input ports 153-154, and stores the two enhanced images F1'-F2' in its local volatile memory 170 according to the data type of the packet header (e.g. 0x34); at step S416, the encoding and transmitting unit 134 of the processor element 120 merges the two enhanced images F1'-F2' with the enhanced image F0' to form a single bitstream, encodes the single bitstream into a single encoded video stream en, and transmits the single encoded video stream en to the receiver 180.

[0139] Please note that, since the multi-processor systems 400 and 800 are used to generate four working face images to form a panoramic image, the processor elements thereof are connected in a ring topology in the transmission stages one to three; in particular, for the multi-processor system 400, the processor elements thereof are connected to form a unidirectional ring topology in the transmission stages one and two, and are connected to form a bidirectional ring topology in the transmission stage three. In contrast, the three-processor system 900 is used to generate three working face images to form a panoramic image as shown in FIG. 1, and the processor elements thereof are connected in a ring topology in the transmission stages one to three. FIG. 9BIn the first to third transmission stages, the plurality of processor elements are connected in a linear topology; in particular, in the first and second transmission stages, the plurality of processor elements are connected in a unidirectional linear topology, and in the third transmission stage, the plurality of processor elements are connected in a bidirectional linear topology. In the first and second transmission stages, the direction of data transmission between the plurality of processor elements is opposite.

[0140] The above descriptions are only the preferred embodiments of the present application, not intended to limit the scope of the application. Any equivalent changes or modifications made without departing from the spirit of the present application shall be included within the scope of the application.

Claims

1. A multiprocessor system, characterized in that, Include: Multiple processor elements are coupled to a multi-lens camera, wherein the multi-lens camera captures a field of view having an X-degree horizontal field of view and a Y-degree vertical field of view, and each processor element includes: Multiple input / output (I / O) ports; and A processing unit j is coupled to the plurality of input / output ports; and Multiple links, each link connecting one of the multiple input / output ports of one of the multiple processor elements to one of the multiple input / output ports of another of the multiple processor elements, such that each processor element is connected to one or two adjacent processor elements by two or more links, each link being configured to transmit data in a unidirectional direction, wherein X <= 360 and Y < 180; A processing unit j includes: An image signal processor is used to acquire n images captured from the multi-lens camera. j Each lens image, and selectively transmitted with the n j Each lens image and zero or more first edge data responsible for outputting overlapping areas are given to a neighboring processor element; as well as An image processing unit, coupled to the image signal processor, is configured to perform a set of operations comprising: (1) selectively receiving first edge data input from another adjacent processor element; and (2) processing the data according to a first vertex sublist, wherein the n... j (3) Selectively transmit and receive multiple input and output joining coefficients from one or two neighboring processor elements, and (4) Based on the first vertex list, the multiple optimal joining coefficients, the input joining coefficients, the input first edge data, and the n... j Each shot produces n images. j n working face images, where n j >=1; Specifically, based on the plurality of responsible control regions, the plurality of output bonding coefficients are selected from the plurality of optimal bonding coefficients; The first vertex sublist contains multiple first vertices with a first data structure, and the multiple first data structures define the n. j First vertex mapping between a lens image and a projected image; Wherein, the plurality of optimal bonding coefficients represent different bonding degrees of the plurality of responsible control regions; as well as The projected image is a working surface image from all processor elements.

2. The system according to claim 1, characterized in that, If the projected image is a panoramic image, the plurality of processor elements are connected to form a ring topology; and if the projected image is a wide-angle image, the plurality of processor elements are connected to form a linear topology.

3. The system according to claim 1, characterized in that, The output first edge data is located at n j The first edge of each lens image, and the responsible control area having the plurality of output joining coefficients, are located in the n j The second edge of a lens image, wherein the first edge is located opposite the second edge.

4. The system according to claim 1, characterized in that, The size of the output first edge data is related to the size of each overlapping region, and the size of each overlapping region changes with the field of view of the multi-lens camera, the resolution of the lens sensor, and the lens mounting angle.

5. The system according to claim 1, characterized in that, The processing unit j further includes: An encoding and transmission unit is used to encode and transmit the n j Each working face image is encoded into an encoded video stream, and the encoded video stream is transmitted.

6. The system according to claim 1, characterized in that, The plurality of processor elements includes a primary processor element and at least one auxiliary processor element, and each auxiliary processor element is further connected to the primary processor element via one of the plurality of links, wherein the image processing unit of each auxiliary processor element is used to transmit at least one working surface image to the primary processor element, and wherein the processing unit of the primary processor element further includes: An encoding and transmission unit is configured to receive at least one input working surface image from the at least one auxiliary processor element, encode the at least one input working surface image and the working surface image from the image processing unit of the main processor element into a single encoded video stream, and transmit the single encoded video stream.

7. The system according to claim 1, characterized in that, The group operation also includes: Selectively transmit the output second edge data and the output tile histogram to the one or two adjacent processor elements, wherein the output second edge data and the output tile histogram are selected from the n j A working surface image, wherein the processing unit j further comprises: An image quality enhancement unit is configured to receive input second edge data and an input tile histogram from one or two adjacent processor elements, and to enhance the image quality of the n elements based on the input second edge data and the input tile histogram. j Image quality enhancement is performed on each working face image to produce n j An enhanced image; Wherein, the output second edge data is located in the n j One or both of the leftmost and rightmost edges of the working surface image, and the output tile histogram containing the n j One or both of the brick histograms of the leftmost edge and the rightmost edge of the working face image.

8. The system according to claim 7, characterized in that, The processing unit j further includes: An encoding and transmission unit is coupled to the image quality enhancement unit for transmitting the n... j An enhanced image is encoded into an encoded video stream, and the encoded video stream is transmitted.

9. The system according to claim 7, characterized in that, The plurality of processor elements includes a primary processor element and at least one auxiliary processor element, and each auxiliary processor element is further connected to the primary processor element via one of the plurality of links, wherein the image processing unit of each auxiliary processor element is used to transmit at least one enhanced image to the primary processor element, and wherein the processing unit of the primary processor element further includes: An encoding and transmission unit is configured to receive at least one input enhanced image from the at least one auxiliary processor element, encode the at least one input enhanced image and an enhanced image from the image processing unit of the main processor element into a single encoded video stream, and transmit the single encoded video stream.

10. The system according to claim 1, characterized in that, The operation of determining the plurality of optimal bonding coefficients in (2) includes: Multiple test bonding coefficients are determined based on the offset of the center of one lens of the multi-lens camera relative to the system center of the multi-lens camera; Based on the multiple test bonding coefficients, the texture coordinates of each first vertex in the first vertex list in each lens image are corrected to generate a second vertex list. Based on the second vertex list, the input first edge data, and n j Each lens image forms a two-dimensional error table, which includes different values ​​of the multiple test bonding coefficients and the corresponding cumulative pixel value error in the multiple control areas. as well as The optimal bonding coefficient for each responsible control area is determined based on at least one local minimum of the cumulative pixel value error of one or two nearest control areas in the two-dimensional error table. The second vertex list contains multiple second vertices with a second data structure, wherein the multiple second data structures define the n. j The second vertex mapping between the lens image and the projected image.

11. The system according to claim 1, characterized in that, The (4) generates the n j The operations for each working surface image include: Based on the plurality of optimal bonding coefficients and the plurality of input bonding coefficients, the texture coordinates of each first vertex in the first vertex list in each lens image are corrected to generate a third vertex list. as well as According to the n j For each shot image, rasterization, texture mapping, and blending operations are performed on each point within each polygon formed by each group of third vertices from the third vertex list to generate the n... j Image of the work surface; The third vertex list contains multiple third vertices with third data structures, and the multiple third data structures define the n. j The third vertex mapping between the lens image and the projected image.

12. The system according to claim 1, characterized in that, Each working surface image is a preset projection of a corresponding lens image from the multi-lens camera.

13. The system according to claim 12, characterized in that, The preset projection is one of the following: equidistant rectangular projection, Miller projection, Mercator projection, Lambert cylindrical equal-area projection, and Panini projection.

14. The system according to claim 1, characterized in that, Each overlapping region contains P1 control regions arranged in a row, and P1>=3.

15. An image processing method, characterized in that, Suitable for a multiprocessor system coupled to a multilens camera, the multilens camera capturing a field of view having an X-degree horizontal field of view and a Y-degree vertical field of view, the multiprocessor system comprising multiple processor elements j and multiple links, each processor element j being connected to one or two adjacent processor elements j by two or more links, each link being configured to transmit data in a unidirectional direction, the method comprising: For a processor element j: Obtain n captured from the multi-lens camera j Individual shots; In a first transmission phase, selectively send and receive with the n j A lens image and zero or more first edge data related to the input and output of the overlapping area, and from one or two adjacent processor elements j; According to a first vertex order list, n j The image from each lens and the input first edge data determine the optimal joining coefficient for multiple control areas within the overlapping area; In a second transmission phase, selectively transmit and receive input and output junction coefficients from one or two adjacent processor elements j; as well as Based on the first vertex sublist, the plurality of optimal joining coefficients, the input joining coefficients, the input first edge data, and n j Each shot produces n images. j n working face images, where n j >=1, X<=360 and Y<180; Specifically, based on the plurality of responsible control regions, the plurality of output bonding coefficients are selected from the plurality of optimal bonding coefficients; The first vertex sublist contains multiple first vertices with a first data structure, and the multiple first data structures define the n. j The first vertex mapping between a camera image and a projected image; and The projected image is a working surface image from all processor elements j; Wherein, the plurality of optimal bonding coefficients represent different bonding degrees of the plurality of responsible control regions; Each working surface image is a preset projection derived from a corresponding lens image of the multi-lens camera.

16. The method according to claim 15, characterized in that, The output first edge data is located at n j The first edge of each lens image, and the control region responsible for the output joining coefficient, are located in the n... j The second edge of a lens image, wherein the first edge is opposite to the second edge.

17. The method according to claim 15, characterized in that, The size of the output first edge data is related to the size of each overlapping region, and the size of each overlapping region changes with the field of view of the multi-lens camera, the resolution of the lens sensor, and the lens mounting angle.

18. The method according to claim 15, characterized in that, Also includes: Regarding the processor element j: The n j Each work surface image is encoded into an encoded video stream; and Transmit the encoded video stream.

19. The method according to claim 15, characterized in that, Also includes: For each of at least one auxiliary processor element: Transmit at least one working surface image to a main processor element; Regarding the main processor element: Receive at least one input working surface image from the at least one auxiliary processor element; Encode the at least one input work surface image and the work surface image generated by the main processor element into a single encoded video stream; and Transmit the single encoded video stream; The plurality of processor elements j includes the main processor element and the at least one auxiliary processor element, and each auxiliary processor element is further connected to the main processor element by one of the plurality of links.

20. The method according to claim 15, characterized in that, In the first transmission phase and the second transmission phase, the data transmission directions between the plurality of processor elements j are opposite.

21. The method according to claim 15, characterized in that, Also includes: Regarding the processor element j: Selectively transmit and receive input and output second edge data and input and output brick histograms from one or two neighboring processor elements j; as well as Based on the input second edge data and the input brick histogram, for the n j Image quality enhancement is performed on each working face image to produce the n... j An enhanced image; Wherein, the output second edge image data is located in the n j One or both of the leftmost and rightmost edges of the working surface image, and the output brick histogram containing the n j One or both of the histograms of the leftmost and rightmost edges of the working face image. The image quality enhancement includes contrast enhancement, low-pass filtering, and image sharpening.

22. The method according to claim 21, characterized in that, Also includes: Regarding the processor element j: The n j Encode an enhanced image into a single encoded video stream; and Transmit the encoded video stream.

23. The method according to claim 21, characterized in that, Also includes: For each of at least one auxiliary processor element: Transmit at least one enhanced image to a main processor element; Regarding the main processor element: Receive at least one input enhanced image from the at least one auxiliary processor element; Encode the at least one input enhanced image and the enhanced image generated by the main processor element into a single encoded video stream; as well as Transmit the single encoded video stream; The plurality of processor elements j includes the main processor element and the at least one auxiliary processor element, and each auxiliary processor element is further connected to the main processor element by one of the plurality of links.

24. The method according to claim 15, characterized in that, The step of determining the plurality of optimal bonding coefficients includes: Multiple test bonding coefficients are determined based on the offset of the center of one lens of the multi-lens camera relative to the system center of the multi-lens camera; Based on the multiple test bonding coefficients, the texture coordinates of each first vertex in the first vertex list in each lens image are corrected to generate a second vertex list. Based on the second vertex list, the input first edge data, and n j Each lens image forms a two-dimensional error table, which includes different values ​​of the multiple test bonding coefficients and the corresponding cumulative pixel value error of the multiple control areas. as well as The optimal bonding coefficient for each responsible control area is determined based on at least one local minimum of the cumulative pixel value error of one or two nearest control areas in the two-dimensional error table. The second vertex list contains multiple second vertices with a second data structure, wherein the multiple second data structures define the n. j The second vertex mapping between the lens image and the projected image.

25. The method according to claim 15, characterized in that, The generation of n j The steps involved in creating a working surface image include: Based on the plurality of optimal bonding coefficients and the plurality of input bonding coefficients, the texture coordinates of each first vertex in the first vertex list in each lens image are corrected to generate a third vertex list. as well as According to the n j For each shot image, rasterization, texture mapping, and blending operations are performed on each point within each polygon formed by each group of third vertices from the third vertex list to generate the n... j Image of the work surface; The third vertex list contains multiple third vertices with third data structures, and the multiple third data structures define the n. j The third vertex mapping between the lens image and the projected image.

26. The method according to claim 15, characterized in that, Each overlapping region contains P1 control regions arranged in a row, and P1>=3.

27. The method according to claim 15, characterized in that, The projected image is either a panoramic image or a wide-angle image.

Citation Information

Patent Citations

  • Capturing and aligning panoramic image and depth data

    US20180139431A1

  • Graphics processing systems with multiple processors connected in a ring topology

    US7623131B1