Image processing method and image processor based on Gaussian splashing

By converting a two-dimensional Gaussian ellipse into a two-dimensional circle under its intrinsic coordinate system and optimizing it during overlap detection and sorting, the problems of uneven calculation load, many false positives and low rendering efficiency in the rendering process are solved, and more efficient rendering performance is achieved.

CN119991904APending Publication Date: 2025-05-13PEKING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510064538.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the rendering process, three-dimensional Gaussian splashing has problems such as uneven calculation load distribution, many false positives in overlap detection, and inefficient rendering caused by object occlusion relationship.

Method used

By converting a 2D Gaussian ellipse into a 2D circle under its eigencoal system, using the perfect circle in the eigencoal system as the collision box for overlap tests, and depth prediction and triple sorting tree processing are performed during the sorting process to reduce false positives and improve rendering efficiency.

Benefits of technology

It realizes the reduction of parallel thread computing complexity, improves hardware rendering efficiency, reduces computing energy consumption and hardware overhead, and improves the utilization rate of rendering pipelines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991904A_ABST
    Figure CN119991904A_ABST
Patent Text Reader

Abstract

According to the image processing scheme based on Gaussian splashing, pixel block information of an image covered by each of a plurality of two-dimensional Gaussian ellipses is determined by using two-dimensional circles corresponding to each of the plurality of two-dimensional Gaussian ellipses; determining at least one two-dimensional Gaussian ellipse associated with each pixel block in the image according to the pixel block information covered by each two-dimensional Gaussian ellipse; and rendering a pixel block associated with the at least one two-dimensional Gaussian ellipse in the image by using the at least one two-dimensional Gaussian ellipse. According to the scheme, the two-dimensional Gaussian ellipse is converted into the corresponding two-dimensional circle under the intrinsic coordinate system to participate in the subsequent image processing process, so that decoupling between coordinates in subsequent processing calculation is realized, the parallel thread calculation complexity is effectively reduced, and the processing efficiency is improved. Therefore, the hardware can be rendered at a smaller pixel block scale during rendering processing, the utilization rate is improved, and the hardware overhead and the calculation energy consumption are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to an image processing method and an image processor based on Gaussian splatting. Background Art

[0002] High-quality and realistic image rendering is crucial for many applications. In order to ensure the quality of image rendering, 3D Gaussian Splatting is a widely used rendering algorithm. Although 3D Gaussian Splatting has better rendering performance than traditional neural mesh algorithms represented by neural radiation fields, there are still some shortcomings in the rendering process, such as: uneven calculation of pressure distribution; the use of square collision boxes in the overlap detection process easily leads to a large number of false positives in overlap detection, which increases the amount of data required to be processed in the rendering process; and due to the occlusion relationship of objects in complex scenes, a considerable number of Gaussian imprints will not be rendered, but these Gaussian imprints still need to be preprocessed and sorted for rendering, which significantly reduces the utilization of the rendering pipeline. Therefore, how to improve the above-mentioned 3D Gaussian splatting to improve rendering effects and efficiency is an urgent problem to be solved. Summary of the invention

[0003] In view of the above problems mentioned in the background technology, the present application provides an image processing method and an image processor based on Gaussian splattering, and accordingly, also provides an electronic device and a readable storage medium.

[0004] In a first embodiment, the present application provides an image processing method based on Gaussian splatting, the method comprising:

[0005] Determine the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses;

[0006] Using the two-dimensional circles corresponding to the two-dimensional Gaussian ellipses, respectively, to determine pixel block information of an image covered by the two-dimensional Gaussian ellipses;

[0007] Determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image according to pixel block information respectively covered by the multiple two-dimensional Gaussian ellipses;

[0008] Using the at least one two-dimensional Gaussian ellipse, a pixel block in the image associated with the at least one two-dimensional Gaussian ellipse is rendered.

[0009] In a second embodiment, the present application provides an image processor based on Gaussian splatting, the image processor comprising:

[0010] A preprocessing module, used to determine the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses; using the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses, determine the pixel block information of an image covered by each of the multiple two-dimensional Gaussian ellipses; and according to the pixel block information covered by each of the multiple two-dimensional Gaussian ellipses, determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image;

[0011] A rendering module is used to render a pixel block associated with the at least one two-dimensional Gaussian ellipse in the image using the at least one two-dimensional Gaussian ellipse.

[0012] In a third embodiment, the present application provides an electronic device. The electronic device includes a memory and the image processor provided in the second embodiment, wherein the memory is used to store a program; and the image processor is coupled to the memory and is used to execute the program stored in the memory to implement each method embodiment provided in the present application.

[0013] In a fourth embodiment, the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the steps in each method embodiment provided by the present application can be implemented.

[0014] The technical solution provided by the embodiment of the present application is to use the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses to determine the pixel block information of an image covered by each of the multiple two-dimensional Gaussian ellipses; then determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image based on the pixel block information covered by each of the multiple two-dimensional Gaussian ellipses; and then use at least one two-dimensional Gaussian ellipse to render a pixel block in the image associated with the at least one two-dimensional Gaussian ellipse. This solution, by converting the two-dimensional Gaussian ellipse into the corresponding two-dimensional circle in its intrinsic coordinate system for participating in the subsequent image processing process, realizes the decoupling between coordinates in the subsequent processing calculation, which is conducive to effectively reducing the computational complexity of parallel threads, so that the hardware can render at a smaller pixel block scale when performing rendering processing, improves utilization, and reduces hardware overhead and computing energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1A schematic diagram of a process flow of an image processing method based on Gaussian splatting provided in one embodiment of the present application;

[0017] Figure 2 A schematic diagram of the structure of an image processor based on Gaussian splattering provided in one embodiment of the present application;

[0018] Figure 3 An example diagram of the change of the intrinsic coordinates provided in one embodiment of the present application;

[0019] Figure 4 The collision box intention of the two-dimensional Gaussian ellipse provided in the embodiment of the present application;

[0020] Figures 5 to 7 An example diagram of rendering performance comparison provided in one embodiment of the present application. DETAILED DESCRIPTION

[0021] In the field of computer vision and graphics, novel view synthesis (also known as novel view synthesis, Novel View Synthesis, NVS) has always been a problem that has received much attention. Novel view synthesis aims to generate scene images from new, unseen perspectives based on a set of images from known perspectives. 3D Gaussian Splatting (3DGS) is a new type of 3D scene representation and rendering method (also known as a SOTA algorithm) that has recently emerged in the task of novel view synthesis. It has high-quality rendering capabilities and fast rendering speed. Among them, SOTA (State Of The Art) algorithm refers to the algorithm, model or technology used to describe the current best algorithm, model or technology in a specific field or problem. Compared with traditional novel view synthesis algorithms such as neural network algorithms represented by Neural Radiance Field (NeRF), 3D Gaussian Splatting does not contain a neural network structure, and it also exceeds the neural network algorithm in training speed, rendering speed and rendering quality. Specifically, the 3D Gaussian splashing algorithm models a scene as an explicit set of 3D Gaussian ellipsoids, describes the scene features through the size, shape, and transparency of the Gaussian ellipsoids, and uses the fourth-order spherical harmonic coefficients to describe the visual effects that depend on the viewing direction. During the rendering process, the 3D Gaussian splashing algorithm first removes the ellipsoids outside the field of view light cone, and then "splashes" the 3D Gaussian ellipsoids within the field of view onto the corresponding camera plane, converting the 3D features into 2D, and extracting the color of the corresponding viewing direction from the spherical harmonic coefficients to form a 2D Gaussian imprint. After that, the 3D Gaussian splashing algorithm confirms the influence range of each Gaussian ellipsoid through a simple overlap test, and performs subsequent local sorting and rendering operations in units of 16×16 pixel blocks. Local sorting sorts the Gaussian imprints by the depth of the Gaussian imprints relative to the camera plane to ensure the correct occlusion order relationship solution during rendering. Finally, pixel-level α-calculation and α-blending are performed to achieve parallel rendering in local pixel blocks to generate the final color image.

[0022] The above-mentioned 3D Gaussian splash mainly realizes real-time new perspective rendering on graphics hardware at the server and PC level. However, it still faces great challenges and many problems to realize real-time rendering on consumer-level and edge graphics hardware (such as Jetson NX, a high-performance, low-power module system designed for edge computing). The main problems are as follows:

[0023] 1) There is an uneven distribution of computational load and too much useless data in the rendering process. In the rendering pipeline of 3D Gaussian splatter, the rendering calculation step accounts for about 65% of the total execution delay, of which the most important contribution comes from the large number of floating-point calculations performed in the pixel-level α-calculation and α-blending. These calculations are assigned to each pixel execution thread, resulting in a large number of repeated calculations by the underlying computing hardware. Moreover, most of the calculations are concentrated in the pixel-level parallel rendering threads, and a large number of floating-point calculations are used at the same time, resulting in the rendering performance being severely limited by the hardware computing power, limiting the real-time rendering application of the algorithm on edge graphics hardware.

[0024] 2) During the rendering process, Gaussian traces usually involve the surface of objects in the scene (specifically, they involve approximating the surface of objects in the scene with a series of 3D Gaussian distributions). Due to the occlusion relationship between objects, a considerable number of Gaussian traces will not contribute to the rendering process. Specifically, in the rendering of a single pixel block, due to the occlusion relationship between objects in a complex scene, a considerable number of Gaussian traces will not be rendered, but these Gaussian traces still need to go through the preprocessing and sorting rendering process, which significantly reduces the utilization rate of the rendering pipeline.

[0025] 3) Since a square collision box is used in the overlap detection process, and the Gaussian footprint is usually an ellipse with a large eccentricity, this leads to a large number of false positives in the overlap detection, which increases the amount of data to be processed in the rendering process and increases the data pressure in the sorting and rendering steps. In other words, the square collision box used in the overlap detection in the three-dimensional Gaussian splash is too large, surrounding a large amount of invalid space around the Gaussian ellipsoid, which leads to a large number of false positive overlap test results in the preprocessing process, increasing the amount of data to be processed in the sorting and rendering process.

[0026] In view of the above problems, the present application provides a data processing solution, in which an improved three-dimensional Gaussian splashing is used to process images. Among them, in the present application solution, the following improvements are made to the three-dimensional Gaussian splashing:

[0027] a) During the Gaussian feature preprocessing, the elliptical Gaussian imprint is transformed into a perfect circle by coordinate transformation to its intrinsic coordinate system.

[0028] b) In the overlap calculation process, the circumscribed square of the perfect circle in the intrinsic coordinate system is used as the new collision box in the overlap test, which effectively reduces the area covered by the collision box, reduces the probability of false positives, and speeds up the entire rendering process.

[0029] c) Before the sorting process, the depth of each rendered pixel block is predicted, and the Gaussian imprint is sorted according to the predicted depth using a triple sorting tree of bucket sorting, radix sorting, and CAS sorting. Gaussian imprints of different depths are processed in blocks and terminated early in a timely manner, thereby reducing the decrease in hardware utilization caused by complex spatial occlusion relationships.

[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0031] In some processes described in the specification, claims and the above-mentioned figures of this application, multiple operations appearing in a specific order are included, and these operations may not be executed or executed in parallel in the order in which they appear in this article. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to different types. The term "or / and" in this application is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example: A or / and B, indicating that A can exist alone, A and B can exist at the same time, and B can exist alone; the character " / " in this application generally indicates that the front and back associated objects are an "or" relationship. It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a product or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such product or system. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the product or system including the elements. In addition, the following embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0032] The technical solutions provided by each embodiment of the present application are introduced and explained below.

[0033] In this application, the various method embodiments provided are applied to Figure 2The Gaussian splash-based image processor is shown. The image processor is also called a Gaussian rendering accelerator, which mainly includes the following three modules: a preprocessing module (a Gaussian preprocessing module), a sorting module (a depth-guided Gaussian sorting module), and a rendering module (a LUT rendering module). The preprocessing module includes a culling unit (FCU), an ellipse transformation unit (SNU), and an overlap detection unit (also called an overlap test unit, ITU). The main functions of each unit included in the preprocessing module are as follows:

[0034] The FCU is used to remove the three-dimensional Gaussian ellipsoid outside the view cone, and for this reason the FCU is also called the view cone culling unit. In the original Gaussian rendering pipeline, only the three-dimensional Gaussian ellipsoid that is very close to the camera's field of view will be culled, which results in a large amount of preprocessing work and low efficiency. The FCU design in this application adopts a more stringent culling method: that is, to remove the Gaussian ellipsoid that exceeds 120% of the view cone size, so as to reduce the amount of data processing while retaining background information as much as possible.

[0035] The ellipse transformation unit (SNU) is used to transform the three-dimensional Gaussian feature into two dimensions, and transform the two-dimensional Gaussian ellipse into a two-dimensional circle (such as a unit circle) in its intrinsic coordinate system. The intrinsic coordinate system refers to a coordinate system in which the covariance matrix of the Gaussian distribution is presented in a diagonalized form. The two-dimensional Gaussian ellipse described in this application is also called a two-dimensional Gaussian footprint, which is obtained by projecting the three-dimensional Gaussian ellipsoid in the three-dimensional coordinate system onto a planar two-dimensional coordinate system, and the two-dimensional coordinate system can be, for example, a world coordinate system. For this reason, the above-mentioned ellipse transformation unit can also be called a footprint transformation unit.

[0036] The overlap test unit (ITU) is used to locate the pixel blocks in the image that intersect with the two-dimensional Gaussian ellipse. For the detailed description of the pixel blocks, please refer to the relevant content in other embodiments, and no further details will be given here. In the present application, when processing an image, in order to reduce the computational cost of performing Gaussian operations on each pixel in the image, the image will first be divided into a plurality of non-overlapping pixel blocks of the same size. Therefore, it can be understood that the pixel block described in the present application refers to a certain pixel group in the image to be processed, wherein the size of the pixel block may be, for example, 16*16 (i.e., the pixel block contains 16*16 pixels).

[0037] related Figure 2 The specific implementation of the functions of each module / unit included in the image processor shown will be described in detail in the following other embodiments.

[0038] Figure 1 FIG. 2 shows a schematic flow chart of an image processing method based on Gaussian splattering provided by an embodiment of the present application. Figure 1 As shown, the image processing method comprises the following steps:

[0039] 101. Determine two-dimensional circles corresponding to multiple two-dimensional Gaussian ellipses;

[0040] 102. Determine pixel block information of an image covered by each of the plurality of two-dimensional Gaussian ellipses by using the two-dimensional circles corresponding to each of the plurality of two-dimensional Gaussian ellipses;

[0041] 103. Determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image according to information of pixel blocks covered by each of the multiple two-dimensional Gaussian ellipses;

[0042] 104. Use the at least one two-dimensional Gaussian ellipse to render a pixel block associated with the at least one two-dimensional Gaussian ellipse in the image.

[0043] In the above 101, the multiple two-dimensional Gaussian ellipses are obtained by transforming multiple three-dimensional Gaussian ellipsoids within the viewing cone. The viewing cone defines the field of view of the camera / person, that is, the viewing cone is equivalent to defining the upper, lower, left, and right viewing boundaries of a camera / person. The above multiple three-dimensional Gaussian ellipsoids are obtained by eliminating a set of input three-dimensional Gaussian ellipsoids using the viewing cone. The input set of three-dimensional Gaussian ellipsoids comes from the input pre-trained three-dimensional Gaussian splash (3DGS) model. The three-dimensional Gaussian splash model uses a set of three-dimensional Gaussian ellipsoids to simulate three-dimensional scenes to achieve high-quality real-time rendering of complex scenes.

[0044] For example, when a pre-trained three-dimensional Gaussian splash model (a machine learning model including a set of three-dimensional Gaussian ellipsoids) is input into Figure 2 When the image processor shown is used, the entire model is generally directly input. In this case, a lot of information that is not actually used in the model is also input. For example, since the camera / human perspective has an upper, lower, left, and right boundary, the part beyond this boundary will not be seen, and naturally will not be presented in the image. Therefore, the three-dimensional Gaussian ellipsoids that exceed this boundary in the input set of three-dimensional Gaussian ellipsoids are information that is not actually used in the image rendering process.

[0045] For this, see Figure 2 As shown, a rejection unit is provided in the image processor provided in the present application. The rejection unit is used to determine which three-dimensional Gaussian ellipsoids are outside the viewing cone, thereby rejecting the three-dimensional Gaussian ellipsoids outside the viewing cone so that they do not participate in subsequent calculations, thereby saving computing resources.

[0046] In specific implementation, after the ellipsoid characteristic information of any one of a group of three-dimensional Gaussian ellipsoids is input into the image processor, the elimination unit will determine whether the three-dimensional Gaussian ellipsoid is located in the viewing cone by performing the following steps: obtaining the center coordinates of the three-dimensional Gaussian ellipsoid from the ellipsoid characteristic information; then, normalizing the center coordinates to obtain the normalized value of the center coordinates; if the normalized value is not within the set data range (such as greater than 1 or less than 0), it will be determined that it is not within the viewing cone range (outside the field of view), so it will be terminated in advance to no longer perform subsequent processing, and the three-dimensional Gaussian ellipsoid will be eliminated. If the normalized value is within the set value range (such as [0, 1] range), it will be determined that the three-dimensional Gaussian ellipsoid is within the viewing cone range; for the three-dimensional Gaussian ellipsoid within the viewing cone range, coordinate conversion will be performed to achieve the conversion of its three-dimensional spatial features into corresponding two-dimensional spatial features, thereby achieving the conversion of the three-dimensional Gaussian ellipsoid into a two-dimensional Gaussian ellipse (also called a two-dimensional Gaussian footprint).

[0047] Based on this, in addition to executing the above step 101, the method provided in this embodiment may further include the following steps:

[0048] 100a. According to the ellipsoid characteristic information of each three-dimensional Gaussian ellipsoid in a group of three-dimensional Gaussian ellipsoids, determine a three-dimensional Gaussian ellipsoid in the group of three-dimensional Gaussian ellipsoids that is located within a viewing cone; wherein the viewing cone is used to describe a specified field of view;

[0049] 100b, converting multiple three-dimensional Gaussian ellipsoids located within the viewing cone in the group of three-dimensional Gaussian ellipsoids to obtain two-dimensional Gaussian ellipses corresponding to each of the multiple three-dimensional Gaussian ellipsoids, thereby triggering the execution of the above step 101.

[0050] Furthermore, the first three-dimensional Gaussian ellipsoid is any three-dimensional Gaussian ellipsoid in the above group of three-dimensional Gaussian ellipsoids. For the first three-dimensional Gaussian ellipsoid, the implementation of the above step 100a includes:

[0051] 100a1. Acquire the center coordinates of the first three-dimensional Gaussian ellipsoid from the ellipsoid feature information of the first three-dimensional Gaussian ellipsoid;

[0052] 100a2. Normalizing the center coordinates to obtain corresponding normalized values;

[0053] 100a3. If the normalized value is within a set value range, the first three-dimensional Gaussian ellipse is located within the viewing cone;

[0054] 100a4. If the normalized value is not within the set value range, the first three-dimensional Gaussian ellipse is outside the viewing cone.

[0055] Wherein, the ellipsoid feature information described above includes but is not limited to the following information: center coordinates (i.e., the position information of the center point of the ellipsoid, including the three-dimensional coordinate information of the x-axis, y-axis and z-axis), spherical harmonic coefficients (SH), rotation features and scaling features, opacity (o). The above-mentioned spherical harmonic coefficients are a group of 27 floating-point numbers, which describe the color of the ellipsoid at different viewing angles. The spherical harmonic coefficients are used to store the color information of the ellipsoid. The rotation feature is a group of rotation quaternions (which can also be the corresponding four-dimensional rotation matrix), and the scaling feature is a scaling matrix (such as a 3-dimensional diagonal matrix). Based on the above-mentioned rotation quaternion array (matrix) and scaling matrix, the covariance matrix of the three-dimensional Gaussian ellipsoid can be determined.

[0056] It can be known from the above content that the two-dimensional Gaussian ellipse in the above 101 is specifically obtained by converting the corresponding three-dimensional Gaussian ellipsoid in the three-dimensional coordinate system to the two-dimensional coordinate system. That is, the two-dimensional Gaussian ellipse is an ellipse in the two-dimensional coordinate system, and the two-dimensional coordinate system can be but not limited to the world coordinate system. Further, after the conversion is completed, the two-dimensional Gaussian ellipse will be converted to a two-dimensional circle in its intrinsic coordinate system. Here, the two-dimensional Gaussian ellipse is transformed into a two-dimensional circle (a perfect circle) in its intrinsic coordinate system through coordinate transformation, which can realize the decoupling of X and Y coordinates in the rendering calculation, so that the α-calculation and α-mixing process can be pre-calculated in units of two-dimensional Gaussian ellipses, and quantitative incremental calculations can be performed at the pixel level, thereby effectively reducing the computational complexity of parallel threads, enabling the hardware to render at a smaller pixel block scale, improving utilization, and reducing hardware overhead and computing energy consumption.

[0057] Figure 3 , a schematic diagram of the process of transforming a two-dimensional Gaussian ellipse in the world coordinate system into a two-dimensional circle in its intrinsic coordinate system is shown in . Wherein, in the intrinsic coordinate system, the covariance matrix of the Gaussian distribution is a diagonal matrix.

[0058] In practical applications, the α calculation formula for Gaussian splashing is generally as follows: Among them, α i is the final opacity, which is the learned opacity o i The product of the two-dimensional Gaussian ellipse (i.e., the two-dimensional Gaussian footprint) and the Gaussian distribution; Δ is the position vector from the two-dimensional Gaussian ellipse (i.e., the two-dimensional Gaussian footprint) to the center of the pixel; Σ is the covariance matrix corresponding to the two-dimensional Gaussian ellipse. However, under normal circumstances, the covariance matrix Σ is not a diagonal matrix, and there is a coupling phenomenon between the X and Y coordinates. Therefore, different pixels cannot be simply derived based on the calculation results of adjacent pixels, but must be recalculated for each pixel. The above covariance matrix is ​​defined as: Σ=TT T, T is the transformation matrix from the two-dimensional Gaussian ellipse intrinsic coordinate system to the world coordinate system. Therefore, if the basis vectors i, j of the coordinate system are projected into the intrinsic coordinate system through the inverse transformation of T, the covariance matrix can be transformed into the unit matrix, and the two-dimensional Gaussian kernel can be transformed into the product of two one-dimensional Gaussian kernels, thereby realizing the decoupling of the X and Y coordinates. That is, after the coordinate transformation, the α (opacity) calculation formula will be changed to the following formula: α = o· Where G represents the Gaussian kernel function.

[0059] Subsequently, the final color rendering can be performed using the two-dimensional Gaussian ellipse through the following α-blending formula: C = ∑(T i c i α i ). Where T is the transmittance accumulated from the first two-dimensional Gaussian ellipse in the pixel block to the two-dimensional Gaussian ellipse in the current rendering, T i =Π(1-α i ) represents the color contribution of the current two-dimensional Gaussian ellipse to the pixel block under the spatial occlusion relationship, c i Represents the color vector decoded from the spherical harmonic coefficients.

[0060] That is, the implementation of the above step 101 can be understood as: normalizing the two-dimensional Gaussian ellipse in the world coordinate system to normalize it to the corresponding unit circle (that is, a two-dimensional circle) in the intrinsic coordinate system, so as to assist the subsequent quantitative calculation and rendering process.

[0061] Also, see Figure 2 This step 101 is implemented by the ellipse transformation unit in the preprocessing module. The subsequent quantization calculation and rendering process can be assisted by the corresponding collision box difference.

[0062] Figure 4 The difference between the collision box used by the two-dimensional Gaussian ellipse in the intrinsic coordinate system and the collision box used in the world coordinate system is shown in Figure 1. Figure 4 In the three-dimensional Gaussian splash algorithm, the collision box is depicted as a circumscribed square of a circle with the major semiaxis of the two-dimensional Gaussian ellipse as the radius; while the collision box in the intrinsic coordinate system (in Figure 4 In the 3D Gaussian algorithm, most of the 2D Gaussian traces are ellipses with relatively high eccentricity. Therefore, the collision box in the intrinsic coordinate system will have significantly fewer intersecting pixel blocks than the collision box in the original algorithm. Figure 4The red pixel blocks in the figure are the number of intersecting pixel blocks reduced after the collision box in the intrinsic coordinate system is used. It can be seen that in the subsequent overlap detection, the collision box in the eigen coordinate system (that is, the circumscribed square of the two-dimensional circle (perfect circle) in the intrinsic coordinate system is used as the new collision box in the overlap detection) can effectively reduce the area covered by the collision box, which can significantly reduce the number of false positive pixel blocks in the subsequent overlap detection (that is, the probability of occurrence), thereby reducing the data and calculation pressure of the underlying rendering thread, and speeding up the entire rendering process.

[0063] The above combination Figure 4 The subsequent overlap detection refers to the implementation of the above step 102.

[0064] In the above 102, the image is an image to be rendered. The image is divided into a plurality of non-overlapping pixel blocks of the same size, each pixel block including 16*16 pixels. Figure 2 This step 102 is implemented by an overlap detection unit (ITU). Specifically, the overlap detection unit determines whether the corresponding two-dimensional Gaussian ellipse affects the pixel block (i.e., covers the pixel block) by detecting the intersection of the two sides of the circumscribed rectangle of the two-dimensional circle closest to the center of the pixel block and the pixel block.

[0065] For example, see Figure 4 If the pixel block in the first row and second column is closest to the left long side and the upper wide side of the intrinsic collision box (i.e., the circumscribed rectangle of a two-dimensional circle in the intrinsic coordinate system), it will be further determined whether the lower right corner of the pixel block has entered the range of the left long side and the upper wide side of the intrinsic collision box; if it has entered, it is determined that the pixel block intersects with the intrinsic collision box, and thus it is determined that the two-dimensional Gaussian ellipse corresponding to the two-dimensional circle covers the pixel block. On the contrary, if it has not entered, it is determined that the pixel block does not intersect with the intrinsic collision box, and thus it is determined that the two-dimensional Gaussian ellipse corresponding to the two-dimensional circle does not cover the pixel block.

[0066] By analogy, the overlap detection unit can detect the pixel block information of an image covered by each of the two-dimensional Gaussian ellipses in the multiple two-dimensional Gaussian ellipses.

[0067] Thus, in a specific implementation, if the first two-dimensional Gaussian ellipse is one of the multiple two-dimensional Gaussian ellipses, then correspondingly, in the above 102, “using the two-dimensional circle corresponding to the first two-dimensional Gaussian ellipse to determine the pixel block information of an image covered by each of the first two-dimensional Gaussian ellipses” may include:

[0068] 1021. Determine a circumscribed rectangle of the two-dimensional circle, wherein all sides of the circumscribed rectangle are tangent to the two-dimensional circle;

[0069] 1022. Determine pixel block information of an image covered by the first two-dimensional Gaussian ellipse using the circumscribed rectangle of the two-dimensional circle.

[0070] Furthermore, based on the pixel block information of an image covered by each detected two-dimensional Gaussian ellipse, the above step 103 will be triggered to determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image, that is, all Gaussian traces affecting a certain pixel block are grouped together for subsequent sorting and rendering processes.

[0071] Among them, considering that a two-dimensional Gaussian ellipse may cover multiple pixel blocks in the image, in order to realize the association of subsequent pixel blocks with corresponding one or more two-dimensional Gaussian ellipses, one processing method is: according to the number of pixel blocks covered by a two-dimensional Gaussian ellipse, the two-dimensional Gaussian ellipse is copied, and a unique identifier is assigned to the copied two-dimensional Gaussian ellipse, which can be the block identification (ID) of the pixel block intersecting with it. In this way, each pixel block can be associated with at least one two-dimensional Gaussian ellipse.

[0072] For example, if the pixel blocks of an image covered by the two-dimensional Gaussian ellipse 1 include: pixel block p1, pixel block p2, pixel block p3, pixel block p4, then four copies of the two-dimensional Gaussian ellipse 1 will be copied, such as the four copied two-dimensional Gaussian ellipses 1 include: two-dimensional Gaussian ellipse 1_1, two-dimensional Gaussian ellipse 1_2, two-dimensional Gaussian ellipse 1_3 and two-dimensional Gaussian ellipse 1_4, then the block identifier of pixel block 1 can be assigned to the two-dimensional Gaussian ellipse 1_1, the block identifier of pixel block p2 can be assigned to the two-dimensional Gaussian ellipse 1_2, the block identifier of pixel block p3 can be assigned to the two-dimensional Gaussian ellipse 1_2, and the block identifier of pixel block p4 can be assigned to the two-dimensional Gaussian ellipse 1_2. Similarly, if the pixel blocks of an image covered by the two-dimensional Gaussian ellipse 2 include: pixel block p2, pixel block p3, pixel block p5. Then the two-dimensional Gaussian ellipse 2 will be copied into three copies, and the block identifiers of pixel block p2, pixel block p3, and pixel block p5 can be assigned to the three copied two-dimensional Gaussian ellipses 2 respectively. In this way, for example, for pixel block p2 in the image, at least one two-dimensional Gaussian ellipse associated with it includes: two-dimensional Gaussian ellipse 1 and two-dimensional Gaussian ellipse 2, that is, for pixel block p2, two-dimensional Gaussian ellipse 1 and two-dimensional Gaussian ellipse 2 are grouped together.

[0073] After the rule grouping is completed, each two-dimensional Gaussian ellipse will be further sent to the corresponding cache (also called pixel block cache). One cache corresponds to one pixel block, so the cache can also be called pixel block cache.

[0074] From the above, it can be understood that: in the overlap detection unit, if a collision box of a Gaussian imprint overlaps with 10 pixel blocks, it will be copied into 10 copies and sent to 10 caches respectively.

[0075] In addition, during the overlap detection process, the preprocessing module will also record the depth value of the two-dimensional Gaussian ellipse with opacity higher than the set threshold for each pixel block to form a rough depth map (also called pixel block depth map) and send it to the corresponding cache. Among them, the depth of the two-dimensional Gaussian ellipse can be understood as the distance between the corresponding pixel block and the two-dimensional Gaussian ellipse. In specific implementation, the depth of the center pixel of the pixel block relative to the center of the imprint (that is, the center of the two-dimensional Gaussian ellipse) is usually taken as the depth of the two-dimensional Gaussian ellipse, but it is not limited to using other methods such as the nearest point method, the farthest point method, the average depth method, the weighted average depth method, the median depth method, etc. to determine a representative depth value for the two-dimensional Gaussian ellipse. And the reason for constructing the depth map here is: considering that in a complex scene, the front object will cover the back object. In order to reduce the processing of the obscured object, the depth map constructed can be used to roughly determine where the boundary of the object is.

[0076] For example, after detection, a pixel block P in the image i The at least one correspondingly associated two-dimensional Gaussian ellipse includes: two-dimensional Gaussian ellipse 1, two-dimensional Gaussian ellipse 2, ..., two-dimensional Gaussian ellipse n, wherein the opacity of two-dimensional Gaussian ellipse 2, two-dimensional Gaussian ellipse nj to two-dimensional Gaussian ellipse n is greater than or equal to the set threshold, then for this pixel block P i The depth values ​​of the two-dimensional Gaussian ellipse 2 and the two-dimensional Gaussian ellipse nj to the two-dimensional Gaussian ellipse n are recorded, thereby forming a map corresponding to the pixel block P i The associated depth map is output to the corresponding buffer.

[0077] It is necessary to add that: Figure 2 The preprocessing module shown in the output corresponding to the information in the cache may also include: the color of each two-dimensional Gaussian ellipse, the color of a two-dimensional Gaussian ellipsoid is calculated from the spherical harmonic coefficients of the two-dimensional Gaussian ellipse to be used in the subsequent pixel block rendering process.

[0078] In the above 104, at least one two-dimensional Gaussian ellipse may be sorted first, and then the sorting result may be used to trigger the execution of rendering a pixel block in the image associated with the at least one two-dimensional Gaussian ellipse using the at least one two-dimensional Gaussian ellipse. The sorting is implemented in conjunction with the corresponding depth map. Based on this, and in conjunction with the above description of the depth map, let the first pixel block be a pixel block in the image, then the "rendering the first pixel block in the image using at least one two-dimensional Gaussian ellipse associated with the first pixel block" contained in the above step 104 may include:

[0079] 1041. Determine, from the at least one two-dimensional Gaussian ellipse, a two-dimensional Gaussian ellipse whose features meet a preset condition; wherein the feature information of the two-dimensional Gaussian ellipse includes opacity, and the feature meeting the preset condition includes that the opacity is greater than or equal to a set threshold;

[0080] 1042. Record the depth value of the two-dimensional Gaussian ellipse whose features meet the preset conditions to form a depth map associated with the first pixel block;

[0081] 1043. Sort the at least one two-dimensional Gaussian ellipse according to the depth map to obtain a sorting result;

[0082] 1044. Render the first pixel block using the at least one two-dimensional Gaussian ellipse according to the sorting result.

[0083] For the implementation of the above 1041-1042, please refer to the related contents described in other embodiments. It should be supplemented here that the opacity of the two-dimensional Gaussian ellipse described in the above 1031 is determined according to the opacity of the corresponding three-dimensional Gaussian ellipse.

[0084] The sorting in the above 1043 is done by Figure 2 In the present application, this sorting module adopts a bucket-radix-CAS three-level hierarchical Gaussian sorting method to reduce unnecessary costs introduced by scene occlusion. For example, at least one two-dimensional Gaussian ellipse is subjected to a bucket-radix-CAS triple sorting tree according to the corresponding depth map, and two-dimensional Gaussian ellipses of different depths can be processed in blocks and terminated in time, thereby reducing the decrease in hardware utilization caused by complex spatial occlusion relationships. Specifically, as shown in Figure 2 The sorting process performed within the sorting module shown in FIG. 1 mainly includes the following:

[0085] First, the input two-dimensional Gaussian ellipses and their depth values ​​are sent to the bucket sorter, which roughly aggregates the two-dimensional Gaussian ellipses into several surface-centered sorting buckets (such as bucket 0 to bucket 3) and two additional sorting buckets (such as background bucket and foreground bucket) describing the foreground and background. The bucket sorter divides (i.e., allocates) each two-dimensional Gaussian ellipse into the corresponding bucket based on the depth value of each two-dimensional Gaussian ellipse, using the depth recorded in the corresponding depth map as the threshold. The bucket sorter uses this method of dividing the sorting buckets using the depth recorded in the depth map as the threshold, which can more accurately indicate the Gaussian distribution than the traditional predefined or randomly selected threshold. Moreover, since the surfaces of objects with similar features have similar Gaussian densities, the memory pressure of each bucket can also be evenly distributed.

[0086] Next, the buckets are further divided using a radix sort through a bucket splitter until the number of two-dimensional Gaussian ellipses in the bucket drops to, for example, less than 64. Radix sorting includes selecting sorting bits and bitwise partitioning. Selecting sorting bits means: based on a certain attribute of the two-dimensional Gaussian ellipse (such as depth value, opacity, etc.), a suitable sorting bit is selected for processing; if a floating point number is used, it can be converted into a binary representation and processed bit by bit. Bitwise partitioning means: using the idea of ​​radix sorting, the Gaussian imprint in the bucket is redistributed to new sub-buckets according to the value of the current bit. For example, if the current bit is the least significant bit, 10 sub-buckets (for decimal numbers) or more sub-buckets (for other radixes) can be created.

[0087] Finally, the CAS bitune sorter (for Figure 2 The CAS precise sorter shown in FIG. 1 performs precise depth sorting on the divided small-scale two-dimensional Gaussian ellipse groups, and sequentially sends the two-dimensional Gaussian ellipses to the input buffer of the rendering unit for subsequent rendering process. Among them, CAS (Compare-And-Swap) is an operation used to implement concurrent bitonic sorting (BitonicSort).

[0088] It should be noted that in the sorting algorithm, a bucket is a container for storing data. Each bucket has a certain interval range and can contain one or more elements. Figure 2 The buckets that appear in the description of the sorting module are essentially data sets or data groups.

[0089] Based on the above content, the above bucket is called an ellipse set (a data set for storing two-dimensional Gaussian ellipses). And accordingly, in an implementable solution, the above 1043 "sorting the at least one two-dimensional Gaussian ellipse according to the depth map to obtain a sorting result" includes:

[0090] 10431. According to the depth value of the at least one two-dimensional Gaussian ellipse, using the depth value included in the depth map as a threshold, assigning the at least one two-dimensional Gaussian ellipse to a plurality of first ellipse sets; wherein one two-dimensional Gaussian ellipse is in one ellipse set;

[0091] 10432. Perform segmentation on the multiple first ellipse sets to obtain multiple second ellipse sets;

[0092] 10433. Perform concurrent bitonic sorting on the multiple second ellipse sets (this is achieved by using a CAS bitonic sorter) to obtain a sorting result of the at least one two-dimensional Gaussian ellipse.

[0093] In the above 10432, the segmentation method includes radix sorting, and after an ellipse set is processed by radix sorting, the two-dimensional Gaussian ellipses in it will be redistributed into multiple new sub-ellipse sets. And, a specific implementation method of the above 10332 includes the following steps:

[0094] 104321. Determine a target first ellipse set to be segmented from the plurality of first ellipse sets according to the number of two-dimensional Gaussian ellipses included in each first ellipse set;

[0095] 104322. Based on the depth value of the two-dimensional Gaussian ellipse included in the target first ellipse set, segment the two-dimensional Gaussian ellipse included in the target first ellipse set by using a radix sorting method to obtain a plurality of sub-ellipse sets;

[0096] 104323. The plurality of second ellipse sets include the plurality of sub-ellipse sets and remaining first ellipse sets among the plurality of first ellipse sets except a target first ellipse set.

[0097] To facilitate understanding of the above 10431 to 10433, an example is given below for illustration.

[0098] Assume that the depth associated with the first pixel block p1 is Figure 1 contains the following depth values: depth value v1, depth value v2, depth value v3, wherein depth value v1 is greater than depth value v2, and depth value v2 is greater than depth value v3. Based on this, four first ellipse sets (empty ellipse sets, i.e., four empty buckets are set, each empty bucket is an empty ellipse set) can be set in advance, and the four first ellipse sets include: ellipse set 0 (foreground set), ellipse set 1, ellipse set 2, ellipse set 3 (for background set), and the depth value range corresponding to ellipse set 0 is [v1, +∞), the depth value range corresponding to ellipse set 1 is [v2, v1), the depth value range corresponding to ellipse set 2 is [v3, v2), and the depth value range corresponding to ellipse set 3 is (-∞, v3). And, assuming that the two-dimensional Gaussian ellipse associated with the first pixel block includes: two-dimensional Gaussian ellipse e1, two-dimensional Gaussian ellipse e2, two-dimensional Gaussian ellipse e3, ..., two-dimensional Gaussian ellipse en-1, two-dimensional Gaussian ellipse en, wherein the depth value associated with the first pixel block p1 is Figure 1 It is obtained by recording the depth value of the two-dimensional Gaussian ellipse with opacity less than the set threshold in the two-dimensional Gaussian ellipse associated with the first pixel block p1. Then: if the depth value of the two-dimensional Gaussian ellipse e1 is greater than v1, the two-dimensional Gaussian ellipse will be assigned to the ellipse set 0; if the depth value of the two-dimensional Gaussian ellipse e2 is greater than v3 and less than v2, the two-dimensional Gaussian ellipse will be assigned to the ellipse set 2, and so on. The other two-dimensional Gaussian ellipses from the two-dimensional Gaussian ellipse e3 to the two-dimensional Gaussian ellipse en associated with the first pixel block p1 can also be assigned to the corresponding ellipse sets in the ellipse set 0 to the ellipse set 3.

[0099] After the allocation is completed, a segmentation operation can be performed on the above-mentioned ellipse set 0 to ellipse set 3 to obtain multiple second ellipse sets. In specific implementation, the segmentation process can be performed on the first Gaussian ellipse only when it is determined that the number of two-dimensional Gaussian ellipses contained in a first ellipse set is higher than the preset number threshold T. For example, the number of two-dimensional Gaussian ellipses in ellipse set 2 is higher than the preset number threshold T, and further, radix sorting will be used to segment ellipse set 2. Specifically, several sub-ellipse sets (empty ellipse sets) can be set for ellipse set 2 in advance, and then, each two-dimensional Gaussian ellipse contained in ellipse set 2 can be allocated to the corresponding sub-ellipse set, but is not limited to using the units of the depth value as the base, etc. For example, ellipse set 2 is provided with 10 sub-ellipse sets, namely, sub-ellipse set 2_0, sub-ellipse set 2_1, sub-ellipse set 2_2, ..., and sub-ellipse set 2_9. If the unit digit of the depth value of a two-dimensional Gaussian ellipse contained in ellipse set 2 is 2, then the two-dimensional Gaussian ellipse is allocated to the sub-ellipse set 2_2. Similarly, other two-dimensional Gaussian ellipses in ellipse set 2 may also be allocated to corresponding sub-ellipse sets in advance. It should be additionally explained here that if the number of two-dimensional Gaussian ellipses in the sub-ellipse set is still higher than the preset number threshold, the radix sorting may continue to be used for segmentation.

[0100] Assume that after the segmentation operation, the multiple second ellipse sets finally obtained include ellipse set 0, ellipse set 1, several sub-ellipse sets corresponding to sub-ellipse set 2, and ellipse set 3. After that, the two-dimensional Gaussian ellipses in each second ellipse set can be sorted in parallel according to the size of the depth value. After the sorting is completed, these second ellipse sets are gradually merged and sorted to obtain the sorting result. For example, when merging and sorting step by step, because the depth value range corresponding to ellipse set 0 is larger than the depth value range corresponding to ellipse set 1, the ordered two-dimensional Gaussian ellipse queue contained in ellipse set 0 will be placed in front of the ordered micro-Gaussian ellipse queue contained in ellipse 1.

[0101] In the above 1044, see Figure 2 As shown, after obtaining the sorting result of at least one two-dimensional Gaussian ellipse associated with the first pixel block through step 1033, at least one two-dimensional Gaussian ellipse can be sent to the input buffer of the rendering module according to the sorting result. Specifically, the ellipse feature information (such as color, rotation feature, scaling feature, etc.) of the two-dimensional Gaussian ellipse is sent to the rendering module so that the rendering module can perform rendering.

[0102] The above rendering module includes multiple rendering unit LUTs. The rendering unit LUT is also called a rendering core. Each rendering unit has multiple auxiliary rendering blocks, which are also called lookup table auxiliary rendering units (LRUs). For example, each rendering unit contains 4*4 LRUs, which are interconnected into an LRU array. Each LRU is responsible for rendering a single pixel in a pixel block. When a batch of sorted two-dimensional Gaussian ellipses are loaded into the input buffer, the leftmost LRU in each row will obtain the corresponding transformed basis vectors i′, j′, Gaussian position vector Δ, opacity o and color, etc., and start the rendering process. The rendering unit runs in a full pipeline form. In each cycle, a new two-dimensional Gaussian ellipse is sent to the leftmost of the LRU array, and the two-dimensional Gaussian ellipse features input in the previous cycle will be propagated to the adjacent unit on the right until it reaches the rightmost of the array. This linear topology reduces the interconnection complexity and does not introduce global delay.

[0103] In specific implementation, the above LRU uses mixed data precision to perform rendering tasks: its input coordinate vector is represented by two 16-bit signed fixed-point numbers. The 11-bit integer part is used to describe the corresponding position to be searched in the one-dimensional Gaussian lookup table, and the 4-bit decimal part is used to prevent error accumulation and precision loss caused by repeated accumulation. Since other eigenvalues, such as transmittance T, opacity o, and color c rendered by a single two-dimensional Gaussian ellipse are all between [0:1], in the LRU, these variables are quantized to INT8 or INT12 precision depending on whether they need to be accumulated between different two-dimensional Gaussian ellipses. INT8 and INT12 refer to two different precision representations of integer data types. INT8 is an 8-bit integer type that can represent integer values ​​from -128 to 127. INT12 is a 12-bit integer type that can represent integer values ​​from -2048 to 2047.

[0104] During the rendering process, LRU mainly performs three tasks: α-calculation, α-blending and color rendering. In α-calculation, the normalized coordinates of the pixel center are calculated step by step through the linear combination of the feature vectors. The corresponding calculation formula is: Δ * (m,n) =Δ * base +m·i′+n·j′; where m and n are the IDs of the current pixel in the rendering pixel block (size is 16×16, for example), Δ * base is the position vector of the pixel block center relative to the center of the two-dimensional Gaussian ellipse after the intrinsic coordinate system transformation. In this step, the coordinates can be exceeded by the lookup table (such as Figure 2The two-dimensional Gaussian ellipse that does not cover the pixel point is identified and eliminated by using the long axis lookup table and short axis lookup table) range shown in the figure. After that, the rendering unit calculates the α value by multiplying the Gaussian opacity with the two one-dimensional Gaussian kernel LUT query results distributed along the long axis and short axis of the two-dimensional Gaussian ellipse, and compares the α value with the predefined threshold. Once the α value is less than the threshold, the Gaussian point is considered too transparent, and the remaining rendering process is skipped. Using the α value, the rendering unit can calculate and update the color and transmittance of the corresponding pixel until the transmittance reaches zero. In a simple non-quantized pipeline, this update process requires four cycles to complete the transmittance update of the same pixel, which forces a rendering unit to process four pixels in sequence, resulting in an increase in memory budget and hardware computing pressure. The present application breaks this data dependency by using an effective quantization calculation method, thereby reducing the size of the rendered pixel block processed at one time by approximately 75% without reducing performance.

[0105] Here are some additional explanations: Figure 2 The major axis lookup table and the minor axis lookup table shown in the figure are used to simplify the calculation of α. For example, in combination with the α calculation formula given in the other embodiments above: If the α calculation is performed normally, several matrix calculations need to be performed, and then addition calculations need to be performed. After the addition, exponential calculations need to be performed, etc. These calculations all need to be calculated using floating-point structures. Then this application simplifies this α calculation into In this case, since G (Gaussian kernel function) is a known operator, its size depends entirely on the input Δ * , so the size of G can be found through the lookup table, that is, it can be understood that a lookup table can be used to replace the corresponding G (Δ * ),in, and Respectively refers to Figure 1 The minor axis and major axis of the two-dimensional Gaussian ellipse shown in FIG. 1 can be replaced by the major axis lookup table accordingly. and replace it with a short axis lookup table This is to facilitate the direct quantitative calculation of the Gaussian kernel, thereby simplifying the α calculation. In addition, the early termination of LRU means: under the above quantization accuracy, if the color concentration is equal to 0, it means that it is completely transparent from the perspective. If an object is completely transparent, it will not affect the rendering result, so it can be skipped directly without subsequent processing and calculation. Similarly, the T register is the total transmittance of the light after passing through the two-dimensional Gaussian ellipse. It can be understood that it is the amount of light remaining from the ray emitted from the camera to the current position. If this remainder is equal to 0, it means that no light has been transmitted forward to the current position. If the light no longer transmits forward, it will not be affected by the two-dimensional Gaussian ellipsoid behind it, so it can be terminated early without subsequent processing and calculation.

[0106] Based on the above content, the present application also provides an image processor based on Gaussian splatting (GSNorm accelerator). Figure 2 FIG. 4 shows a schematic diagram of the structure of the image processor. Figure 2 As shown, the image processor includes: a pre-processing module and a rendering module.

[0107] A preprocessing module, used to determine the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses; using the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses, determine the pixel block information of an image covered by each of the multiple two-dimensional Gaussian ellipses; and according to the pixel block information covered by each of the multiple two-dimensional Gaussian ellipses, determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image;

[0108] A rendering module is used to render a pixel block associated with the at least one two-dimensional Gaussian ellipse in the image using the at least one two-dimensional Gaussian ellipse.

[0109] Furthermore, the above-mentioned preprocessing module includes: a transformation unit and an overlap detection unit. The transformation unit is used to perform a two-dimensional transformation on a plurality of three-dimensional Gaussian ellipses located in a visual cone to obtain the plurality of two-dimensional Gaussian ellipses; and determine the two-dimensional circles corresponding to the plurality of two-dimensional Gaussian ellipses; wherein the visual cone is used to describe a specified field of view. The overlap detection unit is used to determine the pixel block information of an image covered by each of the plurality of two-dimensional Gaussian ellipses using the two-dimensional circles corresponding to the plurality of two-dimensional Gaussian ellipses; and, further used to: determine a two-dimensional Gaussian ellipse whose features meet a preset condition from the at least one two-dimensional Gaussian ellipse associated with the first pixel block; wherein the feature information of the two-dimensional Gaussian ellipse includes opacity, and the feature meeting the preset condition includes opacity being greater than or equal to a set threshold; record the depth value of the two-dimensional Gaussian ellipse whose features meet the preset condition to form a depth map associated with the first pixel block; wherein the first pixel block is a pixel block in the image.

[0110] In addition, the preprocessing module may further include a rejection unit. For detailed descriptions of the units and functions included in the preprocessing module, please refer to the relevant contents in other embodiments, and no further details will be given here.

[0111] Furthermore, the image processor further includes: a sorting module. The sorting module is used to sort at least one two-dimensional Gaussian ellipse associated with the first pixel block according to the depth map associated with the first pixel block to obtain a sorting result; and according to the sorting result, send characteristic information of the at least one two-dimensional high-speed ellipse to the rendering module.

[0112] And, accordingly, the rendering module, when used to render the first pixel block in the image using at least one two-dimensional Gaussian ellipse associated with the first pixel block, is specifically used to: based on feature information of the at least one two-dimensional Gaussian ellipse, trigger the execution of rendering the first pixel block in the image using the at least one two-dimensional Gaussian ellipse.

[0113] In a specific implementation, the rendering module includes a plurality of rendering units, and a plurality of auxiliary rendering blocks are integrated on a rendering unit. The plurality of auxiliary rendering blocks are interconnected to form a rendering block array of size M*N. For example, a rendering array of size 4*4 can be formed, and in this case, a rendering unit includes 4*4 auxiliary rendering modules. The auxiliary rendering blocks are also called lookup table auxiliary rendering units (LRUs). The specific implementation of rendering performed by the rendering module can refer to the relevant contents in other embodiments, which will not be repeated here.

[0114] For detailed description of the functions of each module / unit in the image processor, please refer to the relevant contents in other embodiments.

[0115] The image processor provided by the present application is also called a Gaussian rendering accelerator (GSNorm). Therefore, the image processor provided by the present application is represented by GSNorm below.

[0116] In order to verify that the GSNorm provided by this application has good rendering performance, this application also evaluates the image rendering performance of GSNorm, where the evaluated image rendering performance includes: rendering quality, rendering speed, rendering time, etc. When evaluating the image rendering quality effect, since the GSNorm used in this application is rendered with floating-point inputs through online quantization and mixed INT8-INT12 precision, the first step is to evaluate the quality of rendered images in different scenarios through PSNR (Peak Signal-to-Noise Ratio).

[0117] Figure 4FIG. 2 shows a schematic diagram comparing the rendering quality of the GSNorm used in this application with that of a traditional image processor GPU. Figure 4 As shown in the figure, in the Train rendering scene, GSNorm achieves an average PSNR of 25.07, with negligible quality loss, compared to the 25.11 PSNR obtained by traditionally using full-precision FP32 operations on a graphics processing unit (GPU). FP32 refers to the Single-Precision Floating-Point format. Figure 4 The article further provides PSNR rendering results in other different scenarios such as Truck, Playroom, Drjohnson, etc.

[0118] In addition, to evaluate the performance, this application also compares GSNorm with the desktop RTX 3060 GPU and the consumer-grade edge Jetson Xavier NX. Specifically, this application first evaluates the overall performance improvement of the designed GSNorm. Figure 5 The rendering speed comparison diagram shown in the implementation experiment of this application shows that it takes about 16ms, or 67.8FPS, for the RTX 3060 to render a single image. However, the edge Jetson Xavier NX is about 7 times slower, that is, a single rendering requires 100ms and a throughput of 9.8FPS. In contrast, the GSNorm accelerator proposed in this application can achieve an average of 9.5 times the speed increase compared to Jetson NX in the same rendering task, and a 1.2 times the speed increase compared to the RTX3060GPU. In contrast, GSNorm performs better on dense data sets such as Drjohnson or Playground. This is mainly because the pipeline rendering process of the GSNorm accelerator reduces the storage pressure of the rendering unit and avoids repeated off-chip memory accesses in the GPU rendering process, thereby achieving higher performance improvements.

[0119] Furthermore, this application also analyzes the energy consumption of GSNorm based on the P&R layout results. The full name of P&R is Place&Route. Based on the logic synthesis results of Cadence Genus (a logic synthesis tool) and the layout and routing results of Cadence Innovus (a comprehensive physical design tool, mainly used for the layout and routing of digital integrated circuits (ICs)), the GSNorm accelerator consumes 284mW at an operating frequency of 500MHz, of which the LUT rendering unit contributes about 80% of the total energy consumption. For single-frame image rendering, the GSNorm accelerator consumes about 3.01mJ of energy in 10.4ms. In comparison, the estimated energy consumption of RTX 3060 and Jetson NX is 2.11J per frame and 3.0J per frame, respectively, which means that the energy efficiency of GSNorm is 3 orders of magnitude higher than that of general-purpose computing platforms.

[0120] Furthermore, the present application also evaluates the effectiveness of two-dimensional Gaussian ellipse in reducing rendering calculations through intrinsic transformation (i.e., transformation to the intrinsic coordinate system), and analyzes the changes in rendering time under different optimization strategies. Figure 7 The schematic diagram of the effect of different optimization strategies on rendering time shows that the normalized pipeline improves rendering efficiency by 50% on average with negligible increase in preprocessing time (i.e., 5%), which is very consistent with the computational distribution estimated by this application. This application also notes that in complex scenes (such as Playroom or Drjohnson), the speed improvement is more obvious, which is consistent with the performance results of this application. This further proves the effectiveness of the rendering accelerator designed by this application.

[0121] The embodiment of the present application also provides a structural schematic diagram of an electronic device. The electronic device includes: a memory and the image processor provided by the present application. The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. Specifically,

[0122] The above-mentioned memory is used to store programs;

[0123] The above-mentioned image processor is coupled to the memory 7 and is used to execute the program stored in the memory for the steps or functions in each method provided in each embodiment of the present application.

[0124] Furthermore, the electronic device also includes: communication components, power components, audio components and other components. The components included in the electronic device described above are only schematically shown as part of the components, and do not mean that the electronic device only includes the components described above.

[0125] In specific implementation, the electronic device can be a smart phone, a laptop computer, a desktop EEG, a tablet, or other devices with logic processing capabilities.

[0126] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a computer, the method steps provided in the above embodiments can be implemented.

[0127] An embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the processor is enabled to implement the method steps or functions provided in the above embodiments.

[0128] Through the description of the above implementation modes, those skilled in the art can clearly understand that each implementation mode can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on such an understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An image processing method based on Gaussian splashing, characterized in that: include: Determine the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses; Using the two-dimensional circles corresponding to the two-dimensional Gaussian ellipses, respectively, to determine pixel block information of an image covered by the two-dimensional Gaussian ellipses; Determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image according to pixel block information respectively covered by the multiple two-dimensional Gaussian ellipses; Using the at least one two-dimensional Gaussian ellipse, a pixel block in the image associated with the at least one two-dimensional Gaussian ellipse is rendered.

2. The method according to claim 1, characterized in that The first two-dimensional Gaussian ellipse is one of the plurality of two-dimensional Gaussian ellipses; as well as, Determining pixel block information of an image covered by the first two-dimensional Gaussian ellipse by using the two-dimensional circle corresponding to the first two-dimensional Gaussian ellipse includes: Determine a circumscribed rectangle of the two-dimensional circle; The pixel block information of an image covered by the first two-dimensional Gaussian ellipse is determined by using the circumscribed rectangle of the two-dimensional circle.

3. The method according to claim 1 or 2, characterized in that: The first pixel block is a pixel block in the image; as well as Rendering a first pixel block in the image using at least one two-dimensional Gaussian ellipse associated with the first pixel block includes: Determine, from the at least one two-dimensional Gaussian ellipse, a two-dimensional Gaussian ellipse whose features meet a preset condition; wherein the feature information of the two-dimensional Gaussian ellipse includes opacity, and the feature meeting the preset condition includes that the opacity is greater than or equal to a set threshold; Recording depth values ​​of a two-dimensional Gaussian ellipse whose features meet a preset condition to form a depth map associated with the first pixel block; sorting the at least one two-dimensional Gaussian ellipse according to the depth map to obtain a sorting result; According to the sorting result, the first pixel block is rendered using the at least one two-dimensional Gaussian ellipse.

4. The method according to claim 3, characterized in that Sorting the at least one two-dimensional Gaussian ellipse according to the depth map to obtain a sorting result, including: According to the depth value of the at least one two-dimensional Gaussian ellipse, the at least one two-dimensional Gaussian ellipse is assigned to a plurality of first ellipse sets using the depth value included in the depth map as a threshold value; wherein one two-dimensional Gaussian ellipse is in one ellipse set; Performing a segmentation operation on the plurality of first ellipse sets to obtain a plurality of second ellipse sets; Performing concurrent bitonic sorting on the plurality of second ellipse sets to obtain a sorting result of the at least one two-dimensional Gaussian ellipse; The segmentation method includes radix sorting. After an ellipse set is processed by radix sorting, the two-dimensional Gaussian ellipses in the ellipse set will be redistributed into multiple new sub-ellipse sets.

5. The method according to claim 4, characterized in that Performing a segmentation operation on the plurality of first ellipse sets to obtain a plurality of second ellipse sets includes: Determining a target first ellipse set to be segmented among the plurality of first ellipse sets according to the number of two-dimensional Gaussian ellipses contained in each first ellipse set; Based on the depth value of the two-dimensional Gaussian ellipse contained in the target first ellipse set, the two-dimensional Gaussian ellipse contained in the target first ellipse set is segmented by using a radix sorting method to obtain a plurality of sub-ellipse sets; The plurality of second ellipse sets include the plurality of sub-ellipse sets and remaining first ellipse sets among the plurality of first ellipse sets except a target first ellipse set.

6. An image processor based on Gaussian splashing, characterized in that: include: A preprocessing module, used to determine the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses; using the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses, determine the pixel block information of an image covered by each of the multiple two-dimensional Gaussian ellipses; and according to the pixel block information covered by each of the multiple two-dimensional Gaussian ellipses, determine at least one two-dimensional Gaussian ellipse associated with each pixel block in the image; A rendering module is used to render a pixel block associated with the at least one two-dimensional Gaussian ellipse in the image using the at least one two-dimensional Gaussian ellipse.

7. The image processor according to claim 6, characterized in that: The preprocessing module includes: a transformation unit and an overlap detection unit; The transformation unit is used to perform a two-dimensional transformation on a plurality of three-dimensional Gaussian ellipses located within a viewing cone to obtain the plurality of two-dimensional Gaussian ellipses; and determine a two-dimensional circle corresponding to each of the plurality of two-dimensional Gaussian ellipses; wherein the viewing cone is used to describe a specified field of view; An overlap detection unit is used to determine the pixel block information of an image covered by each of the multiple two-dimensional Gaussian ellipses using the two-dimensional circles corresponding to each of the multiple two-dimensional Gaussian ellipses; and is also used to: determine a two-dimensional Gaussian ellipse whose features meet preset conditions from the at least one two-dimensional Gaussian ellipse associated with the first pixel block; wherein the feature information of the two-dimensional Gaussian ellipse includes opacity, and the feature meeting the preset condition includes opacity being greater than or equal to a set threshold; record the depth value of the two-dimensional Gaussian ellipse whose features meet the preset condition to form a depth map associated with the first pixel block; wherein the first pixel block is a pixel block in the image.

8. The image processor according to claim 7, characterized in that: Also includes: Sorting module; The sorting module is used to sort at least one two-dimensional Gaussian ellipse associated with the first pixel block according to the depth map associated with the first pixel block to obtain a sorting result; according to the sorting result, send characteristic information of the at least one two-dimensional high-speed ellipse to the rendering module; The rendering module, when used to render the first pixel block in the image using at least one two-dimensional Gaussian ellipse associated with the first pixel block, is specifically used to: based on feature information of the at least one two-dimensional Gaussian ellipse, trigger the execution of rendering the first pixel block in the image using the at least one two-dimensional Gaussian ellipse.

9. An electronic device, characterized in that: include: A memory and an image processor as claimed in any one of claims 6 to 7; wherein: The memory is used to store programs; The image processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the image processing method according to any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the steps in the image processing method according to any one of claims 1 to 5 can be implemented.

Citation Information

Cited By

  • Dynamic deletion method and system for three-dimensional Gaussian spattering

    CN120635289A