Primitive encoding method, apparatus, device, storage medium and program product

By merging and storing shared vertex data as shared vertices in the structure, the problem of high video memory usage in primitive encoding is solved, improving the efficiency of computations such as ray tracing and video memory utilization.

CN121120809BActive Publication Date: 2026-02-10MOORE THREADS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511653519.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-10
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

In existing technologies, primitive encoding has a high memory usage rate when constructing BVH, which affects the efficiency of rendering operations such as ray tracing.

Method used

By encoding at least two primitives to be encoded based on the first encoding method, the shared vertex data is merged and stored as a shared vertex in the structure, reducing the storage of duplicate vertices, realizing compact encoding of continuous primitives, and reducing the storage space occupancy rate.

Benefits of technology

It effectively reduces storage overhead and improves the coding efficiency and memory utilization of computations such as ray tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120809B_ABST
    Figure CN121120809B_ABST
Patent Text Reader

Abstract

This application discloses a primitive encoding method, apparatus, device, storage medium, and program product. The method includes: encoding at least two primitives to be encoded based on a first encoding method to obtain a structure of the first encoding method; the structure of the first encoding method includes a shared vertex stored once, the shared vertex being shared by at least two of the primitives to be encoded, and the primitive indices of the two primitives to be encoded being consecutive in the geometry. This application can reduce the video memory usage during BVH construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of graphics processor technology, and in particular to a primitive encoding method, apparatus, device, storage medium, and computer program product. Background Technology

[0002] In the field of computer graphics, when calculating global illumination for large-scale scenes, the Bounding Volume Hierarchy (BVH) data volume is extremely large, consuming a significant amount of video memory. Primitive encoding is a crucial step in constructing the BVH, directly impacting the efficiency of rendering operations such as ray tracing. Current primitive encoding technologies exhibit high video memory usage. Summary of the Invention

[0003] In view of this, embodiments of this application provide at least one primitive encoding method, apparatus, device, storage medium, and program product.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] On one hand, embodiments of this application provide a primitive encoding method, the method comprising: encoding at least two primitives to be encoded based on a first encoding method to obtain a structure of the first encoding method; the structure of the first encoding method includes a shared vertex stored once, the shared vertex being shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded being consecutive in the geometry.

[0006] On the other hand, embodiments of this application provide a primitive encoding device, the device comprising: an encoding module, configured to encode at least two primitives to be encoded based on a first encoding method to obtain a structure of the first encoding method; the structure of the first encoding method includes a shared vertex stored once, the shared vertex being shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded being consecutive in the geometry.

[0007] In another aspect, embodiments of this application provide a computer device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.

[0008] In another aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.

[0009] In another aspect, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement some or all of the steps in the above-described method.

[0010] Based on the above embodiments disclosed in this application, by encoding at least two consecutive primitive indices of primitives to be encoded using the first encoding method, shared vertex data can be merged and stored as a common vertex in the structure, reducing the amount of redundant data caused by repeated vertex storage. At the same time, by storing the common vertex only once in the structure and recording the starting primitive index, compact encoding of consecutive primitives can be achieved, reducing the storage space occupancy rate after encoding a single primitive, effectively reducing storage overhead, and improving the encoding efficiency and video memory utilization of calculations such as ray tracing.

[0011] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this application. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0013] Figure 1 A schematic diagram of a BVH construction process provided in an embodiment of this application;

[0014] Figure 2 A schematic diagram of a geometric mesh provided in an embodiment of this application;

[0015] Figure 3 This is a schematic diagram of another geometric mesh provided in an embodiment of this application;

[0016] Figure 4 A schematic diagram illustrating the implementation process of a primitive encoding method provided in this application embodiment;

[0017] Figure 5 A schematic diagram of a structure provided in an embodiment of this application. Figure 1 ;

[0018] Figure 6 A schematic diagram of a structure provided in an embodiment of this application. Figure 2 ;

[0019] Figure 7 A schematic diagram of a structure provided in an embodiment of this application. Figure 3 ;

[0020] Figure 8 An illustration of index information for a structure provided in this application embodiment. Figure 1 ;

[0021] Figure 9 An illustration of index information for a structure provided in this application embodiment. Figure 2 ;

[0022] Figure 10 An illustration of index information for a structure provided in this application embodiment. Figure 3 ;

[0023] Figure 11 This application provides a schematic diagram illustrating the correspondence between threads and primitives in an embodiment of the present application.

[0024] Figure 12 A schematic diagram of a shared memory provided in an embodiment of this application;

[0025] Figure 13 A schematic diagram of an encoding process provided in this application embodiment. Figure 1 ;

[0026] Figure 14 A schematic diagram of an encoding process provided in this application embodiment. Figure 2 ;

[0027] Figure 15 A schematic diagram of an encoding process provided in this application embodiment. Figure 3 ;

[0028] Figure 16 This is a schematic diagram of the composition structure of a primitive encoding device provided in an embodiment of this application;

[0029] Figure 17 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0031] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.

[0033] Besides the number of rays in parallel traversal, the speed of ray traversal, and the calculation of intersections between rays, bounding boxes, and triangles, hardware-accelerated ray traversal technology may also encounter bottlenecks related to video memory. This is especially true when calculating global illumination for large-scale scenes, where the Bounding Volume Hierarchy (BVH) data volume can be extremely large, consuming significant amounts of video memory. A BVH is a tree-like structure used to accelerate scene traversal. Intermediate nodes contain bounding boxes and pointers to child or leaf nodes. Leaf nodes store one or more geometric primitives. The BVH construction process is as follows... Figure 1 As shown, in stage 1, the geometric primitives of the entire scene are divided into two bounding boxes: one containing geometric primitives of type 11 and the other containing geometric primitives of type 12. Correspondingly, as shown in the tree diagram of stage 1, the root node of the entire scene is divided into two child nodes, labeled as child nodes of type 11 and child nodes of type 12, respectively. Next, in stage 2, the bounding boxes of geometric primitives of type 11 are further divided into bounding boxes of type 21 and type 22. Correspondingly, as shown in the tree diagram of stage 2, the nodes originally labeled as type 11 are further divided into two child nodes, labeled as child nodes of type 21 and child nodes of type 22, respectively. Finally, in stage 3, the bounding boxes of geometric primitives of type 21 are further divided into bounding boxes of type 31 and type 32. Correspondingly, as shown in the tree diagram of stage 4, the nodes originally labeled as type 21 are further divided into two child nodes, labeled as child nodes of type 31 and child nodes of type 32, respectively.

[0034] The above scheme uses a triangle list to store geometric primitives in BVH leaf nodes. This storage structure is relatively loose and can lead to wasted video memory. In fact, within a geometric mesh, the same vertex can be shared by multiple triangles, such as... Figure 2 and Figure 3 As shown, this application proposes to encode multiple triangles into the same BVH leaf node using an IndexedTriangle List, making the BVH structure more compact and reducing memory pressure.

[0035] This application provides a primitive encoding method, which can be executed by a processor of a computer device. The computer device refers to a device with data processing capabilities, such as a server, laptop, tablet, desktop computer, smart TV, set-top box, or mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device).

[0036] This application provides a primitive encoding method, which includes: encoding at least two primitives to be encoded based on a first encoding method to obtain a structure of the first encoding method; the structure of the first encoding method includes a shared vertex stored once, the shared vertex being shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded being consecutive in the geometry.

[0037] Please see Figure 4 A schematic diagram illustrating the implementation process of a primitive encoding method is provided, as follows: Figure 4 As shown, the method includes the following steps S401 and S402:

[0038] Step S401: Encode at least two primitives to be encoded based on the first encoding method.

[0039] Step S402: After encoding, the structure of the first encoding method is obtained; the structure of the first encoding method includes a shared vertex stored once, the shared vertex is shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded are consecutive in the geometry.

[0040] The first encoding method is used to encode at least two primitives to be encoded. The structure of the first encoding method includes a shared vertex that is stored once. The shared vertex is shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded are consecutive in the geometry.

[0041] In some embodiments, the primitives to be stored are basic graphic units constituting the geometry (such as triangles), which need to be stored in the leaf nodes of the acceleration structure (BVH) through an encoding process to support computations such as ray tracing. The index values ​​of multiple primitives to be encoded in the geometry are arranged in ascending order, and adjacent primitives to be encoded have a continuous logical positional relationship during storage or processing.

[0042] The first encoding method is a primitive encoding algorithm provided in this application embodiment, used to encode at least two primitives with consecutive primitive indices in a geometry. It achieves vertex data reuse by storing shared vertices (vertices shared by at least two primitives) once, reducing redundant storage. In some embodiments, the structure corresponding to this first encoding method includes shared vertex data and a starting primitive index, used for compact encoding of consecutive primitives.

[0043] Based on the above embodiments disclosed in this application, by encoding at least two consecutive primitive indices of primitives to be encoded using the first encoding method, shared vertex data can be merged and stored as a common vertex in the structure, reducing the amount of redundant data caused by repeated vertex storage. At the same time, by storing the common vertex only once in the structure and recording the starting primitive index, compact encoding of consecutive primitives can be achieved, reducing the storage space occupancy rate after encoding a single primitive, effectively reducing storage overhead, and improving the encoding efficiency and video memory utilization of calculations such as ray tracing.

[0044] In some embodiments, the structure includes encoding information, index information, and vertex information; the encoding information is used to indicate the number of primitives and the encoding method of the primitives to be encoded in the primitive node; the vertex information includes the vertex data of each vertex in the primitive to be encoded.

[0045] In the structure of the first encoding method, the index information includes a first primitive index and vertex indices of each vertex in the primitive to be encoded; the first primitive index is the smallest or largest primitive index among the at least two primitives to be encoded; the vertex index is used to locate the corresponding vertex data in the vertex information.

[0046] In some embodiments, in the structure of the second encoding method, the index information includes at least one second primitive index, the second primitive index corresponds one-to-one with the primitive to be encoded, and the position of the second primitive index of the primitive to be encoded in the index information corresponds to the position of the vertex data of the primitive to be encoded in the vertex data.

[0047] Please see Figure 5 It shows a schematic diagram of a structure, such as Figure 5 As shown, the structure 40 may include encoding information 41, index information 42, and vertex information 43.

[0048] The encoding information 41 is used to indicate the number of elements and the encoding method of the elements to be encoded in the element node. For example, the encoding information 41 may include a first encoding parameter and a second encoding parameter. When the first encoding parameter is set to a first value, it indicates that the structure 40 is a structure generated by a first encoding method; when the first encoding parameter is set to a second value, it indicates that the structure 40 is a structure generated by a second encoding method. The second encoding parameter can be used to indicate the number of elements to be encoded in the element node.

[0049] In some embodiments, the first encoding parameter is represented by a single binary bit to distinguish the encoding method of the structure: when the bit is a first preset value (e.g., 0), it indicates that the structure is generated using the first encoding method; when the bit is a second preset value (e.g., 1), it indicates that the structure is generated using the second encoding method. The number of bits or the range of values ​​of the second encoding parameter are dynamically adjusted based on the maximum number of primitives supported by the structure (primitive number threshold). For example, when the structure can store a maximum of N primitives, the second encoding parameter needs to use log2(N+1) bits (rounded up) to cover the value range of 0 to N, thereby accurately recording the actual number of primitives to be encoded.

[0050] For example, assuming the maximum number of primitives supported by the structure (primitive number threshold) is 9, then the second encoding parameter needs to use 4 bits (because 2 4 =16, covering the value range of 0-15). In this case, the structure's encoding information includes: First encoding parameter: 1 bit, 0 indicates the first encoding method, 1 indicates the second encoding method; Second encoding parameter: 4 bits, used to record the number of primitives to be encoded. If the structure stores 5 consecutive primitives with shared vertices using the first encoding method, then the first encoding parameter is set to 0, and the second encoding parameter is set to 5 (binary 0101); if it stores 3 independent primitives using the second encoding method, then the first encoding parameter is set to 1, and the second encoding parameter is set to 3 (binary 0011). Through this design, the encoding information can flexibly adapt to primitive groups of different sizes, while controlling the storage space occupied by the parameters.

[0051] The encoding information 42 is used at least to indicate the primitive index of one or more primitives to be encoded in the primitive node.

[0052] The vertex information 43 is used to indicate the vertex data of all vertices corresponding to the primitives to be encoded in the primitive node. In the structure of the first encoding method, the vertex data of one vertex is shared by at least two primitives to be encoded. In the structure of the second encoding method, the vertex data of each vertex corresponds to only one primitive to be encoded.

[0053] For example, suppose the geometry contains 8 triangular primitives (indices 0-7), where primitives 0-3 are indexed consecutively and their vertices can be shared. For example, primitive 0 contains vertices A, B, and C; primitive 1 contains vertices B, C, and D; primitive 2 contains vertices C, D, and E; and primitive 3 contains vertices D, E, and F. In this case, vertices B, C, D, and E are shared by at least two primitives. Primitives 4-7 are indexed consecutively but their vertices are not shared. For example, primitive 4 contains vertices G, H, and I; primitive 5 contains vertices J, K, and L; primitive 6 contains vertices M, N, and O; and primitive 7 contains vertices P, Q, and R. All vertices are used by only one primitive. When using the first encoding method for primitives 0-3, the vertex information stores the shared vertices A, B, C, D, E, and F (stored only once); when using the second encoding method for primitives 4-7, the vertex information stores the independent vertex data of each primitive (G, H, and I correspond to primitive 4, J, K, and L correspond to primitive 5, and so on), and the data of each vertex corresponds to only one primitive.

[0054] In some embodiments, in the structure of the first encoding method, the index information includes a first primitive index and a vertex index: the first primitive index is the minimum or maximum primitive index in a continuous primitive group, used to mark the start or end position of the primitive group, so as to facilitate quick location of the logical range of continuous primitives; the vertex index is the position identifier in the vertex information, and each vertex index corresponds to a vertex data in the vertex information. Vertex data shared by at least two primitives can be reused through the vertex index, reducing redundant storage.

[0055] For example, please refer to Figure 6 Suppose the geometry contains three consecutive primitives (indices 0-2), where, for example, primitive 0 contains vertices A, B, and C; primitive 1 contains vertices B, C, and D; and primitive 2 contains vertices C, D, and E. Using the first encoding method, in the structure's index information, the first primitive index is 0 (minimum index) or 2 (maximum index). The vertex indices of primitive 0 are 0, 1, and 2, corresponding to vertices A, B, and C in the vertex information, respectively; the vertex indices of primitive 1 are 1, 2, and 3, corresponding to vertices B, C, and D in the vertex information, respectively, and so on.

[0056] In some embodiments, the structure of the second encoding method includes at least one second primitive index in the index information. Each second primitive index corresponds one-to-one with the primitive to be encoded, and the arrangement order of the second primitive indexes of the primitive to be encoded in the index information strictly corresponds to the storage position of its vertex data in the vertex information (i.e., the i-th second primitive index in the index information corresponds to the i-th group of vertex data in the vertex information). Through this positional association, the corresponding vertex data can be quickly located directly according to the index position of the second primitive index without additional searching. This is suitable for scenarios where primitive indexes are not continuous or vertices are not shared, ensuring access efficiency when data is stored independently.

[0057] For example, suppose the geometry contains three primitives (indices 0-2) that do not satisfy the first encoding condition, where the vertex data of primitive 0 is A, B, and C, the vertex data of primitive 1 is D, E, and F, and the vertex data of primitive 2 is G, H, and I. When using the second encoding method, the structure's index information includes second primitive indices 0, 1, and 2 (arranged in order), and the vertex information stores three sets of vertex data in order (the first set is A, B, and C; the second set is D, E, and F; and the third set is G, H, and I). In this case, the first second primitive index 0 in the index information corresponds to the first set of vertex data (A, B, and C) in the vertex information, the second second primitive index 1 corresponds to the second set of vertex data (D, E, and F), and so on. The corresponding vertex data can be directly located through the index position, enabling fast access to independent primitive data.

[0058] In some embodiments of this application, the method further includes: determining at least two primitives to be encoded from a plurality of primitives to be stored that satisfy a first encoding method. It is understood that the current embodiments can be combined with any of the foregoing embodiments.

[0059] The conditions of the first encoding method that at least two primitives to be encoded satisfy can include at least one of the following: (1) Primitive indices are continuous, that is, the index values ​​of the primitives to be encoded in the geometry must be arranged in order (e.g., indices 0, 1, 2, ..., n). (2) Vertices are shareable, that is, multiple primitives to be encoded need to share at least one vertex (e.g., three triangle primitives with indices 0, 1, and 2 may share vertices A, B, and C). (3) The number of primitives is ≥ 2, that is, at least two consecutive primitives that can share vertices are required to trigger the first encoding method. If a single primitive cannot reuse vertex data, the second encoding method must be used instead. (4) Multiple primitives to be encoded satisfy the encoding amount threshold of the first encoding method, that is, the encoded structure is sufficient to store all the data of multiple primitives to be encoded.

[0060] For example, suppose there are 12 primitives to be stored (index 0-11), where primitives 0-7 have consecutive indices (0,1,2,...,7) and each pair of adjacent primitives shares at least one vertex (e.g., primitive 0 shares vertex A with primitive 1, primitive 1 shares vertex B with primitive 2, primitive 2 shares vertex C with primitive 3, and so on), satisfying the conditions of the first encoding method; primitives 8-11 have consecutive indices (8,9,10,11), but no two primitives share a vertex (e.g., each primitive's vertex is independent), not satisfying the conditions of the first encoding method. In the above embodiment, primitive groups that meet the conditions of the first encoding method are first filtered: if primitives 0-7 are found to have consecutive indices and each pair of adjacent primitives shares at least one vertex (e.g., vertices A, B, C, etc.), satisfying the conditions of the first encoding method, then primitives 0-7 can be encoded as 1 TriangleNode, marked as encoded, and the remaining primitives are 8-11. Next, the remaining primitives are processed: Primitives 8-11 are detected to be consecutive but without shared vertices, meaning that all remaining primitives to be stored do not meet the conditions of the first encoding method. The remaining primitives to be stored are then encoded using the second encoding method, meaning that each primitive to be stored independently stores 3 vertices.

[0061] Based on the above embodiments disclosed in this application, by dynamically selecting continuous groups of primitives that meet the conditions of the first encoding method and share at least one vertex from multiple primitives to be stored for encoding, shared vertex data can be reused (stored only once) and redundant storage can be reduced.

[0062] In some embodiments, determining at least two primitives to be encoded that satisfy the first encoding method among a plurality of primitives to be stored includes: allocating a corresponding thread for each primitive to be stored;

[0063] For each thread, the graph data of the element to be stored is obtained through the thread, and the graph data of at least one other element is obtained based on the order of the element index.

[0064] The target thread with the largest number of coded primitives and the at least one other primitive are determined as the at least two primitives to be encoded; the number of coded primitives is the number of primitives acquired, including the primitives to be stored and the at least one other primitive.

[0065] Wherein, the primitive to be stored acquired by the thread and the at least one other primitive satisfy the encoding amount threshold of the first encoding method.

[0066] In some embodiments, threads are dynamically created or allocated from a thread pool based on the number of primitives to be stored, ensuring that each thread corresponds one-to-one with a primitive to be stored; after the thread allocation is completed, each thread independently performs subsequent primitive data acquisition and encoding calculations.

[0067] For example, if there are 1000 triangle primitives to be stored, a thread can be allocated to each triangle, creating a total of 1000 threads. Each thread is responsible for processing the corresponding triangle primitive and independently performs subsequent data acquisition and encoding operations.

[0068] In some embodiments, the order of the primitive indices to be stored is the same as the order of the thread indices of the corresponding threads. For example, primitives 0 to n are assigned to threads 0 to n, respectively.

[0069] Here, to determine which primitives can be combined using the first encoding method, it is necessary to analyze the correlation between the primitives to be stored from aspects such as whether the indices are continuous and whether the vertices are shared.

[0070] The aforementioned enable quantity can be the number of primitives that the current thread can encode using the first encoding method, and must meet the encoding quantity threshold of the first encoding method. This encoding quantity threshold is used to control that the primitives to be stored and at least one other primitive acquired by the thread will not exceed the storage limit of the structure, for example, it will not exceed the maximum number of primitives that the structure can store, and / or, it will not exceed the maximum number of vertices that the structure can store.

[0071] In some embodiments, each thread first acquires the metadata of the primitives it is responsible for storing; then, based on the order of the primitive indices (e.g., ascending or descending), it acquires the metadata of each other primitive, starting from the first adjacent primitive; after acquiring each other primitive, it immediately checks whether the currently acquired primitive group (including its own primitive and the acquired other primitives) meets the encoding threshold of the first encoding method (e.g., not exceeding the maximum number of primitives or vertices that the structure can store); if it meets the threshold, it continues to acquire the next adjacent primitive and repeats the check; if it does not meet the threshold, it stops acquiring, records the number of primitives that currently meet the conditions as the coded quantity, and records the maximum or minimum primitive index as the first primitive index.

[0072] For example, suppose thread A is responsible for primitive X (primitive index 5, vertex data A, B, C). The structure requires storing a maximum of 3 primitives and no more than 5 vertices. Thread A first retrieves primitive Y (index 6, vertex data B, C, D) in ascending order of index. It finds that X and Y have consecutive indices, vertices are shared (B and C are shared), and the number of primitives (2) and vertices (A, B, C, D) do not exceed the threshold, thus satisfying the conditions. It then retrieves primitive Z (index 7, vertex data C, D, E). At this point, the number of primitives is 3 and the number of vertices (A, B, C, D, E) is 5, still satisfying the threshold. It then tries to retrieve primitive W (index 8, vertex data D, E, F). At this point, the number of primitives is 4 (exceeding the maximum number of primitives in the structure, 3), or the number of vertices is 6 (exceeding the maximum number of vertices, 5), which does not satisfy the threshold, so it stops retrieving primitives. Finally, the number of primitives that can be encoded is 3 (primitives X, Y, Z), with the first primitive index being 5 or 7.

[0073] It is important to understand that the above steps are executed in parallel by various threads. Since different threads are assigned to process different primitives to be stored, the parallel execution of multiple threads can fully cover all primitives to be stored and their adjacent primitives, avoiding the omission of possible encoding combinations. Each thread independently obtains and checks primitive data one by one, which can dynamically adapt to the storage limitations of the structure and accurately determine the number of coded elements. Parallel processing can significantly reduce the overall computation time and improve the utilization of computing resources.

[0074] In some embodiments, the encoding threshold includes a primitive number threshold and a vertex number threshold; the step of obtaining the primitive data of the primitive to be stored through the thread, and obtaining the primitive data of at least one other primitive based on the primitive index order, includes: obtaining the primitive data of the primitive to be stored through the thread, and updating the number of primitives to be encoded and the number of vertices to be encoded of the thread; obtaining the other primitives sequentially based on the primitive index order, and continuing to update the number of primitives to be encoded and the number of vertices to be encoded based on the obtained other primitives, until the stopping condition is met.

[0075] The stopping conditions include at least one of the following:

[0076] After obtaining the next other graphic element, the number of updated graphic elements to be encoded does not meet the graphic element number threshold.

[0077] After obtaining the next other primitive, the updated number of vertices to be encoded does not meet the vertex number threshold.

[0078] Here, the primitive number threshold is the maximum number of primitives that can be stored in the structure of the first encoding method, used to limit the number of primitives in a single structure to prevent exceeding the storage limit; the vertex number threshold is the maximum number of vertices that can be stored in the structure of the first encoding method, used to limit the number of vertices in a single structure to prevent exceeding the storage limit; the number of primitives to be encoded is the number of primitives that the thread has currently acquired, used to dynamically track the cumulative number of primitives in the current structure; the number of vertices to be encoded is the number of vertices that the thread has currently acquired, used to dynamically track the cumulative number of vertices in the current structure.

[0079] In some embodiments, all threads perform the following operations in parallel: First, each thread obtains the metadata of the primitive to be stored that it is responsible for, and initializes the number of primitives to be encoded to 1 and the number of vertices to be encoded to the number of vertices of the primitive to be stored.

[0080] For example, suppose thread A is responsible for primitive X (primitive index is 5, vertex data is A, B, C). After obtaining primitive X, thread A can update the number of primitives to be encoded to 1 and the number of vertices to be encoded to 3.

[0081] In some embodiments, the method further includes: in response to satisfying the stopping condition, determining the number of primitives to be encoded in the thread as the number of coded primitives in the thread.

[0082] In some embodiments, other primitives are acquired sequentially according to their primitive indices. For each acquired primitive, its vertex count is immediately added to the number of vertices to be encoded, and the number of primitives to be encoded is incremented by 1. After each update, a stop condition is immediately checked: if the number of updated primitives exceeds a preset threshold, or the number of vertices exceeds a preset threshold, acquisition is terminated, and the number of primitives to be encoded before the last update is taken as the number of coded primitives.

[0083] Here, if the updated number of primitives to be encoded does not meet the primitive number threshold, it can be that the updated number of primitives to be encoded is greater than the primitive number threshold; if the updated number of vertices to be encoded does not meet the vertex number threshold, it can be that the updated number of vertices to be encoded is greater than the vertex number threshold.

[0084] For example, continuing the above example, assume the structure requires storing a maximum of 4 primitives and no more than 5 vertices. After obtaining primitive X, thread A updates the number of primitives to be encoded to 1 and the number of vertices to be encoded to 3. Thread A, following the ascending index order, first obtains primitive Y (index 6, vertex data B, C, D), updating the number of primitives to be encoded to 2 and the number of vertices to 4 (A, B, C, D); thread A continues to obtain primitive Z (index 7, vertex data D, E, F), updating the number of primitives to be encoded to 3 and the number of vertices to 6 (A, B, C, D, E, F), triggering the stopping condition that the number of vertices cannot exceed 5, and the final number of readable primitives is 2 (primitives X and Y).

[0085] Based on the above embodiments disclosed in this application, the efficiency of primitive processing can be improved by executing the above operations in parallel through multiple threads. At the same time, the dynamic condition termination mechanism can flexibly adapt to the storage limit requirements under different scenarios, and ultimately maximize the number of coded primitives under resource constraints, thereby optimizing the structural rationality of the coded data block.

[0086] In some embodiments, encoding at least two primitives to be encoded based on a first encoding method to obtain a structure of the first encoding method includes: encoding at least two primitives to be encoded of the target thread based on the first primitive index and the first encoding method using the target thread with the largest number of encoding possibilities to obtain a structure of the first encoding method; the first primitive index is the primitive index of the primitive to be stored.

[0087] Among them, the target thread is the thread with the largest number of coded primitives (the number of primitives that meet the encoding threshold of the first encoding method) among all parallel executing threads. The primitive group corresponding to the target thread (including the primitives to be stored under its own responsibility and other adjacent primitives) can achieve maximum vertex data reuse through the first encoding method.

[0088] In this embodiment of the application, in order to maximize the reuse of shared vertex data and reduce redundant storage, the thread that can encode the most primitives (i.e. the thread with the largest number of primitives to be encoded) is processed first. By comparing the number of primitives that can be encoded by each thread, the optimal target thread is selected to encode the corresponding primitives to be encoded. After processing, the number of primitives that can be encoded by other threads is updated (after deducting the primitives that have already been encoded). The remaining primitives to be stored are processed using the second encoding method.

[0089] In some embodiments, the method further includes: in response to completing the encoding process of the at least two primitives to be encoded, updating the number of coded elements for each thread; based on the updated number of coded elements for each thread, executing the encoding process of other primitives to be encoded that satisfy the first encoding method and updating the number of coded elements for each thread, until there is no number of coded elements that meet the quantity requirement.

[0090] In some embodiments, the quantity requirement can be an integer of 2 or greater. This quantity requirement is related to the encoding amount threshold of the first encoding method; in some embodiments, the quantity requirement is less than the encoding amount threshold.

[0091] In some embodiments, updating the number of coded elements for each thread may include: updating the encoding state of each primitive to be stored in the shared memory; and updating the number of coded elements for each thread based on the encoding state of each primitive to be stored in the shared memory in response to all threads completing the update of their encoding states.

[0092] In some embodiments, the number of coded elements calculated by all threads can be centrally obtained; all coded elements are compared, and the thread with the largest value is selected as the target thread; based on the first encoding method, at least two primitives to be encoded corresponding to the target thread (including the primitives to be stored that the target thread is responsible for and other primitives obtained) are encoded to generate a structure of the first encoding method; subsequently, the number of coded elements of other threads is updated (excluding primitives already encoded by the target thread); this process is repeated (the number of coded elements of the remaining threads is compared again, and the largest one is selected for encoding) until the number of coded elements of all threads no longer satisfies the first encoding method; finally, the remaining primitives to be stored are encoded using the second encoding method to obtain a structure of the second encoding method.

[0093] It should be noted that this step is performed in parallel by multiple threads. That is, all threads perform the following operations in parallel: Each thread first calculates its own coded quantity; then, by accessing shared memory or performing atomic operations, it compares the coded quantities of all current threads to determine whether it is the target thread with the largest coded quantity; if so, it encodes at least two primitives to be encoded (including the primitives to be stored that it is responsible for and other primitives that it has acquired) based on the first encoding method, generates a structure of the first encoding method, and updates the coded quantities of other threads through broadcasting or shared memory (subtracting the primitives that have already been encoded); if not, it waits for the current target thread to complete the encoding; this process is repeated (all threads re-determine whether they are the new target thread) until the coded quantities of all threads do not meet the encoding quantity threshold of the first encoding method; finally, the remaining primitives to be stored (which do not meet the first encoding condition) are encoded using the second encoding method.

[0094] For example, assume that the number of coded elements of threads A, B, C, and D are 3, 2, 1, and 1 respectively, and the corresponding primitives are X, Y, Z, and O respectively. All threads execute in parallel. Thread A determines that it has the largest number of coded elements (3), and as the target thread, it performs the encoding process of the primitive X it is responsible for and the other primitives Y and Z it has obtained using the first encoding method to generate the structure of the first encoding method. Subsequently, thread A updates its own number of coded elements through the encoding status of the shared memory primitives X, Y, and Z. That is, after deducting X, Y, and Z, the number of coded elements of threads B and C becomes 0, and the number of coded elements of thread D becomes 1, which is less than 2. Finally, the remaining uncoded primitive O is processed using the second encoding method.

[0095] Based on the above embodiments disclosed in this application, the thread with the largest number of coded elements can be selected first for encoding, generating a structure and updating the number of coded elements of other threads. This can maximize the reuse of shared vertex data, reduce redundant storage, and store the remaining primitives independently through the second encoding method. This can significantly improve primitive encoding efficiency, reduce video memory usage, and optimize and accelerate the overall performance of structure construction while ensuring data correctness.

[0096] In some embodiments, the method further includes encoding the remaining primitives to be stored based on a second encoding method.

[0097] The first encoding method differs from the second encoding method. The second encoding method is another primitive encoding algorithm provided in this application embodiment, used to encode the remaining primitives to be encoded that do not meet the conditions of the first encoding method. Each primitive independently stores vertex data, which is used when the primitive index is not continuous or vertices cannot be shared.

[0098] It should be noted that the structure using the first encoding method and the structure using the second encoding method occupy the same amount of memory or storage space. In some embodiments, the structure using the first encoding method and the structure using the second encoding method may include fields with the same function, but the specific content stored in the fields may be the same or different.

[0099] In some embodiments, during the encoding of multiple primitives to be stored, it is necessary to repeatedly filter out primitive groups (including at least two primitives to be encoded) that meet the conditions of the first encoding method from the unencoded primitives to be stored, until the remaining primitives to be stored do not meet the conditions of the first encoding method, then the second encoding method is used for encoding and storage.

[0100] In some embodiments, encoding the remaining primitives to be stored based on the second encoding method includes: for each thread, under the condition that the encoding conditions of the second encoding method are met, obtaining a threshold number of unencoded primitives in the first direction based on the order of the primitive indexes, and encoding the obtained unencoded primitives based on the second encoding method to obtain a structure of the second encoding method.

[0101] The encoding conditions of the second encoding method include: the primitives of the thread are not encoded; and in the second direction based on the order of the primitive index, the number of unencoded primitives is an integer multiple of the primitive number threshold of the second encoding method.

[0102] Wherein, the number of unencoded primitives obtained is less than or equal to the primitive number threshold of the second encoding method, and the first direction is opposite to the second direction.

[0103] In some embodiments, the primitive number threshold of the second encoding method is the maximum number of primitives allowed to be encoded and stored when the structure is stored based on the second encoding method. For example, the maximum number of primitives can be 3.

[0104] Here, the first direction can be either the increasing direction of the primitive index or the decreasing direction of the primitive index; correspondingly, the first direction is opposite to the second direction.

[0105] In some possible implementations, for each thread, the thread verifies whether its associated primitives are unencoded and counts the total number of unencoded primitives by indexing in the second direction. If the total number of unencoded primitives is an integer multiple of the primitive quantity threshold of the second encoding method and its associated primitives are unencoded, then the thread obtains the primitive quantity threshold number of unencoded primitives of the second encoding method by indexing in the first direction, and encodes these unencoded primitives by applying the second encoding method to obtain the structure of the second encoding method.

[0106] For example, suppose thread 2 detects that there were 6 unencoded primitives in the previous thread (satisfying the multiple of 3 condition), then it starts scanning from the current thread 2. Suppose that the primitives corresponding to threads 5 and 9 are not encoded in the next thread, then thread 2 can encode the primitives corresponding to threads 2, 5 and 9 into a structure with a second encoding method. From the perspective of thread 5, since there were 7 unencoded primitives in the previous thread (not satisfying the multiple of 3 condition), thread 5 will not perform the encoding process.

[0107] Based on the above embodiments disclosed in this application, by defining an encoding trigger condition (the number of unencoded primitives counted along the second direction is an integer multiple of a threshold), the set of primitives that meet the batch encoding conditions can be accurately filtered during multi-threaded parallel processing, reducing the occupation of computing resources by invalid encoding operations; at the same time, by obtaining unencoded primitives not exceeding the threshold in the indexing order along the first direction (opposite to the second direction) and applying the second encoding method to generate a structure, the standardized compressed storage of discrete primitives can be achieved, reducing the storage space occupancy rate and improving the efficiency of subsequent data parsing.

[0108] In some embodiments, the method further includes:

[0109] The maximum number of coded threads is determined based on the number of coded threads described above;

[0110] Among at least one thread corresponding to the maximum number of coded elements, the thread with the smallest thread index is determined as the target thread; wherein the order of the thread indices of the thread is the same as or the opposite of the order of the element indices of the corresponding elements to be stored.

[0111] In some possible implementations, all threads perform the following operations in parallel: each thread independently calculates its own coded quantity; based on the coded quantity of each thread, the maximum coded quantity is determined, and then a mask is generated using the maximum coded quantity to mark all threads with that maximum coded quantity; if there are multiple marked threads, the thread with the smallest or largest index in the mask is found.

[0112] For example, suppose there are 4 threads (indices 0 to 3) with 3, 3, 2, and 1 coded elements respectively, and the thread index order is the same as the element index order of the corresponding elements to be stored (i.e., index 0 corresponds to element 0, index 1 corresponds to element 1, index 2 corresponds to element 2, and index 3 corresponds to element 3), or the thread index order is the reverse of the element index order of the corresponding elements to be stored (i.e., index 0 corresponds to element 3, index 1 corresponds to element 2, index 2 corresponds to element 1, and index 3 corresponds to element 0). The same or reversed scenarios do not affect the implementation process of this application. All threads perform the following operations in parallel: Each thread determines its own coded element count (e.g., thread 0 is 3, thread 1 is 3, thread 2 is 2, and thread 3 is 1). Based on the coded element count of each thread, all threads determine the maximum coded element count as 3. Based on the maximum value of 3, a binary mask is generated (e.g., 1100, indicating that the coded element counts of threads 0 and 1 are equal to the maximum value). If the mask contains multiple flag bits (e.g., 1100), the target thread is selected by finding the first 1 bit in the mask (i.e., thread 0 with the smallest index); or, the target thread is selected by choosing the last 1 bit in the mask (i.e., thread 1 with the largest index). It can be seen that using a binary mask (e.g., 1100) efficiently marks all candidate threads with the maximum number of coded threads, avoiding the overhead of traversing all threads. Simultaneously, bit operations on the mask (such as finding the first or last 1 bit) can quickly determine the index of the target thread, reducing global synchronization and computational complexity, and improving parallel coding efficiency.

[0113] Based on the embodiments provided in this application, the coordinated execution of the above steps can reduce thread contention and improve coding efficiency while ensuring coding correctness. At the same time, by prioritizing the coding of the primitive group with the most reusable vertex data, storage redundancy can be reduced and the overall performance during structure construction can be optimized and accelerated.

[0114] In some embodiments, the step of sequentially acquiring the other primitives based on the primitive index, and continuing to update the number of primitives to be encoded and the number of vertices to be encoded based on the acquired other primitives, includes: acquiring the primitive data of the current primitive from other threads; the current primitive is determined based on the order of the primitive index; if it is determined based on the primitive data of the current primitive that there is a target vertex in the current primitive that overlaps with a candidate vertex in the candidate vertex set, assigning a vertex index of the candidate vertex to the target vertex; the candidate vertex set is used to store candidate vertices and their corresponding vertex indices; if it is determined based on the number of vertices in the candidate vertex set, the vertex number threshold, and the number of vertices in the current primitive that have not been assigned vertex indices that the candidate vertex set can accommodate all vertices of the current primitive, generating and assigning new vertex indices to the vertices in the current primitive that have not been assigned vertex indices, and updating the candidate vertex set, the number of primitives to be encoded, and the number of vertices to be encoded.

[0115] In this embodiment, each element to be stored is assigned to a corresponding thread and acquired by that thread. To reduce the pressure on memory access, a thread-to-thread data synchronization mechanism can be used, allowing the current thread to acquire other elements from other threads corresponding to other elements. This embodiment describes the acquisition process for any one of the other elements during the acquisition of at least one other element. Here, "current element" refers to any one of the other elements.

[0116] For example, suppose thread A processes primitive X (index 5) and needs to obtain the data for the next consecutive primitive Y (index 6). Thread A sends a request to thread B, which is processing primitive Y, through an inter-thread synchronization mechanism. Thread B returns the vertex data of primitive Y (such as vertices B, C, and D) and primitive index 6 to thread A, and thread A completes the acquisition of primitive Y data.

[0117] Here, the candidate vertex set is a temporary data structure used to store vertices with assigned vertex indices and their corresponding indices. This is used to quickly determine whether a new primitive's vertex already has a shared vertex, thus enabling vertex data reuse. By comparing the vertices of the current primitive with those in the candidate vertex set, if a common vertex (target vertex) exists, the vertex index corresponding to that vertex in the candidate vertex set is directly assigned, achieving vertex data sharing and reuse.

[0118] In some embodiments, the current thread parses the vertex data of the current primitive obtained from other threads and compares them one by one with the vertices in the candidate vertex set. If a vertex (target vertex) of the current primitive is found to be completely consistent with a vertex (candidate vertex) in the candidate vertex set, the vertex index of the candidate vertex is assigned to the target vertex.

[0119] For example, the candidate vertex set already stores vertex A (index 0), vertex B (index 1), and vertex C (index 2). The vertex data of the current primitive Y are B, C, and D, where vertices B and C coincide with the vertices in the candidate vertex set. Thread A assigns index 1 to vertex B and index 2 to vertex C.

[0120] In some embodiments, it can be determined whether the candidate vertex set can accommodate all vertices of the current primitive based on the following method: the current thread calculates the remaining number of vertices in the candidate vertex set, i.e., the vertex number threshold minus the current number of vertices in the candidate vertex set; if the remaining number of vertices is greater than or equal to the number of vertices in the current primitive that are not assigned vertex indices, then it is determined that the candidate vertex set can accommodate all vertices (vertex indices) of the current primitive; if the remaining number of vertices is less than the number of vertices in the current primitive that are not assigned vertex indices, then it is determined that the candidate vertex set cannot accommodate all vertices (vertex indices) of the current primitive.

[0121] In some embodiments, if it is determined that the candidate vertex set can accommodate all vertices (vertex indices) of the current primitive, new vertex indices can be generated for vertices without assigned indices (e.g., assigned in ascending order), the new vertices and their corresponding vertex indices can be added to the candidate vertex set, and the number of primitives to be encoded (incremented by 1) and the number of vertices to be encoded (incremented by the number of new vertices).

[0122] For example, suppose the candidate vertex set currently stores 3 vertices (vertices A, B, and C), the vertex count threshold is 5, and the current primitive contains 2 vertices (D and E) without assigned vertex indices; the current thread calculates the remaining vertex count as 2 (5-3=2), finds that this remaining count is equal to the number of unassigned vertices in the current primitive, and therefore determines that the candidate vertex set can accommodate all vertices of the current primitive; then, new vertex indices (such as 3 and 4) are generated for vertices D and E, D and E and their indices are added to the candidate vertex set, and the number of primitives to be encoded is updated from 1 to 2, and the number of vertices to be encoded is increased from 3 to 5, completing the vertex index allocation and data update.

[0123] Based on the above embodiments disclosed in this application, by obtaining the current primitive's primitive data from other threads through the inter-thread data synchronization mechanism, the frequent access to memory can be reduced, data access latency can be lowered, and data acquisition efficiency during multi-threaded parallel processing can be improved. At the same time, when it is detected that the vertex of the current primitive coincides with the vertex in the candidate vertex set, an existing vertex index is assigned to the coinciding vertex. By reusing the index information in the candidate vertex set, the duplicate storage of the same vertex data can be avoided, reducing redundant storage of vertex information. By calculating the remaining number of vertices in the candidate vertex set and comparing it with the number of vertices in the current primitive that have not been assigned an index, a new index is generated and the set and encoding number are updated when there is room for storage. This ensures that vertex storage does not exceed the capacity limit and maximizes the utilization of the vertex storage space of the structure.

[0124] In some embodiments, the plurality of primitives to be stored correspond to a thread group, and the plurality of threads in the thread group execute the encoding process of the plurality of primitives to be stored in parallel; the step of allocating a corresponding thread for each primitive to be stored includes: allocating a thread for each primitive to be stored among the plurality of threads corresponding to the plurality of primitives to be stored, according to the order of the first primitive index of each primitive to be stored in the geometry.

[0125] In some embodiments, a corresponding thread can be allocated for each primitive to be stored based on the order (e.g., ascending or descending) of its first primitive index in the geometry, with a one-to-one mapping between thread indices and primitive indices. For example, if the geometry contains primitives 0 to n, then threads 0 to n will process the corresponding primitives 0 to n respectively, ensuring that the thread execution order is strictly consistent with the primitive logical order. This facilitates subsequent encoding processes, allowing the current thread to obtain other primitives with potentially consecutive first primitive indices from adjacent threads.

[0126] In some embodiments, the method further includes: obtaining the number of primitives in the geometry; determining at least one thread group and the primitives corresponding to each thread group based on the number of primitives in the geometry and a preset number of threads; wherein the thread group includes the number of threads, and the number of threads in the thread group is greater than or equal to the number of primitives corresponding to the thread group.

[0127] A thread group is a collection of multiple threads that execute in parallel. Each thread group contains a preset number of threads used to perform the encoding process of the multiple primitives to be stored in parallel. The preset number of threads is the number of threads contained in each thread group (e.g., 64), used to control the computational granularity of a single thread group and balance load and resource utilization.

[0128] In some embodiments, the total number of primitives to be stored in the geometry is obtained, i.e., the number of primitives; based on the preset number of threads in each thread group (e.g., 64), the number of thread groups is calculated by rounding up; a continuous range of primitives is allocated to each thread group; if the number of primitives in the last thread group is less than the preset number of threads, some of its threads remain idle to ensure that all primitives are covered.

[0129] In some embodiments, the rounding formula can be: N = (number of primitives + number of preset threads - 1) / number of preset threads.

[0130] In some embodiments, the range of primitives processed by the i-th thread group can be: i * preset number of threads to (i+1) * preset number of threads - 1.

[0131] For example, suppose the geometry contains 100 primitives to be stored (number of primitives = 100), and the preset number of threads in each thread group is 64 (preset number of threads = 64); the number of thread groups N is calculated by the formula N = (100 + 64 - 1) / 64 = 163 / 64 ≈ 2 (rounded up), that is, divided into 2 thread groups; the 0th thread group (Group0) contains 64 threads, processing primitives 0-63; the 1st thread group (Group1) contains 64 threads, processing primitives 64-99 (a total of 36 primitives), of which threads 36-63 (a total of 28 threads) are idle because there are no corresponding primitives.

[0132] Based on the above embodiments disclosed in this application, by dynamically calculating the number of thread groups and evenly distributing tasks according to the number of primitives, it is possible to adapt to the coding needs of geometric objects of different scales and avoid wasting thread resources. At the same time, by controlling the calculation granularity of a single thread group by preset thread number, the load can be balanced and idle threads can be reduced, thereby improving the overall efficiency of parallel computing.

[0133] In some embodiments, encoding at least two primitives of the target thread based on the first primitive index and the first encoding method to obtain the structure of the first encoding method includes: encoding at least two primitives of the target thread based on the first encoding method, the number of encodeable primitives, the first primitive index, and the candidate vertex set to obtain the structure of the first encoding method.

[0134] In some embodiments, based on the first encoding method, the number of coded elements, the first primitive index, and the candidate vertex set, the identifier of the first encoding method and the number of coded elements can be recorded in the encoding information; the first primitive index (such as the minimum or maximum primitive index) can be recorded in the index information, and a corresponding vertex index can be assigned to each vertex of the primitive to be encoded based on the vertex index of the candidate vertex set; the vertex data (such as vertex coordinates, normal vectors, etc.) in the candidate vertex set can be stored sequentially in the vertex information to complete the construction of the first encoding method structure.

[0135] In some embodiments, the above-described structure for encoding at least two primitives of the target thread based on the first encoding method, the number of encodeable primitives, the first primitive index, and the candidate vertex set to obtain the first encoding method may include: generating the encoding information based on the first encoding method and the number of encodeable primitives; determining the index information based on the first primitive index and the vertex index of the candidate vertex in the candidate vertex set; and generating the vertex information by encoding the corresponding vertex data based on the vertex index of the candidate vertex in the candidate vertex set.

[0136] Specifically, based on the first encoding method, the encoding method identifier (e.g., setting 1 binary bit to 0) is stored in the encoding information; based on the coded quantity, the number of primitives is converted into a binary number (e.g., 4 bits) and stored in the encoding information, thus completing the generation of the encoding information. For example, assuming the coded quantity is 3 and the first encoding method identifier is 0, in the encoding information, 1 bit of the identifier is set to 0, and the 4-bit quantity parameter is set to 0011 (binary 3), together constituting the encoding information.

[0137] The process involves recording the first primitive index in the index information. Then, based on the vertex indices in the candidate vertex set, a corresponding vertex index is assigned to each vertex of the primitive to be encoded (e.g., vertex A of primitive 0 corresponds to index 0, and vertex B corresponds to index 1). These vertex indices are then stored in the index information in primitive order, completing the construction of the index information. For example, assuming the first primitive index is 0, and the candidate vertex set includes vertices A (index 0), B (index 1), and C (index 2), the index information records the first primitive index 0, and assigns indices 0, 1, and 2 to vertices A, B, and C of primitive 0, and indices 1, 2, and 3 to vertices B, C, and D of primitive 1, thus completing the generation of the index information.

[0138] The process involves iterating through each vertex in the candidate vertex set and storing its data sequentially in the vertex information based on its vertex index. For example, vertex index 0 corresponds to the data of vertex A, index 1 corresponds to the data of vertex B, and so on, thus generating the vertex information. For instance, suppose the candidate vertex set contains vertices A (index 0, data is (0,0,0)), B (index 1, data is (1,0,0)), and C (index 2, data is (1,1,0)). The vertex information stores the coordinate data of these three vertices sequentially, forming the vertex information portion.

[0139] Based on the above embodiments disclosed in this application, a first encoding method structure with a compact structure and strong data correlation can be generated, thereby improving the efficiency of primitive encoding, reducing storage redundancy, and improving the utilization rate of video memory.

[0140] In some embodiments, the index information includes a first index parameter bit to a fourth index parameter bit, and the index parameter bit is 32 bits; the vertex information includes 9 vertex parameter bits; in the structure of the first encoding method, the primitive number threshold is 8, the vertex number threshold is 9, the first index parameter bit is used to store the first primitive index; the second index parameter bit to the fourth index parameter bit are used to store the vertex index of 32 vertices; the vertex index is 4 bits; in the structure of the second encoding method, each index parameter bit from the first index parameter bit to the third index parameter bit is used to store a primitive index.

[0141] In some possible implementations, the index parameter bits are 32-bit binary fields in the structure used to store primitive or vertex indices, and multi-primitive or multi-vertex index storage is achieved through bit allocation. A vertex parameter bit is a field in the structure used to store the vertex data of a single vertex.

[0142] like Figure 8 As shown, the index information of the structure may include the first index parameter bit 91, the second index parameter bit 92, the third index parameter bit 93 and the fourth index parameter bit 94; the vertex information 43 may include 9 vertex parameter bits, which store the vertex data of vertices 0 to 8 respectively.

[0143] In some embodiments, please refer to Figure 9For the first encoding method, the first index parameter bit 91 stores the index of the first primitive; the second index parameter bit 92, the third index parameter bit 93, and the fourth index parameter bit 94 are used to store the vertex indices of 32 vertices. Since the vertex index ranges from 0 to 8, it requires 4 bits to represent. The aforementioned second to fourth index parameter bits are each divided into eight 4-bit units according to the index range, each corresponding to one vertex of eight primitives. Since one primitive includes three vertices, the second index parameter bit can sequentially store the vertex indices of the first vertices of the eight primitives; the third index parameter bit can sequentially store the vertex indices of the second vertices of the eight primitives; and the fourth index parameter bit can sequentially store the vertex indices of the third vertices of the eight primitives. Figure 9 As shown, taking the storage of the first primitive in the structure as an example, the vertex index of the first vertex of the first primitive can be stored in the 0th to the 3rd bit of the second index parameter bit 92, i.e., 951; the vertex index of the second vertex of the first primitive can be stored in the 0th to the 3rd bit of the third index parameter bit 93, i.e., 952; the vertex index of the third vertex of the first primitive can be stored in the 0th to the 3rd bit of the fourth index parameter bit 93, i.e., 953; thus, 950 is the vertex index of all the first primitives, and so on. It can be seen that the structure can store the primitive number threshold (8) primitives through the second index parameter bit 92, the third index parameter bit 93 and the fourth index parameter bit 94.

[0144] In some embodiments, please refer to Figure 10 For the second encoding method, the first to third index parameters are used to store the second primitive indexes of the three primitives. In vertex information 43, the nine vertex parameter bits can be divided into three groups, each corresponding to one of the three primitives. Figure 10 An example of grouping is given, in which the first to third vertex parameter bits (961) are assigned to the first primitive, the fourth to sixth vertex parameter bits (962) are assigned to the second primitive, and the seventh to ninth vertex parameter bits (963) are assigned to the third primitive.

[0145] Based on the embodiments disclosed in this application, by designing the index parameter bits to be 32 bits and using bit-grouping to store the vertex index (such as storing the vertex index of 8 primitives in a group of 4 bits in the first encoding method), and storing vertex data in groups as needed (such as sharing vertices in the first encoding method and storing them independently in the second encoding method), the structure design of the two encoding methods can be compatible, improving encoding flexibility. Through unified bit allocation rules and threshold control (such as the primitive number threshold 8 and the vertex number threshold 9), the storage space utilization can be optimized and redundant data can be reduced. At the same time, by storing the index and vertices separately or in a shared manner, the encoding requirements of different scenarios can be adapted, data access latency can be reduced, and the encoding efficiency and computational performance during accelerated structure construction can be improved.

[0146] In some embodiments, the method further includes: obtaining the number of generated structures; determining a target storage address based on the number of generated structures, historical storage addresses, and the size of the structures; and storing the structures obtained using the first encoding method or the second encoding method based on the target storage address.

[0147] The number of generated structures refers to the total number of structures created and stored during the storage process using either the first or second encoding method; that is, the sum of the number of structures using the first encoding method and the number of structures using the second encoding method. It should be noted that structures using the first and second encoding methods occupy the same number of bytes in storage space.

[0148] In some possible implementations, the number of generated structures can be obtained, and combined with historical storage addresses and structure sizes, the starting storage address of the next structure, i.e., the target storage address, can be calculated using an address calculation formula. The newly generated structure using either the first or second encoding method is then stored based on this target storage address. Here, the address calculation formula can be: Target storage address = Historical storage address + Number of generated structures × Structure size.

[0149] For example, assuming five structures have been generated, each with a uniform size of 128 bytes, and the historical storage address is 0x1000, then the target storage address is 0x1000 + 5 × 128 = 0x1400, and the next structure (the sixth one) will be stored at address 0x1400. If the sixth structure uses the first encoding method, its encoding information, index information, and vertex information will be written into the 128 bytes of memory starting at address 0x1400 according to the rules.

[0150] The following describes the application of the primitive encoding method provided in the embodiments of this application in a real-world scenario.

[0151] This application provides a method for encoding BVH leaf nodes, the implementation of which is as follows:

[0152] 1. On the CPU side, the Driver parses the D3D12_BUILD_RAYTRACING_ACCELERATION_STRUCTURE_DESC parameter passed from the App. For each Geometry, it first calculates the number of primitives it contains, NumPrimitives. Then, the Driver internally starts a Compute Pipeline to encode the triangle primitives (hereinafter referred to as triangles) contained in the Geometry into BVH leaf nodes in parallel. The specific definition of the leaf nodes is as follows:

[0153] struct TriangleNode {

[0154] uint32 info;

[0155] uint32 primIndex0;

[0156] uint32 primIndex1_or_index0;

[0157] uint32 primIndex2_or_index1;

[0158] uint32 index2;

[0159] float3 vtx0;

[0160] float3 vtx1;

[0161] float3 vtx2;

[0162] float3 vtx3;

[0163] float3 vtx4;

[0164] float3 vtx5;

[0165] float3 vtx6;

[0166] float3 vtx7;

[0167] float3 vtx8;

[0168] };

[0169] The info field uses 4 bits to encode the number of triangular primitives stored in the TriangleNode, and 1 bit to indicate the primitive encoding method, including the first encoding method (hereinafter referred to as Triangle List) and the second encoding method (hereinafter referred to as Indexed Triangle List). Since a TriangleNode can store a maximum of 9 vertex data, the Indexed Triangle List encoding method can encode a maximum of 8 triangles (in the current embodiment, primitives are all represented by triangles), and the Triangle List encoding method can represent a maximum of 3 triangles.

[0170] In some embodiments, when using the Triangle List encoding method, the three parameters primIndex0, primIndex1_or_index0, and primIndex2_or_index1 represent the primitive index of the triangle in the Geometry. vtx0~vtx8 store the vertex data of the primitives in the TriangleNode. In the current embodiment, index2 is not used.

[0171] In some embodiments, when the Indexed Triangle List encoding method is used, primIndex0 represents the primitive index of the starting primitive in Geometry, and the subsequent primitive indices are sequentially increased (8 in a sequential increase). That is, only triangles with consecutive indices in Geometry can be encoded into the same TriangleNode using the Indexed Triangle List method. vtx0-vtx8 store the vertex data of the primitives in the TriangleNode. primIndex1_or_index0, primIndex2_or_index1, and index2 represent the vertex indices of the 8 primitives, and the index range is [0, 8].

[0172] Since index0-index2 are all 32 bits and the index range is 0 to 8, the designed structure can store a maximum of 8 triangles with a total of 9 vertices. The index of each triangle can be represented by 4 bits of data, and 4 bits multiplied by 8 equals 32 bits. Since each triangle primitive requires three vertex indices, index0, index1, and index2 can represent the vertex indices of 8 triangles in total.

[0173] To make it easier to understand, index0, index1, and index2 can be viewed as a two-dimensional array or matrix with 3 rows and 32 columns. Columns 0 to 3 represent the vertex indices of triangle 0. The first row is the vertex index of vertex 0, the second row is the vertex index of vertex 1, and the third row is the vertex index of vertex 2. Thus, 32 columns divided by 4 equals 8, indicating that a total of 8 vertex indices of triangles can be represented.

[0174] 2. Before executing the Compute Pipeline, the Driver binds resources such as the ConstantBuffer and UAVs (Index Buffer / Vertex Buffer / Scratch Buffer) to the RootSignature. The ConstantBuffer stores the following parameter: indexFormat: Vertex index format;

[0175] totalPrimitives: The total number of primitives to be encoded in the current Geometry;

[0176] numPrimsDoneOffset: The offset of the number of primitives that have been encoded in the ScratchBuffer;

[0177] encodedPrimsOffset: The starting address for writing TriangleNode in ScratchBuffer.

[0178] In some embodiments, the encoding of BVH leaf nodes is a parallel computation with primitives as the task unit. The Driver divides the computation into N groups of threads (NumPrimitives) based on the number of primitives in the geometry. Each wave group is allocated 64 threads (THREADGROUP_SIZE), i.e., N = (NumPrimitives + THREADGROUP_SIZE - 1) / THREADGROUP_SIZE. For example, when there are 100 primitives (NumPrimitives), it needs to be divided into N groups, N = (100 + 64 - 1) / 64 = 163 / 64 = 2 (rounded down). Among them, the 64 threads in Group 0 process primitives 0-63; the 64 threads in Group 1 process primitives 64-99 (the remaining 36), and threads 36-63 can be idle. Then each Wave Thread obtains its global thread ID (DispatchThreadID) and uses it as a primitive index to read all vertex data (TriangleData) of a triangle from the Vertex Buffer passed in by the App. For example... Figure 11 As shown, each thread in a set of Wave Threads is used to read the vertex data of a corresponding triangle primitive from the vertex buffer.

[0179] Additionally, each wave group allocates a shared memory block `bool EncodedThread[THREADGROUP_SIZE]` to indicate whether the TriangleData corresponding to the current thread has been encoded into the TriangleNode. See also... Figure 12 The diagram illustrates a shared memory approach, where each bit in the shared memory represents whether the corresponding triangle primitive has been encoded. Through this shared memory, threads within the same thread group can obtain the task progress of other threads in the same group.

[0180] Then, each Wave Thread calls the WaveReadLaneAt() function in the direction of increasing Lane Index (intra-group index) to try to read the TriangleNode in the 8 consecutive Threads on the right, store the non-overlapping vertex data, and assign an index to the corresponding primitive vertex, with the index range being [0, 8].

[0181] In some embodiments, the pseudocode for implementing the above process is as follows:

[0182] struct LoadTriangleConstants{

[0183] uint indexFormat;

[0184] uint totalPrimitives;

[0185] uint numPrimsDoneOffset;

[0186] uint encodedPrimsOffset;

[0187] };

[0188] struct TriangleData{

[0189] float3 vtx[3];

[0190] };

[0191] ConstantBuffer <loadtriangleconstants>Constants: register(b0);

[0192] RWByteAddressBufferScratchBuffer : register(u0);

[0193] RWByteAddressBufferIndexBuffer: register(u1);

[0194] RWByteAddressBufferVertexBuffer: register(u2);

[0195] groupshared bool EncodedThread[THREADGROUP_SIZE];

[0196] [numthreads(THREADGROUP_SIZE, 1, 1)]

[0197] void LoadTrianglePrimitives(

[0198] uint globalId : SV_DispatchThreadID

[0199] uint localId: SV_GroupThreadID)

[0200] {

[0201] if (globalId<Constants.totalPrimitives)

[0202] {

[0203] const uint primitiveIndex = globalId;

[0204] uint3 triangleIndex = LoadTriangleIndex(IndexBuffer, primitiveIndex);

[0205] / / Read a triangle vertex data from the Vertex Buffer

[0206] TriangleData currTriangle = LoadTriangleData(triangleIndex,VertexBuffer);

[0207] const uint maxEncodeVertices = 9; / / A TriangleNode can store a maximum of 9 vertices.

[0208] const uint maxEncodeTriangles = 8; / / A TriangleNode can store a maximum of 8 triangles.

[0209] float3 candidateVertices[maxEncodeVertices]; / / Candidate vertices that can be encoded

[0210] uintcandidateIndices[maxEncodeTriangles][3]; / / Candidate triangles that can be encoded

[0211] uint numCandidateVertices = 0; / / Number of candidate vertices that can be encoded

[0212] uint numCandidateTriangles = 0; / / Number of candidate triangles that can be encoded

[0213] uint laneIndex = WaveGetLaneIndex();

[0214] uint laneCount = WaveGetLaneCount();

[0215] for (uint i = 0; i<8&&((laneIndex + i) <laneCount); i++)

[0216] {

[0217] TriangleData candidateTri = WaveReadLaneAt(currTriangle, laneIndex +i);

[0218] / / Check if the vertices in candidateTri (the candidate triangle) overlap with those in candidateVertices. If they overlap, initialize an index value for each vertex.

[0219] for (uint j = 0; j<3; j++)

[0220] {

[0221] for (uint k = 0; k <numCandidateTriangles; k++)

[0222] {

[0223] if (IsSameVtx(candidateTri.vtx[j], candidateVertices[k]))

[0224] {

[0225] candidateIndices[i][j] = k;

[0226] }

[0227] }

[0228] }

[0229] / / Check if there are any extra spaces in the candidateVertices array to place a new vertex.

[0230] if ((maxEncodeVertices - numCandidateVertices)>= NumInvalidIndices(candidateIndices[i]))

[0231] {

[0232] / / If the remaining space can accommodate candidateTri, then calculate an index value for each vertex.

[0233] numCandidateTriangles++;

[0234] for (uint j = 0; j<3; j++)

[0235] {

[0236] if (candidateIndices[i][j]== INVALID_IDX)

[0237] {

[0238] candidateIndices[i][j]= numCandidateVertices;

[0239] candidateVertices[numCandidateVertices++] = candidateTri.vtx[j];

[0240] }

[0241] }

[0242] }

[0243] }

[0244] / / Identify a continuous segment of triangles that can be compressed and encoded using a tuple (starting primitive index, number of candidate triangles).

[0245] uint2 candidateTriangles(primitiveIndex, numCandidateTriangles);

[0246] / / Select Wave Thread to start coding and store TriangleNode

[0247] }

[0248] }

[0249] 3. After the above process is completed, each thread (Wave Thread) can know the maximum number of triangles that can be encoded in the incrementing direction starting from the current Lane Index. Each time, the tuple candidateTriangles with the largest numCandidateTriangles value is selected to start encoding. If there are multiple tuples with the largest numCandidateTriangles value at the same time, the thread with the smallest Index is selected to encode.

[0250] like Figure 13 As shown, each time the tuple with the largest numCandidateTriangles value is selected to start encoding the BVH leaf node, that is, the tuple (3,6) is selected to start encoding first. Thread 3 is responsible for encoding TriangleData 3 / 4 / 5 / 6 / 7 / 8 into a TriangleNode and writing it to the specified position of ScratchBuffer, updating the shared memory. The range of the EncodedThread array [3,8] indicates that the TriangleData in the corresponding thread has been encoded.

[0251] The remaining threads in the group update the tuple by querying the latest state of EncodedThread, such as... Figure 14 As shown, the encoding will begin with the remaining tuples with the largest numCandidateTriangles value, i.e., thread 0 with the largest numCandidateTriangles value will be selected to encode TriangleData 0 / 1 / 2; accordingly, the shared memory will be updated, and the range of the EncodedThread array [0,2] indicates that the TriangleData in the corresponding thread has been encoded.

[0252] In some embodiments, the pseudocode for implementing the above process is shown below:

[0253] EncodedThread[localId] = false;

[0254] / / Number of triangles that need to be encoded in the current Wave Group

[0255] uint localTrianglesSum = WaveActiveSum(1);

[0256] / / Number of triangles currently encoded in the Wave Group

[0257] uint numEncodedTriangles = 0;

[0258] while(localTrianglesSum>numEncodedTriangles)

[0259] {

[0260] / / Each time, the Wave Thread with the largest numCandidateTriangles value is selected to start coding first.

[0261] uint maxCandidateTriangles = WaveActiveMax(numCandidateTriangles);

[0262] / / Query which threads have a numCandidateTriangles value that is the maximum value within the current Wave Group, represented by a bit mask.

[0263] Uint64 candidateLaneMask= WaveActiveBallot64(numCandidateTriangles ==maxCandidateTriangles);

[0264] {

[0265] / / If multiple active Lanes have a maximum value of numCandidateTriangles, select the LaneIndex with the smallest value and perform TriangleNode encoding.

[0266] uint minLaneIndex = FirstBitLow(candidateLaneMask);

[0267] if (minLaneIndex == WaveGetLaneIndex())

[0268] {

[0269] / / Query the number of completed TriangleNodes in the global Wave Group

[0270] uint numPrimsDone = 0;

[0271] ScratchGlobal.InterlockedAdd(Constants.encodedPrimIndexOffset, 1,numPrimsDone);

[0272] / / Calculate the storage location of TriangleNode in Scratch Buffer based on numPrimsDone

[0273] uint encodePrimOffset = Constants.encodedPrimsOffset + numPrimsDone *sizeof(TriangleNode);

[0274] / / Encode the candidate Triangles into a TriangleNode

[0275] TriangleNode encodedTriangle = EncodeMultiTriangles(primitiveIndex,numCandidateTriangles, candidateVertices);

[0276] / / Write the TriangleNode into the ScratchBuffer

[0277] WriteTriangleNodeIndexed(encodePrimOffset, encodedTriangle);

[0278] / / Update the EncodedThread array, marking the Lane Index that has been encoded.

[0279] for (uint i = 0; i <numCandidateTriangles; i++)

[0280] {

[0281] EncodedThread[localId + i] = true;

[0282] }

[0283] }

[0284] }

[0285] / / All threads within the Wave Group synchronously wait for the previous EncodedThread data write operation to complete at the current execution point.

[0286] GroupMemoryBarrierWithGroupSync();

[0287] / / Each Lane Thread updates the number of candidate coded triangles based on the latest EncodedThread state.

[0288] uint uncodedCountLeft = WavePrefixCountBits(EncodedThread[localId group] == false);

[0289] uint uncodedCountRight = WaveReadLaneAt(uncodedCountLeft, localId +numCandidateTriangles);

[0290] / / Update the number of remaining encodeable triangles in the current Lane Thread

[0291] numCandidateTriangles = uncodedCountRight - uncodedCountLeft;

[0292] / / Update the number of coded triangles within the Wave Group

[0293] numEncodedTriangles = WaveActiveSum(EncodedThread[localId]);

[0294] }

[0295] Fourth, after all the consecutive triangles have been encoded, there may be several isolated triangles remaining in the Wave Group. In order to make full use of the TriangleNode storage space, the remaining non-consecutive triangles are encoded in the Triangle List manner.

[0296] First, query how many uncoded threads precede the current lane and their corresponding lane index. Figure 15 The solid line represents an unencoded TriangleNode. If the number of unencoded Triangles before the current Lane is a multiple of 3, and the triangleData corresponding to the current thread is unencoded, then starting from the current Lane, find three unencoded TriangleData points and write them to a TriangleNode. Figure 15 As shown, thread 2 can encode TriangleData 2 / 5 / 9, and thread 12 can encode TriangleData12.

[0297] In some embodiments, the pseudocode is as follows:

[0298] / / If the remaining TriangleData cannot be encoded consecutively, then use a Triangle List to place 3 non-consecutive triangles in each TriangleNode.

[0299] if (maxCandidateTriangles == 1)

[0300] {

[0301] / / Query how many uncoded threads precede the current Lane

[0302] uintuncodedPrefixCount = WavePrefixCountBits(EncodedThread[localId] == false);

[0303] / / Query the unencoded threads in the current Wave Group, represented by a bit mask.

[0304] uint64_t uncodedThreadMask= WaveActiveBallot64(EncodedThread[localId]== false);

[0305] / / If the number of unencoded triangles in the current lane is a multiple of 3, and the triangleData corresponding to the current thread is not encoded.

[0306] if ((uncodedPrefixCount % 3 == 0)&&

[0307] (EncodedThread[localId] == false))

[0308] {

[0309] / / Starting from the current Lane, search to the right for at most 3 coded Triangles based on uncodedThreadMask.

[0310] uint3uncodedLaneIndex = BitMaskScanForward3(primitiveIndex,uncodedThreadMask);

[0311] TriangleData triangleList[3];

[0312] if (IsValidLaneIndex(uncodedLaneIndex.x))

[0313] {

[0314] EncodedThread[uncodedLaneIndex.x] = true;

[0315] triangleList[0] = currTriangle;

[0316] }

[0317] if (IsValidLaneIndex(uncodedLaneIndex.y))

[0318] {

[0319] EncodedThread[uncodedLaneIndex.y] = true;

[0320] triangleList[1] = WaveReadLaneAt(currTriangle, uncodedLaneIndex.y);

[0321] }

[0322] if (IsValidLaneIndex(uncodedLaneIndex.z))

[0323] {

[0324] EncodedThread[uncodedLaneIndex.z] = true;

[0325] triangleList[2] = WaveReadLaneAt(currTriangle, uncodedLaneIndex.z);

[0326] }

[0327] WriteTriangleNodeList(uncodedLaneIndex, triangleList);

[0328] }

[0329] }

[0330] Normally, 8 triangles require 8 DWords to store the Primitive Index and 72 DWords to store the vertex data. When constructing the BVH, this method first finds a continuous arrangement of triangles based on the index or vertex sequence input by the APP. In the optimal case, TriangleNode can use 9 vertices and a total of 27 DWords to encode 8 triangles and 1 DWord to encode the Primitive Index of the starting triangle, saving up to 65% of Video Memory.

[0331] In this embodiment, geometric primitives are encoded using an Indexed Triangle List method. One DWord is used to store the index of consecutive primitives, and nine vertex data are used to encode eight consecutive triangles. Furthermore, geometric primitives are encoded in the BVH leaf nodes using both Triangle List and Indexed Triangle List methods, maximizing video memory savings. Thus, encoding primitive data using the Indexed Triangle List method and reusing triangle vertices makes the BVH structure more compact, achieving the goal of saving video memory.

[0332] Based on the foregoing embodiments, this application provides a primitive encoding device, which includes the included units and the modules included in each unit. It can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0333] Figure 16 This is a schematic diagram of the composition structure of a primitive encoding device provided in an embodiment of this application, as shown below. Figure 16 As shown, the primitive encoding device 1400 includes: an encoding module 1410, wherein:

[0334] The encoding module 1410 is used to encode at least two primitives to be encoded based on a first encoding method to obtain a structure of the first encoding method; the structure of the first encoding method includes a shared vertex stored once, the shared vertex being shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded being consecutive in the geometry.

[0335] In some embodiments, the structure includes encoding information, index information, and vertex information; the encoding information is used to indicate the number of primitives and the encoding method of the primitives to be encoded in the primitive node; the vertex information includes the vertex data of each vertex in the primitive to be encoded.

[0336] In the structure of the first encoding method, the index information includes a first primitive index and vertex indices of each vertex in the primitive to be encoded; the first primitive index is the smallest or largest primitive index among the at least two primitives to be encoded; the vertex index is used to locate the corresponding vertex data in the vertex information.

[0337] In some embodiments, the encoding module 1410 is configured to determine, from a plurality of primitives to be stored, at least two primitives that satisfy a first encoding method.

[0338] In some embodiments, the encoding module 1410 is configured to allocate a corresponding thread for each of the primitives to be stored; for each thread, obtain the primitive data of the primitive to be stored through the thread, and obtain the primitive data of at least one other primitive based on the order of the primitive index; determine the primitive to be stored and the at least one other primitive obtained by the target thread with the largest number of coded primitives as the at least two primitives to be encoded; the number of coded primitives is the number of primitives obtained for the primitives to be stored and the at least one other primitive; wherein the primitive to be stored and the at least one other primitive obtained by the thread satisfy the encoding amount threshold of the first encoding method.

[0339] In some embodiments, the encoding threshold includes a primitive number threshold and a vertex number threshold; the encoding module 1410 is configured to obtain the primitive data to be stored through the thread, and update the number of primitives to be encoded and the number of vertices to be encoded of the thread; to obtain other primitives sequentially based on the primitive index, and to continue updating the number of primitives to be encoded and the number of vertices to be encoded based on the obtained other primitives, until a stopping condition is met; wherein, the stopping condition includes at least one of the following: after obtaining the next other primitive, the updated number of primitives to be encoded does not meet the primitive number threshold; after obtaining the next other primitive, the updated number of vertices to be encoded does not meet the vertex number threshold.

[0340] In some embodiments, the encoding module 1410 is configured to determine the number of primitives to be encoded in the thread as the number of coded primitives in the thread in response to the stopping condition being met.

[0341] In some embodiments, the encoding module 1410 is configured to encode at least two primitives of the target thread based on the first primitive index and the first encoding method using the target thread with the largest number of encodeable primitives, to obtain a structure of the first encoding method; the first primitive index is the primitive index of the primitive to be stored.

[0342] In some embodiments, the encoding module 1410 is configured to update the number of coded elements of each thread in response to completing the encoding process of the at least two primitives to be encoded; based on the updated number of coded elements of each thread, execute the encoding process of other primitives to be encoded that satisfy the first encoding method and update the number of coded elements of each thread until there is no number of coded elements that meet the quantity requirement.

[0343] In some embodiments, the encoding module 1410 is configured to obtain graph data of the current graph element from other threads; the current graph element is determined based on the order of the graph element indices; if it is determined based on the graph data of the current graph element that there is a target vertex in the current graph element that overlaps with a candidate vertex in the candidate vertex set, the vertex index of the candidate vertex is assigned to the target vertex; the candidate vertex set is used to store candidate vertices and their corresponding vertex indices; if it is determined based on the number of vertices in the candidate vertex set, the vertex number threshold, and the number of vertices in the current graph element that have not been assigned vertex indices, the candidate vertex set can accommodate all vertices of the current graph element, new vertex indices are generated and assigned to the vertices in the current graph element that have not been assigned vertex indices, and the candidate vertex set, the number of graph elements to be encoded, and the number of vertices to be encoded are updated.

[0344] In some embodiments, the plurality of primitives to be stored correspond to a thread group, and the plurality of threads in the thread group execute the encoding process of the plurality of primitives to be stored in parallel; the encoding module 1410 is used to allocate a thread for each primitive to be stored in the plurality of threads corresponding to the plurality of primitives to be stored according to the order of the first primitive index of each primitive to be stored in the geometry.

[0345] In some embodiments, the encoding module 1410 is used to obtain the number of primitives in the geometry; based on the number of primitives in the geometry and a preset number of threads, determine at least one thread group and the primitives corresponding to each thread group; wherein, the thread group includes the number of threads, and the number of threads in the thread group is greater than or equal to the number of primitives corresponding to the thread group.

[0346] In some embodiments, the encoding module 1410 is used to encode at least two primitives of the target thread based on the first encoding method, the number of encodeable primitives, the first primitive index, and the candidate vertex set, to obtain a structure of the first encoding method.

[0347] In some embodiments, the encoding module 1410 is configured to generate the encoding information based on the first encoding method and the number of coded vertices; determine the index information based on the first primitive index and the vertex index of the candidate vertices in the candidate vertex set; and generate the vertex information by encoding the vertex data corresponding to the vertex index of the candidate vertices in the candidate vertex set.

[0348] In some embodiments, the encoding module 1410 is configured to determine a maximum number of coded elements based on the number of coded elements for each thread; and to determine the thread with the smallest thread index as the target thread among at least one thread corresponding to the maximum number of coded elements; wherein the order of the thread indices of the thread is the same as or the opposite of the order of the element indices of the corresponding elements to be stored.

[0349] In some embodiments, the encoding module 1410 is used to encode the remaining primitives to be stored based on a second encoding method.

[0350] In some embodiments, in the structure of the second encoding method, the index information includes at least one second primitive index, the second primitive index corresponds one-to-one with the primitive to be encoded, and the position of the second primitive index of the primitive to be encoded in the index information corresponds to the position of the vertex data of the primitive to be encoded in the vertex data.

[0351] In some embodiments, the encoding module 1410 is configured to, for each thread, under the condition that the encoding conditions of the second encoding method are met, acquire a threshold number of unencoded primitives in a first direction based on the order of the primitive indexes of the second encoding method, and encode the acquired unencoded primitives based on the second encoding method to obtain a structure of the second encoding method; the encoding conditions of the second encoding method include: the primitives of the thread are not encoded; the number of unencoded primitives in a second direction based on the order of the primitive indexes is an integer multiple of the threshold number of primitives in the second encoding method; wherein, the acquired unencoded primitives are less than or equal to the threshold number of primitives in the second encoding method, and the first direction is opposite to the second direction.

[0352] In some embodiments, the encoding module 1410 is configured to: acquire the number of generated structures; determine a target storage address based on the number of generated structures, historical storage addresses, and structure sizes; and store the structures obtained using the first encoding method or the second encoding method based on the target storage address.

[0353] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this application can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0354] It should be noted that, in the embodiments of this application, if the above-described primitive encoding method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0355] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0356] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.

[0357] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0358] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0359] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0360] Figure 17 This application provides a hardware entity diagram of a computer device as an embodiment of the present application, such as... Figure 17 As shown, the hardware entity of the computer device 1500 includes a processor 1501 and a memory 1502, wherein the memory 1502 stores a computer program that can run on the processor 1501, and the processor 1501 executes the program to implement the steps in the method of any of the above embodiments.

[0361] The memory 1502 stores computer programs that can run on the processor. The memory 1502 is configured to store instructions and applications that can be executed by the processor 1501. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1501 and various modules in the computer device 1500. It can be implemented by flash memory or random access memory (RAM).

[0362] The processor 1501 executes the steps of any of the above-mentioned primitive encoding methods when executing a program. The processor 1501 typically controls the overall operation of the computer device 1500.

[0363] This application provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the primitive encoding method as described in any of the above embodiments.

[0364] It should be noted that the descriptions of the storage medium and device embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0365] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.

[0366] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0367] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0368] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0369] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0370] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0371] Furthermore, in the various embodiments of this application, all functional units can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units. Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0372] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0373] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.< / loadtriangleconstants>

Claims

1. A primitive encoding method, characterized in that, The method includes: At least two primitives to be encoded are encoded based on a first encoding method to obtain a structure of the first encoding method; the structure includes index information and vertex information. In the structure of the first encoding method, the index information includes a first primitive index and vertex indices of each vertex in the primitive to be encoded; the first primitive index is the smallest or largest primitive index among the at least two primitives to be encoded; the vertex index is used to locate the corresponding vertex data in the vertex information; the structure of the first encoding method includes a shared vertex stored once, the shared vertex is shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded are consecutive in the geometry.

2. The method according to claim 1, characterized in that, The structure includes encoding information; the encoding information is used to indicate the number of primitives and the encoding method of the primitives to be encoded in the primitive node; the vertex information includes the vertex data of each vertex in the primitive to be encoded.

3. The method according to claim 2, characterized in that, The method further includes: Among a plurality of primitives to be stored, at least two primitives that satisfy the first encoding method are determined.

4. The method according to claim 3, characterized in that, The step of determining at least two primitives to be encoded that satisfy the first encoding method from a plurality of primitives to be stored includes: Assign a corresponding thread to each of the primitives to be stored; For each thread, the graph data of the element to be stored is obtained through the thread, and the graph data of at least one other element is obtained based on the order of the element index. The target thread with the largest number of coded primitives and the at least one other primitive are determined as the at least two primitives to be encoded; the number of coded primitives is the number of primitives acquired, including the primitives to be stored and the at least one other primitive. Wherein, the primitive to be stored acquired by the thread and the at least one other primitive satisfy the encoding amount threshold of the first encoding method.

5. The method according to claim 4, characterized in that, The encoding threshold includes a primitive number threshold and a vertex number threshold; the step of obtaining the primitive data of the primitive to be stored through the thread, and obtaining the primitive data of at least one other primitive based on the order of the primitive index, includes: The thread obtains the graph data of the graph primitives to be stored, and updates the number of graph primitives to be encoded and the number of vertices to be encoded in the thread. The other primitives are obtained sequentially based on the primitive index, and the number of primitives to be encoded and the number of vertices to be encoded are updated based on the obtained primitives until the stopping condition is met. The stopping conditions include at least one of the following: After obtaining the next other graphic element, the number of updated graphic elements to be encoded does not meet the graphic element number threshold. After obtaining the next other primitive, the updated number of vertices to be encoded does not meet the vertex number threshold.

6. The method according to claim 5, characterized in that, The method further includes: In response to the satisfaction of the stopping condition, the number of primitives to be encoded for the thread is determined as the number of coded primitives for the thread.

7. The method according to claim 6, characterized in that, The encoding of at least two primitives to be encoded based on the first encoding method to obtain the structure of the first encoding method includes: Using the target thread with the largest number of encodeable elements, at least two primitives of the target thread are encoded based on the first primitive index and the first encoding method to obtain the structure of the first encoding method; the first primitive index is the primitive index of the primitive to be stored.

8. The method according to claim 6, characterized in that, The method further includes: In response to the completion of the encoding process of the at least two primitives to be encoded, the number of coded primitives in each thread is updated; Based on the updated number of coded elements for each thread, the encoding process for other primitives to be encoded that satisfy the first encoding method is executed, and the number of coded elements for each thread is updated until there is no number of coded elements that meet the requirement.

9. The method according to claim 5, characterized in that, The step of sequentially acquiring the other primitives based on the primitive index, and then updating the number of primitives to be encoded and the number of vertices to be encoded based on the acquired primitives, includes: Obtain the current element's element data from another thread; the current element is determined based on the order of the element index. If, based on the graph data of the current graph element, it is determined that there is a target vertex in the current graph element that overlaps with a candidate vertex in the candidate vertex set, the vertex index of the candidate vertex is assigned to the target vertex; the candidate vertex set is used to store candidate vertices and their corresponding vertex indices. If it is determined that the candidate vertex set can accommodate all vertices of the current primitive based on the number of vertices in the candidate vertex set, the vertex number threshold, and the number of vertices in the current primitive without assigned vertex indices, new vertex indices are generated and assigned to the vertices in the current primitive without assigned vertex indices, and the candidate vertex set, the number of primitives to be encoded, and the number of vertices to be encoded are updated.

10. The method according to claim 4, characterized in that, Each set of multiple primitives to be stored corresponds to a thread group, and multiple threads in the thread group execute the encoding process of the multiple primitives to be stored in parallel. The process of allocating a corresponding thread for each of the primitives to be stored includes: According to the order of the first primitive index of each primitive to be stored in the geometry, a thread is allocated for each primitive to be stored in the multiple threads corresponding to the plurality of primitives to be stored.

11. The method according to claim 10, characterized in that, The method further includes: Obtain the number of primitives in the geometry; Based on the number of primitives in the geometry and the preset number of threads, at least one thread group and the primitives corresponding to each thread group are determined. The thread group includes the specified number of threads, and the number of threads in the thread group is greater than or equal to the number of primitives corresponding to the thread group.

12. The method according to claim 7, characterized in that, The process of encoding at least two primitives of the target thread based on the first primitive index and the first encoding method to obtain a structure of the first encoding method includes: Based on the first encoding method, the number of encodeable primitives, the first primitive index, and the candidate vertex set, at least two primitives to be encoded in the target thread are encoded to obtain the structure of the first encoding method.

13. The method according to claim 12, characterized in that, The step of encoding at least two primitives of the target thread based on the first encoding method, the number of encodeable primitives, the first primitive index, and the candidate vertex set to obtain a structure of the first encoding method includes: The encoded information is generated based on the first encoding method and the number of coded elements. The index information is determined based on the first primitive index and the vertex indices of the candidate vertices in the candidate vertex set; The vertex information is generated based on the vertex data corresponding to the vertex index codes of the candidate vertices in the candidate vertex set.

14. The method according to claim 4, characterized in that, The method further includes: The maximum number of coded threads is determined based on the number of coded threads described above; Among at least one thread corresponding to the maximum number of coded elements, the thread with the smallest thread index is determined as the target thread; wherein the order of the thread indices of the thread is the same as or the opposite of the order of the element indices of the corresponding elements to be stored.

15. The method according to any one of claims 4 to 14, characterized in that, The method further includes: The remaining primitives to be stored are encoded based on the second encoding method.

16. The method according to claim 15, characterized in that, In the structure of the second encoding method, the index information includes at least one second primitive index, which corresponds one-to-one with the primitive to be encoded. The position of the second primitive index of the primitive to be encoded in the index information corresponds to the position of the vertex data of the primitive to be encoded in the vertex data.

17. The method according to claim 15, characterized in that, The encoding of the remaining primitives to be stored based on the second encoding method includes: For each thread, under the condition that the encoding conditions of the second encoding method are met, the number of unencoded primitives of the second encoding method is obtained in the first direction based on the order of the primitive index. The obtained unencoded primitives are then encoded based on the second encoding method to obtain the structure of the second encoding method. The encoding conditions for the second encoding method include: The primitives of the thread are not encoded; In the second direction based on the order of the primitive index, the number of uncoded primitives is an integer multiple of the primitive number threshold of the second encoding method; Wherein, the number of unencoded primitives obtained is less than or equal to the primitive number threshold of the second encoding method, and the first direction is opposite to the second direction.

18. The method according to any one of claims 1 to 14, characterized in that, The method further includes: Get the number of generated structures; The target storage address is determined based on the number of generated structures, historical storage addresses, and structure sizes. The structure obtained by storing the first encoding method or the second encoding method based on the target storage address.

19. A primitive encoding device, characterized in that, The device includes: An encoding module is used to encode at least two primitives to be encoded based on a first encoding method to obtain a structure of the first encoding method; the structure includes index information and vertex information; In the structure of the first encoding method, the index information includes a first primitive index and vertex indices of each vertex in the primitive to be encoded; the first primitive index is the smallest or largest primitive index among the at least two primitives to be encoded; the vertex index is used to locate the corresponding vertex data in the vertex information; the structure of the first encoding method includes a shared vertex stored once, the shared vertex is shared by at least two primitives to be encoded, and the primitive indices of the two primitives to be encoded are consecutive in the geometry.

20. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 18.

21. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 18.

22. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Encoding device, decoding device, encoding method, and decoding method

    CN119301642A