Prefetching Method, Device, Equipment, Storage Medium and Program Product
By dynamically adjusting the prefetch strategy based on the BVH tree's hierarchical information, ray type and cache encoding of the BVH tree during the ray tracing process, the problem of unreasonable prefetch data settings in the existing technology is solved, and the computing efficiency and resource utilization of ray tracing are improved.
Patent Information
- Application Number
- CN202410742926.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-06-07
AI Technical Summary
In the existing ray tracing technology, the data size and storage location of the prefetched data are unreasonable, which affects the computing efficiency and resource utilization of the graphics processor.
During the ray tracing process based on the hierarchical enclosure structure BVH tree, the cache prefetch size and cache location of the data that needs to be prefetched are determined by obtaining prefetched parameters (including the hierarchical information of the target node, the ray type and the cache encoding set by the user).
By dynamically adjusting the prefetch strategy, we can balance data response efficiency and cache space utilization, and improve overall computing efficiency and resource utilization.
Smart Images

Figure CN118733487B_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of data processing technologies, and in particular, to a prefetching method, apparatus, device, storage medium, and program product. Background Art
[0002] In computer graphics, ray tracing is a technique for simulating the propagation and reflection of light in a three-dimensional space, and is commonly used to achieve complex light interaction effects, such as shadows, reflections, and refractions. However, as the complexity of the scene increases, the amount of data to be processed during the ray tracing process also increases significantly, imposing a huge pressure on memory and caches. To improve the efficiency of ray tracing, researchers have proposed various prefetching strategies aimed at pre-reading the data that may be needed to reduce memory access latency and cache misses. However, in related technologies, the data size and storage location of the prefetched data are not reasonably set, affecting the computing efficiency and resource utilization rate of the graphics processor. Summary of the Invention
[0003] In view of this, embodiments of this application at least provide a prefetching method, apparatus, device, storage medium, and program product.
[0004] The technical solution of the embodiments of this application is implemented as follows:
[0005] On the one hand, embodiments of this application provide a prefetching method, the method including: during the ray tracing process of the current ray based on a hierarchical bounding volume structure BVH tree, in the case where the node information of the target node does not exist in the cache module, obtaining a prefetching parameter, and determining, based on the prefetching parameter, the cache prefetch size of the second data to be prefetched and / or the cache location of the second data in the cache module; the node information is used to determine whether the bounding box corresponding to the target node intersects with the current ray; reading first data including the node information from the main memory to the cache module, and in the case where the second data exists, prefetching the second data from the main memory to the cache module based on the cache prefetch size of the second data and / or the cache location of the second data in the cache module; wherein, the prefetching parameter includes at least one of the following: the hierarchical information of the target node in the BVH tree, the ray type of the current ray, and the cache encoding set by the user.
[0006] In some embodiments, the hierarchical information of the target node in the BVH tree is used to characterize the level of spatial locality of the node information of the target node; the spatial locality of the node information corresponding to different ray types is different; wherein, the cache prefetch size of the second data is positively correlated with the spatial locality; the storage distance is negatively correlated with the spatial locality, and the storage distance is the distance between the cache location and the processor that initiates the ray tracing.
[0007] In some embodiments, the cache module includes at least one level of cache units, and the distances between different levels of cache units and the processor are different; the hierarchical information and / or the ray type are used to determine a target cache unit for storing the second data in the at least one level of cache units, and the target cache unit is the cache location.
[0008] In some embodiments, the hierarchical information includes at least one of the following: first hierarchical information and second hierarchical information; the first hierarchical information is used to characterize the number of levels between a target node and the root node of the BVH tree, and the second hierarchical information is used to characterize the number of levels between the target node and the leaf node of the BVH tree; wherein, the number of levels included in the first hierarchical information is positively correlated with the spatial locality, and the number of levels included in the second hierarchical information is negatively correlated with the spatial locality.
[0009] In some embodiments, the ray type of the current ray includes a primary ray type, a shadow ray type, a reflection ray type, and a refraction ray type; wherein, when other prefetch parameters are the same, the cache prefetch size when the ray type of the current ray is the primary ray type or the shadow ray type is greater than the cache prefetch size when the ray type of the current ray is the reflection ray type or the refraction ray type; and / or, the distance between the cache location and the processor when the ray type of the current ray is the primary ray type or the shadow ray type is less than the distance between the cache location and the processor when the ray type of the current ray is the reflection ray type or the refraction ray type.
[0010] In some embodiments, obtaining the prefetch parameters and determining the cache prefetch size of the second data to be prefetched and / or the cache location of the second data in the cache module based on the prefetch parameters includes: querying the cache encoding set by the user for the ray tracing process; and determining the cache prefetch size and / or the cache location based on the cache encoding.
[0011] In some embodiments, obtaining the prefetch parameters and determining the cache prefetch size of the second data to be prefetched and / or the cache location of the second data in the cache module based on the prefetch parameters includes: querying the cache encoding set by the user for the ray tracing process; in the case where there is a cache encoding set by the user for the ray tracing process, determining the cache prefetch size and / or the cache location based on the cache encoding; in the case where there is no cache encoding set by the user for the ray tracing process, obtaining the hierarchical information and / or the ray type, and determining the cache prefetch size and / or the cache location based on the hierarchical information and / or the ray type.
[0012] In some embodiments, the method further includes: not prefetching when it is determined that the second data does not exist.
[0013] On the other hand, an embodiment of the present application provides a prefetching device, which includes:
[0014] An acquisition module, configured to, during the ray tracing process of the current ray based on the hierarchical bounding volume structure BVH tree, when the node information of the target node does not exist in the cache module, acquire prefetch parameters, and determine the cache prefetch size of the second data to be prefetched and / or the cache position of the second data in the cache module based on the prefetch parameters; the node information is used to determine whether the bounding box corresponding to the target node intersects with the current ray;
[0015] A read prefetch module, configured to read first data including the node information from the main memory to the cache module, and when it is determined that the second data exists based on the cache prefetch size of the second data, prefetch the second data from the main memory to the cache module based on the cache prefetch size of the second data and / or the cache position of the second data in the cache module;
[0016] Wherein, the prefetch parameters are used to determine at least one of the following: the hierarchical information of the target node in the BVH tree, the ray type of the current ray, and the cache encoding set by the user.
[0017] In yet another aspect, an embodiment of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and the processor implements some or all of the steps in the above method when executing the program.
[0018] In still another aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented.
[0019] In the embodiments of the present application, by at least one of the hierarchical information of the target node in the BVH tree, the ray type of the current ray, and the cache encoding set by the user, the cache prefetch size and cache position of the second data for each prefetch operation are reasonably adjusted. Furthermore, by dynamically adjusting the prefetch strategy for each prefetch operation, a balance can be achieved between the data response efficiency and the utilization rate of the cache space. Overall, not only can the data response efficiency be improved, but also the utilization rate of the cache space can be improved.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of the present application. Description of the Drawings
[0021] The accompanying drawings here are incorporated into the specification and form a part of this specification. These drawings illustrate embodiments consistent with this application and, together with the specification, are used to explain the technical solutions of this application.
[0022] Figure 1 It is a schematic diagram of the implementation process of a prefetching method provided for an embodiment of this application;
[0023] Figure 2 It is a schematic diagram of querying node information in a scenario where a cache module provided for an embodiment of this application includes two cache units;
[0024] Figure 3 It is a schematic diagram of a BVH tree provided for an embodiment of this application;
[0025] Figure 4A It is a schematic diagram of the implementation process of determining the cache prefetch size and / or cache location based on prefetch parameters provided for an embodiment of this application;
[0026] Figure 4B It is another schematic diagram of the implementation process of determining the cache prefetch size and / or cache location based on prefetch parameters provided for an embodiment of this application;
[0027] Figure 5 It is a schematic diagram of a cache scenario of a two - level cache provided for an embodiment of this application;
[0028] Figure 6 It is a schematic diagram of the composition structure of a prefetching device provided for an embodiment of this application;
[0029] Figure 7 It is a schematic diagram of the hardware entity of a computer device provided for an embodiment of this application. Detailed implementation manners
[0030] In order to make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be further elaborated in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations of this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0031] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It is understood that "first / second / third" may be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.
[0033] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0034] (1) A Bounding Volume Hierarchy (BVH) is a tree structure on a set of geometric objects. The leaf nodes of the tree are all geometric objects wrapped in bounding volumes. These nodes are then grouped into small sets and contained within larger bounding volumes. In turn, these are also grouped recursively and contained within other larger bounding volumes, ultimately forming a tree structure with a single bounding volume at the top of the tree. The bounding volume hierarchy is used to efficiently support various operations on a set of geometric objects, such as in collision detection and ray tracing.
[0035] (2) A cache line is a concept in computer architecture and is the smallest data unit of a cache. A cache line is a set of data stored in the cache, usually of the same size and consisting of multiple words. When the CPU accesses the main memory, the data is read from the main memory into the cache. If the size of the cache is an integer multiple of the cache line size, then the size of each cache line is one integer fraction of the cache size. The introduction of cache lines is to solve the contradiction between cache capacity and access speed. As the cache capacity increases, the access speed of the cache decreases. To improve the access speed of the cache, the cache can be divided into multiple cache lines of the same size, so that fast access can be performed between cache lines.
[0036] (3) Cache Miss: When the CPU needs to access the main memory, it first checks whether the data is in the cache. If the data is not in the cache, a Cache Miss occurs. When a cache miss occurs, the CPU needs to read the data from the main memory and store it in the cache. When reading a cache line, the CPU usually reads the entire cache line into the cache, not just the data needed. The reason for this is that other data in the cache line may be accessed in the future, so reading the entire cache line into the cache in advance can reduce future cache misses. When reading a cache line, the cache line prefetch policy also needs to be considered. The prefetch policy refers to predicting future access patterns based on the data access pattern and prefetching relevant cache lines into the cache in advance. The correct prefetch policy can reduce cache misses and improve the performance of the program.
[0037] (4) Spatial Locality: If a data is accessed, then the data adjacent to it is very likely to be accessed soon. This spatial locality is common in data structures such as arrays and linked lists, as well as consecutive memory address accesses. In a computing system, the spatial locality of the cache can help predict the data to be accessed soon, thus improving the execution efficiency of the program. Correspondingly, temporal locality means that once a data is accessed, it is very likely to be accessed again in the near future. This temporal locality is common in situations such as loop programs and recursive functions that have the same code executed repeatedly. In a computing system, the temporal locality of the cache can help predict the data to be accessed soon, thus improving the execution efficiency of the program.
[0038] (5) Ray Tracing: It is a method of presenting three-dimensional (3D) images on a two-dimensional (2D) screen. The process is to calculate effects such as reflection, refraction, and shadows in the scene by simulating the propagation of light to generate realistic images. Specifically, the ray tracing algorithm generates an image by tracing countless rays emitted from the observer's eye. These rays are shot one by one at each surface in the scene and bounce off the first surface they encounter. Then, the algorithm calculates the effects such as reflection, refraction, and shadows experienced by the rays during the bouncing process to determine the color of each pixel. Since the ray tracing algorithm requires a large amount of computing power, it usually requires high-performance computer hardware and complex algorithm optimizations. However, with the continuous development of computer technology, ray tracing has become the mainstream image rendering technology in fields such as computer games and movie production.
[0039] (6) Tree: In computer data, a tree is a non-linear data structure used to represent data with a hierarchical structure. Each node in a tree can contain data and pointers to its child nodes. The root node of a tree is the topmost node and has no parent node, while other nodes have exactly one parent node. The level of a tree refers to the number of nodes on the path from the root node to a certain node. The root node is at level 0, and its level is 0. The level of each node can be obtained by calculating the number of nodes on the path from the root node to that node. The depth of a tree refers to the maximum level of the tree. In a binary tree, the depth of the root node is 0, and the depth of each child node is equal to the depth of its parent node plus 1.
[0040] Existing prefetching strategies usually only focus on cache misses and ignore the impact of memory access latency. In fact, memory access latency has a great impact on the efficiency of ray tracing. If only cache misses are considered, it may lead to prefetching data that is not needed or having a large data access latency.
[0041] An embodiment of this application provides a prefetching method, which can be executed by a processor of a computer device. Herein, the computer device can refer to devices with data processing capabilities such as servers, laptop computers, tablet computers, desktop computers, smart TVs, set-top boxes, mobile devices (such as mobile phones, portable video players, personal digital assistants, dedicated messaging devices, portable game devices), etc.
[0042] Figure 1 It is a schematic diagram of the implementation process of a prefetching method provided by an embodiment of this application. As Figure 1 shown, this method includes the following steps S101 to S102:
[0043] Step S101: During the ray tracing process of the current ray based on the hierarchical bounding volume structure BVH tree, when the node information of the target node does not exist in the cache module, obtain the prefetching parameters, and determine the cache prefetch size of the second data to be prefetched and / or the cache location of the second data in the cache module based on the prefetching parameters; the node information is used to determine whether the bounding box corresponding to the target node intersects with the current ray.
[0044] In some embodiments, the above ray tracing process of the current ray based on the hierarchical bounding volume structure BVH tree can use the BVH tree to determine the object(s) (for now, taking the current ray as an example) that the ray collides with. Exemplarily, the ray tracing process for the current ray may include: starting the search from the root node of the BVH tree; for each node, calculating whether the current ray intersects with the bounding box of the node; if there is no intersection, the current ray will not intersect with the objects within the node, and directly skip to the next node for judgment; if the current ray intersects with the bounding box of the node, further determine whether the current ray intersects with the corresponding sub-bounding box (the bounding box corresponding to the child node of the node) within the node, that is, for each child node of the node, calculate whether the current ray intersects with the bounding box of the child node; repeat the above steps until it is determined whether the object corresponding to the leaf node of the BVH tree intersects with the current ray. Through the above steps, the BVH tree can be used to quickly determine the object that a ray collides with and record the collision information, so as to perform subsequent rendering operations.
[0045] It can be understood that the present application is actually applied to the following scenario: during the above process of traversing the nodes of the BVH tree, in order to determine whether the current ray intersects with the bounding box corresponding to the currently traversed node, the node information of the currently traversed node (hereinafter, the currently traversed node is simply referred to as the target node) is queried from the cache module; and when the node information of the target node does not exist in the cache module, while obtaining the node information of the target node from the main memory, the cache prefetch size of the second data to be prefetched and / or the cache location of the second data in the cache module are determined based on the prefetch parameter.
[0046] In some embodiments, the prefetch parameter includes at least one of the following: the hierarchical information of the target node in the BVH tree, the ray type of the current ray, and the cache encoding set by the user.
[0047] In some embodiments, the hierarchical information of the target node in the BVH tree is used to characterize the level of spatial locality of the node information of the target node; the spatial locality of the node information corresponding to different ray types is different.
[0048] Among them, the hierarchical information represents the position of the target node in the above BVH tree. It can be understood that different positions of the target node in the BVH tree result in different corresponding spatial locality; for example, the closer the target node is to the leaf node of the BVH tree, due to the closer spatial distance, correspondingly, the higher the corresponding spatial locality; conversely, the closer the target node is to the root node of the BVH tree, due to the farther spatial distance, correspondingly, the lower the corresponding spatial locality. Therefore, the hierarchical information of the target node in the BVH tree can be used as the prefetch parameter.
[0049] Among them, the current rays of different ray types also have different corresponding spatial locality. Usually, during the ray tracing process, the types of rays can include: primary ray, shadow ray, reflection ray, and refracted ray. Among them, the primary ray is the ray emitted from the camera; the shadow ray is the ray emitted from the intersection point and pointing to the light source. If this ray does not intersect any object before pointing to the light source, then this light source has a contribution value to this intersection point; otherwise, this intersection point is in the shadow of this light source; for the reflection ray, if the object surface has the property of reflection, then part of the light will be reflected and continue to move forward in the scene. According to Snell's law, a new ray will be emitted from the intersection point; for the refracted ray, when the object surface has the property of refraction and is partially transparent, part of the light will enter the object and continue to propagate. According to Snell's law, a new ray will be emitted from the intersection point and enter the object. It can be seen that both the primary ray and the shadow ray point to the same object (the camera and the light source), and their changes in space are orderly, that is, the spatial locality is relatively good, while the reflection ray and the refracted ray are affected by the incident angle and the medium, and can be considered disorderly in space, that is, the spatial locality is relatively poor. Therefore, the ray type of the current ray can also be used as the prefetch parameter.
[0050] It can be understood that the cache prefetch size of the second data is positively correlated with the spatial locality; the storage distance is negatively correlated with the spatial locality, and the storage distance is the distance between the cache location and the processor that initiates the ray tracing.
[0051] In an embodiment of the present application, the higher the spatial locality determined based on the level information of the target node in the BVH tree and / or the ray type of the current ray, the higher the correlation between the second data that can be obtained and the first data, and therefore, the larger the cache prefetch size, and accordingly, the shorter the cache distance; on the contrary, the lower the spatial locality determined based on the level information of the target node in the BVH tree and / or the ray type of the current ray, the lower the correlation between the second data that can be obtained and the first data, and therefore, the smaller the cache prefetch size, and accordingly, the longer the cache distance. In other words, the cache prefetch size of the second data is positively correlated with the spatial locality; the storage distance is negatively correlated with the spatial locality, and the storage distance is the distance between the cache location and the processor that initiates ray tracing.
[0052] In some embodiments, the ray type of the current ray includes a principal ray type, a shadow line type, a reflection line type and a refraction line type; wherein, when other prefetch parameters are the same, the cache prefetch size when the ray type of the current ray is the principal ray type or the shadow line type is larger than the cache prefetch size when the ray type of the current ray is the reflection line type or the refraction line type; and / or, when the ray type of the current ray is the principal ray type or the shadow line type, the distance between the cache position and the processor is smaller than the distance between the cache position and the processor when the ray type of the current ray is the reflection line type or the refraction line type.
[0053] In the embodiment of the present application, by obtaining the ray type of the current ray, since different ray types correspond to different spatial localities, a more suitable cache prefetch size and / or cache position can be obtained by considering the ray type of the current ray.
[0054] Among them, the cache encoding set by the user is the encoding information pre-set by the user based on the current graphics processor's operating scenario and the ray tracing algorithm. Different encoding information can correspond to different cache prefetch sizes, and different encoding information can correspond to different cache locations; different encoding information can also correspond to different cache prefetch sizes and different cache locations.
[0055] In some embodiments, when the cache prefetch size of the second data and / or the cache location of the second data in the cache module are determined based on at least two of the above prefetch parameters, for example, when the cache prefetch size is determined based on at least two of the above prefetch parameters, embodiments of the present application can set corresponding weights for each prefetch parameter, determine the corresponding cache prefetch sizes respectively based on the prefetch parameters, and then, based on the weights corresponding to the prefetch parameters, perform weighted averaging on the cache prefetch sizes determined by the prefetch parameters to obtain the final cache prefetch size. Exemplarily, if the current cache prefetch size is determined based on the hierarchical information of the target node in the BVH tree, the ray type of the current ray, and the cache encoding set by the user, where the cache prefetch size determined based on the hierarchical information is 300 bytes, the cache prefetch size determined based on the ray type is 400 bytes, and the cache prefetch size determined based on the cache encoding is 200 bytes, when the weights corresponding to the hierarchical information, ray type, and cache encoding are (2, 2, 4), the weighted average value can be obtained as 275 bytes, and then the final cache prefetch size can be set to 275 bytes, or set to 300 bytes closest to 275 bytes.
[0056] Step S102: Read the first data including the node information from the main memory to the cache module, and when there is the second data, prefetch the second data from the main memory to the cache module based on the cache prefetch size of the second data and / or the cache location of the second data in the cache module.
[0057] In embodiments of the present application, since the current scenario is that the node information of the target node is missing in the cache module, therefore, at least the node information of the target node needs to be read from the main memory to the cache module to facilitate the subsequent node traversal process in the above ray tracing process. At the same time, after reading the node information of the target node into the cache module, it is possible to determine whether there is the second data based on the cache prefetch size of the second data determined to be prefetched, or it can also be determined whether there is the second data based on the cache encoding; when there is the second data, prefetch the second data from the main memory to the cache module based on the cache prefetch size of the second data and / or the cache location of the second data in the cache module.
[0058] In some embodiments, when there is no such second data, no prefetching is performed.
[0059] Among them, when it is determined based on cache encoding that the second data does not exist, prefetching may not be performed; when it is determined based on the cache prefetch size of the second data that prefetching is required and it is determined that the second data does not exist, prefetching may not be performed; when it is determined based on cache encoding that the second data does not exist and it is determined based on the cache prefetch size of the second data that prefetching is required and it is determined that the second data does not exist, prefetching may not be performed.
[0060] In some embodiments, the first data at least includes all node information of the target node. That is to say, when the node information of the target node is 128 bytes, the data size of the first data is at least 128 bytes. In other embodiments, in addition to including the node information of the target node, the first data may further include other information, such as node information of other nodes, etc.
[0061] In some implementation scenarios, during the process of reading data from the main memory to the cache module, the data is stored in the cache module in units of cache lines. Therefore, when the data size of the node information of the target node is an integer multiple of the size of a cache line, the above-mentioned first data can be an integer multiple of the size of a cache line. In this way, the first data (node information of the target node) can be read from the main memory and cached in an integer number of cache lines in the cache module. Correspondingly, when the second data needs to be prefetched, the second data is also cached in the cache module in the form of the above-mentioned cache lines.
[0062] Among them, the cache prefetch size of the second data can be in units of data size, that is, in bytes; the cache prefetch size of the second data can be in units of the smallest data unit in the cache, that is, in units of cache lines; the cache prefetch size of the second data can be in units of the size of the node information of each node, that is, in units of the number of nodes. For example, the node information of N nodes is prefetched into the cache module as the second data.
[0063] In some embodiments, the cache module includes at least one level of cache units, and the distances between different levels of cache units and the processor are different; the hierarchical information and / or the light type are used to determine a target cache unit for storing the second data in the at least one level of cache units, and the target cache unit is the cache location.
[0064] Among them, the cache module may include at least one level of cache units, and the positions of cache units at different levels in the cache module are different, that is, the distances between different cache units and the processor are different. Among them, if the cache module includes two cache units, namely a first-level cache unit and a second-level cache unit, where the distance between the first-level cache unit and the processor is less than the distance between the second-level cache unit and the processor, the above prefetch parameter is used to determine a target cache unit among these two levels of cache units and cache the second data into the target cache unit.
[0065] In the embodiments of the present application, a target cache unit can be determined among these two levels of cache units based on the prefetch parameter, and the second data is cached into the target cache unit. In this way, since the hierarchical information of the target node in the BVH tree and / or the ray type of the current ray reflect the level of spatial locality between the first data and the second data, the cache position of the second data for each prefetch operation can be reasonably adjusted. When the spatial locality is relatively high, the second data can be cached into a cache unit closer to the processor that initiates ray tracing, which can improve the data response efficiency; and when the spatial locality is relatively low, the second data can be cached into a cache unit farther from the processor that initiates ray tracing, which can reduce the unnecessary occupation of high-performance cache space.
[0066] It can be understood that the level of the cache unit in the above cache module is used to represent the access priority of the cache unit. The process of querying the node information of the target node in the cache module in step S101 is a process of sequentially traversing each cache unit according to the levels of each cache unit in the cache module. Correspondingly, in some embodiments, the method further includes: sequentially searching for the node information of the target node in the at least one level of cache units according to the level; and determining that the node information of the target node does not exist in the cache module when the node information of the target node cannot be found in any of the at least one level of cache units.
[0067] Please refer to Figure 2 , which shows a schematic diagram of querying node information in a scenario where a cache module includes two cache units. Among them, the cache module 210 includes a first cache unit 211 and a second cache unit 212, and the level of the first cache unit 211 is lower than that of the second cache unit 212, that is, the distance between the first cache unit 211 and the processor that initiates ray tracing is less than the distance between the second cache unit 212 and the processor that initiates ray tracing. It should be understood that the cache module in the present application may also include other numbers of cache units, and the embodiments of the present application do not limit this.
[0068] In some implementation scenarios, the ray tracing module 220 sends a node information query request to the cache module 210. In response to the node information query request, the cache module 210 sequentially searches for the node information of the target node in the cache units of the at least one level in the order of the levels (from small to large), that is, first queries the node information of the target node in the first cache unit 211. If the node information is found, the node information of the target node is fed back to the ray tracing module 220; if not found, the node information of the target node is queried in the second cache unit 212. If the node information is found, the node information of the target node is fed back to the ray tracing module 220; if not found, it is determined that the cache module does not have the node information of the target node, step S101 is executed to obtain the prefetch parameter, and based on the prefetch parameter, the cache prefetch size of the second data to be prefetched and / or the cache location of the second data in the cache module are determined.
[0069] In the above embodiments, taking the determination of the target cache unit and the determination of the prefetch cache size as examples respectively, the influence of the above prefetch parameter on the prefetch policy is described. In some implementation scenarios, the prefetch parameter can not only affect the target cache unit where the second data is stored, but also affect the prefetch cache size of the second data. Exemplarily, there may be a first prefetch policy, a second prefetch policy, and a third prefetch policy. Among them, the first prefetch policy is used to prefetch 128 bytes of the second data to the second cache unit above, the second prefetch policy is used to prefetch 256 bytes of the second data to the second cache unit above, and the third prefetch policy is used to prefetch 256 bytes of the second data to the first cache unit above. That is to say, based on the above prefetch parameter, the current prefetch policy can be determined among multiple preset prefetch policies.
[0070] In the embodiments of the present application, the level of spatial locality is estimated through the hierarchical information of the target node in the BVH tree and / or the ray type of the current ray, and then the level of spatial locality between the first data to be read and the second data to be read is determined. Therefore, the cache prefetch size and cache location of the second data for each prefetch operation can be reasonably adjusted through the hierarchical information of the target node in the BVH tree and / or the ray type of the current ray; thus, by dynamically adjusting the prefetch policy for each prefetch operation, a balance can be achieved between the data response efficiency and the utilization rate of the cache space. Overall, not only can the data response efficiency be improved, but also the utilization rate of the cache space can be improved.
[0071] In some implementation scenarios, in order to provide a more fine-grained prefetch control scheme, a prefetch policy can be used to indicate caching the second data into cache units at different levels. Refer to the following embodiments. In some embodiments, the cache module includes at least one level of cache units; the hierarchical information of the target node in the BVH tree and / or the ray type of the current ray are used to determine at least one sub-cache unit for storing the second data and the corresponding sub-prefetch size in the at least one cache unit; wherein, the level of at least one of the sub-cache units is negatively correlated with the spatial locality, and / or, the sub-prefetch size corresponding to at least one of the sub-cache units is positively correlated with the spatial locality.
[0072] The following embodiments will illustrate the case where the level of at least one of the sub-cache units is negatively correlated with the spatial locality.
[0073] In the embodiments of the present application, the hierarchical information of the target node in the BVH tree and / or the ray type of the current ray are used to determine at least one cache unit for storing the second data in the above at least one cache unit.
[0074] If there are a first prefetch policy and a second prefetch policy, wherein the spatial locality determined by the hierarchical information and / or ray type corresponding to the first prefetch policy is less than the spatial locality determined by the hierarchical information and / or ray type corresponding to the second prefetch policy; then, the combined level of at least one sub-cache unit corresponding to the first prefetch policy is less than the combined level of at least one sub-cache unit corresponding to the second prefetch policy. Wherein, the combined level can be determined based on the level of each sub-cache unit and the sub-prefetch size corresponding to each sub-cache unit.
[0075] Exemplarily, if there are a first prefetching policy (the sub-prefetch size of the second cache unit is 256 bytes), a second prefetching policy (the sub-prefetch size of the second cache unit is 128 bytes, and the sub-prefetch size of the first cache unit is 128 bytes), and a third prefetching policy (the sub-prefetch size of the first cache unit is 256 bytes). Based on the principle that the larger the sub-prefetch size, the smaller the corresponding weight, and the smaller the sub-prefetch size, the larger the corresponding weight, 256 bytes can be quantified as the weight "1", correspondingly, 128 bytes can be quantified as the weight "2", the level of the first cache unit can be quantified as "1", and the level of the second cache unit can be quantified as "2". Then, it can be determined that the level corresponding to the first prefetching policy is "2", the comprehensive level corresponding to the second prefetching policy is "6", and the comprehensive level corresponding to the third prefetching policy is "1". It can be seen that the lower the above comprehensive level, the more second data can actually be prefetched to the cache unit closer to the processor. Based on the above concept, in order to set a more fine-grained prefetching policy division, different proportions of sub-prefetch sizes (such as 32 bytes, 64 bytes, 128 bytes, etc.) can also be set for the first cache unit and the second cache unit, so as to achieve a more fine-grained prefetching policy division under the condition that the same cache prefetch size remains unchanged.
[0076] In the embodiments of the present application, under the condition that the cache prefetch size remains unchanged, by setting different proportions of sub-prefetch sizes in different cache units, a more fine-grained prefetching policy can be provided, and thus a corresponding prefetching policy can be provided for a more complex scenario.
[0077] The following embodiments will illustrate that the sub-prefetch size corresponding to at least one of the sub-cache units is positively correlated with spatial locality.
[0078] In the embodiments of the present application, the level information of the target node in the BVH tree and / or the ray type of the current ray are used to determine the sub-prefetch size corresponding to at least one of the sub-cache units.
[0079] If there are a first prefetching policy and a second prefetching policy, where the spatial locality determined by the level information and / or ray type corresponding to the first prefetching policy is less than the spatial locality determined by the level information and / or ray type corresponding to the second prefetching policy; then, the comprehensive sub-prefetch size of at least one sub-cache unit corresponding to the first prefetching policy is greater than the comprehensive sub-prefetch size of at least one sub-cache unit corresponding to the second prefetching policy.
[0080] Exemplarily, if there are three preset levels, the sub-prefetch sizes corresponding to the first prefetch policy to the third prefetch policy can be set as follows: the first prefetch policy (the sub-prefetch size of the second cache unit is 128 bytes, and the sub-prefetch size of the first cache unit is 128 bytes), the second prefetch policy (the sub-prefetch size of the second cache unit is 256 bytes, and the sub-prefetch size of the first cache unit is 128 bytes), the third prefetch policy (the sub-prefetch size of the second cache unit is 256 bytes, and the sub-prefetch size of the first cache unit is 256 bytes). Based on the principle that the smaller the level of the cache unit, the greater the corresponding weight, and the larger the level of the cache unit, the smaller the corresponding weight, the first cache unit can be set with a weight of "2", and correspondingly, the second cache unit can be set with a weight of "1". Then, the comprehensive sub-prefetch size corresponding to the first prefetch policy can be determined as "384 bytes", the comprehensive sub-prefetch size corresponding to the second prefetch policy can be determined as "512 bytes", and the comprehensive sub-prefetch size corresponding to the third prefetch policy can be determined as "768 bytes". It can be seen that the larger the above sub-prefetch size, the more second data can actually be prefetched to the cache unit closer to the processor.
[0081] In the embodiments of the present application, the present application can, under the condition that the cache location remains unchanged (cached to the same cache unit), set different sub-prefetch sizes for a prefetch policy in different cache units, thereby providing a more fine-grained prefetch policy, and further providing a corresponding prefetch policy for more complex scenarios.
[0082] In other implementation scenarios, while the level of at least one of the sub-cache units is negatively correlated with the spatial locality, the sub-prefetch size corresponding to at least one of the sub-cache units is positively correlated with the spatial locality. The present application does not limit this.
[0083] In some embodiments, the level information includes at least one of the following: first level information and second level information; the first level information is used to represent the number of levels between the target node and the root node of the BVH tree, and the second level information is used to represent the number of levels between the target node and the leaf node of the BVH tree; wherein, the number of levels included in the first level information is positively correlated with the spatial locality, and the number of levels included in the second level information is negatively correlated with the spatial locality.
[0084] In some embodiments, the first level information can be directly used as the level information; the second level information can also be directly used as the level information; or the comprehensive level information can be determined based on the first level information and the second level information to be used as the level information of the target node in the BVH tree.
[0085] Please refer to Figure 3Schematic diagram of the BVH tree shown. Among them, if the first-level information is used as the level information of this level, the first-level information is used to represent the number of levels between the target node and the root node of the BVH tree. The first-level information corresponding to each node (the level information of the node in the BVH tree) is shown in Table 1.
[0086] Table 1
[0087] Node First-level information Node A 0 Node B 1 Node C 1 Node D 2 Node E 2 Node F 3 Node G 3 Node H 3 Node I 3 Node J 2 Node K 2
[0088] Among them, if the second-level information is used as the level information of this level, the second-level information is used to represent the number of levels between the target node and the leaf node of the BVH tree. The second-level information corresponding to each node (the level information of the node in the BVH tree) is shown in Table 2.
[0089] Table 2
[0090] Node Second-level information Node A 2 Node B 2 Node C 1 Node D 1 Node E 1 Node F 0 Node G 0 Node H 0 Node I 0 Node J 0 Node K 0
[0091] Among them, if the comprehensive level information determined by the first-level information and the second-level information at the same time is used as the level information of this level, since the number of levels included in the first-level information is positively correlated with the spatial locality, and the number of levels included in the second-level information is negatively correlated with the spatial locality, the difference obtained by subtracting the second-level information from the first-level information can be used as the comprehensive level information, and the comprehensive level information is positively correlated with the spatial locality. The comprehensive level information corresponding to each node (the level information of the node in the BVH tree) is shown in Table 3.
[0092] Table 3
[0093] Node First-level information Second-level information Comprehensive-level information Node A 0 2 -2 Node B 1 2 -1 Node C 1 1 0 Node D 2 1 1 Node E 2 1 1 Node F 3 0 3 Node G 3 0 3 Node H 3 0 3 Node I 3 0 3 Node J 2 0 2 Node K 2 0 2
[0094] It can be seen that in Table 3, compared with Table 1 and Table 2, by integrating the first-level information and the second-level information, "nodes F to I" and "nodes J and K" can be effectively distinguished; the differences in the positions of nodes A, B, and C in the BVH tree can also be effectively distinguished.
[0095] In the embodiments of the present application, by obtaining the first-level information representing the number of levels between the target node and the root node of the BVH tree, and / or the second-level information representing the number of levels between the target node and the leaf node of the BVH tree, and further using it as the level information of the target node, the position of the target node in the BVH tree can be accurately represented, and then the obtained level information can accurately reflect the level of spatial locality, providing a data basis for determining the cache prefetch size of the second data and / or the cache position of the second data in the cache module.
[0096] The following will be described separately throughFigure 4A and Figure 4B respectively illustrate the solutions for determining the cache prefetch size and / or cache location based on cache encoding.
[0097] Figure 4A is a schematic diagram of an implementation process for determining the cache prefetch size and / or cache location based on prefetch parameters provided by an embodiment of this application. This method can be executed by a processor of a computer device. Based on Figure 1 , Figure 1 S101 in Figure 4A can be updated to S401 to S402, and will be described in combination with the steps shown in
[0098] Step S401, during the ray tracing process of determining the collision information of the current ray based on the hierarchical bounding volume structure BVH tree, when the node information of the target node does not exist in the cache module, query the cache encoding set by the user for the ray tracing process.
[0099] Among them, the cache encoding set by the user for the ray tracing process is set by the user in advance based on the scene information of the ray tracing. In the embodiments of this application, the above-mentioned scene information of the ray tracing can be the object information and light source information that need to be rendered in the current rendering scene, etc. Different numbers / surface materials of objects and different positions / intensities of light sources correspond to different rendering scenes. Therefore, the user can set different cache encodings for different scenes, and then set the corresponding cache prefetch size and / or cache location for different scenes through this cache encoding; the above-mentioned scene information of the ray tracing can also be the mode scene of the current computer device, for example, the performance mode or the power saving mode, etc. For the performance mode, a larger cache prefetch size and / or a cache location closer to the processor can be set, and for the power saving mode, a smaller cache prefetch size and / or a cache location farther from the processor can be set. By setting different cache encodings for different scene information and then adopting different prefetch strategies, data prefetch operations can be performed with different cache prefetch sizes and / or cache locations in different rendering scenes / device mode scenes.
[0100] In the embodiments of this application, the above-mentioned algorithm information may include at least one of the following: the traversal algorithm information for traversing the BVH tree, the ray generation algorithm for the current ray during the ray tracing process, the calculation algorithm for calculating whether there is an intersection between the current ray and the bounding box corresponding to the target node during the ray tracing process, and the algorithm for determining that the current ray finally intersects with the geometric object in the scene, etc.
[0101] In some embodiments, the user sets different cache encodings for different scenario information and / or algorithm information. Based on the scenario information and algorithm information in the current ray tracing process, the cache encoding corresponding to the current ray tracing process can be determined from the preset cache encodings.
[0102] Exemplarily, if the scenario of the current ray tracing process is determined by the scenario information and algorithm information, the process of querying the cache encoding set by the user for the ray tracing process based on the scenario information and / or algorithm information of the ray tracing process can be determined by the relationship table shown in Table 4.
[0103] Table 4
[0104] Scenario information Algorithm information Cache encoding Scenario information A Algorithm information A 00 Scenario information B Algorithm information A 00 Scenario information C Algorithm information A 00 Scenario information D Algorithm information A 00 Scenario information A Algorithm information B 01 Scenario information B Algorithm information B 01 Scenario information C Algorithm information B 10 Scenario information D Algorithm information B 10 Scenario information A Algorithm information C 11 Scenario information B Algorithm information C 11 Scenario information C Algorithm information C None Scenario information D Algorithm information C None
[0105] It can be seen that a combination of one piece of scenario information and one piece of algorithm information corresponds to at most one cache encoding. That is to say, a cache encoding can be determined based on the scenario information and algorithm information of the current ray tracing process.
[0106] In some embodiments, a combination of one piece of scenario information and one piece of algorithm information may not have a corresponding cache encoding; such as "Scenario Information C and Algorithm Information C" and "Scenario Information D and Algorithm Information C" in Table 1 above, and there is no corresponding cache encoding for these two scenarios of the ray tracing process.
[0107] Step S402: Determine the cache prefetch size and / or cache location based on the cache encoding.
[0108] In some embodiments, the cache prefetch size and / or cache location corresponding to the above cache encoding can be stored in a mapping relationship table.
[0109] Exemplarily, taking the cache encoding including 4 and different preset levels corresponding to different information encodings as an example, the above mapping relationship table can be presented in the form of Table 5. It can be understood that one information encoding corresponds to only one preset level, and one preset level can correspond to one or more information encodings, and the present application does not make any limitations in this regard.
[0110] Table 5
[0111]
[0112] Exemplarily, in the case of determining that the information encoding is "01" based on Table 4 above, based on this Table 5, it can be determined that the cache prefetch size and cache location matching the information encoding "01" are "128 byte stored in the second cache unit".
[0113] Figure 4BFIG. 0 is a schematic implementation flowchart of determining cache prefetch size and / or cache location based on prefetch parameters provided by an embodiment of the present application. This method can be executed by a processor of a computer device. Based on Figure 1 , Figure 1 S101 in [] can be updated to S403 to S405, which will be described in conjunction with the steps shown in Figure 4B .
[0114] Step S403: Query the cache encoding set by the user for the ray tracing process.
[0115] Step S404: When there is a cache encoding set by the user for the ray tracing process, determine the cache prefetch size and / or cache location based on the cache encoding.
[0116] Among them, the specific implementation manner of the above step S403 can refer to the description of step S401 in the foregoing embodiment. Correspondingly, when there is a cache encoding set by the user for the ray tracing process, the specific implementation manner of "determining the cache prefetch size and / or cache location based on the cache encoding" can refer to the description of step S402 in the foregoing embodiment. It will not be repeated here.
[0117] Step S405: When there is no cache encoding set by the user for the ray tracing process, obtain the hierarchical information and / or the ray type, and determine the cache prefetch size and / or cache location based on the hierarchical information and / or the ray type.
[0118] In the current embodiment, if the user does not set a corresponding information encoding for the current ray tracing process (that is, there is no cache encoding set by the user for the ray tracing process), it is necessary to obtain the hierarchical information and / or the ray type, and determine the cache prefetch size and / or cache location based on the hierarchical information and / or the ray type.
[0119] Among them, the hierarchical information represents the position of the target node in the above BVH tree. It can be understood that different positions of the target node in the BVH tree result in different corresponding spatial locality of the target node. Specifically, the closer the target node is to the leaf node of the BVH tree, due to the closer spatial distance, correspondingly, the higher the corresponding spatial locality; conversely, the closer the target node is to the root node of the BVH tree, due to the farther spatial distance, correspondingly, the lower the corresponding spatial locality. For the current rays of different ray types, their corresponding spatial locality is also different. Usually, during the ray tracing process, the types of rays can include: primary rays, shadow rays, reflection rays, and refraction rays. The primary rays and shadow rays both point to the same objects (camera and light source), and they change orderly in space, that is, the spatial locality is relatively good, while the reflection rays and refraction rays are affected by the incident angle and the medium, and can be considered disordered in space, that is, the spatial locality is relatively poor.
[0120] The following describes the application of the prefetching method provided by the embodiments of the present application in an actual scenario.
[0121] In one implementation scenario, the cache prefetch size (including several cache lines) and the cache level (destination level, indicating which level of cache to prefetch to, corresponding to the cache position in the above embodiments) can be adjusted according to the level of the hierarchical bounding volume structure tree (BVH tree level, corresponding to the first hierarchical information in the above embodiments).
[0122] Among them, in the normal ray tracing process, when accessing the BVH tree, the closer to the leaf node, due to the closer spatial distance, the relatively better the spatial locality. Therefore, a larger cache prefetch size (that is, read several more cache lines) can be set when a cache miss occurs, and prefetch to a cache (unit) closer to the module that initiates the ray tracing (corresponding to the processor that initiates the ray tracing in the above embodiments).
[0123] In another implementation scenario, the cache prefetch size and the cache level can be adjusted according to the type of ray during the ray tracing process.
[0124] Among them, usually, the spatial locality of different types of rays is also different. The types of rays include but are not limited to: primary rays, shadow rays, reflection rays, and refraction rays.
[0125] Among them, since the primary rays are all emitted from the camera perspective, the spatial locality of the primary rays passing through adjacent pixels is relatively better; the shadow rays directed to the same light source also have good spatial locality; relatively, the reflected rays and refracted rays are affected by the incident angle and the medium and can be considered spatially disordered.
[0126] Based on the above considerations, in the embodiments of the present application, according to the type of light rays, a relatively large cache prefetch size can be set for the primary rays and shadow rays, but a relatively small cache prefetch size can be set for the reflected rays and refracted rays or no prefetch is performed.
[0127] In some other implementation scenarios, the above cache prefetch size and cache level can be determined based on the cache hint provided by the user (corresponding to the cache encoding in the above embodiments).
[0128] Among them, the user can actively provide the corresponding cache prefetch size and cache level according to different rendering scenarios and algorithms.
[0129] The above are the methods for determining the cache prefetch size and cache level in various implementation scenarios. To facilitate the understanding of the above implementation scenarios, the above implementation scenarios will be respectively described exemplarily below.
[0130] In some implementation scenarios, for the levels of the hierarchical bounding volume structure tree, during the process of constructing the hierarchical bounding volume structure tree, the level information of each node in the hierarchical bounding volume structure tree can be directly added to the node information. Among them, the level information of the hierarchical bounding volume structure tree can include first-level information, and the first-level information indicates the level of the node from the root node; the level information of the hierarchical bounding volume structure tree can include second-level information, and the second-level information indicates the level of the node from the leaf node (the second-level information can be called LevelTo Leaf, abbreviated as LTL).
[0131] Please refer to Figure 3 , the hierarchical bounding volume structure tree includes nodes A to K, where node A is the root node and nodes F to K are leaf nodes.
[0132] Among them, when the level information includes the first-level information, the first-level information corresponding to node A is 0; the first-level information corresponding to nodes B and C is 1; the first-level information corresponding to nodes D, E, J, and K is 2; the first-level information corresponding to nodes F, G, H, and I is 3. When the level information includes the second-level information, the second-level information corresponding to nodes F, G, H, I, J, and K is 0; the second-level information corresponding to nodes C, D, and E is 1; the second-level information corresponding to nodes A and B is 2.
[0133] In hardware, there are usually multiple levels of caches to cache the node information of each node. In traditional designs, each cache miss requires reading a cache line of a fixed length. Please refer to Figure 5 , Figure 5 which shows a schematic diagram of the cache scenario of a two-level cache. Among them, this hierarchical bounding volume structure tree is stored in the main memory. During the ray tracing process, that is, during the process of determining whether a ray to be detected intersects with an object, each node in this hierarchical bounding volume structure tree will be traversed in sequence. Thus, it is necessary to cache the node information of each node of this hierarchical bounding volume structure tree from the main memory. As Figure 5 shown, the node information of the nodes of this hierarchical bounding volume structure tree can be cached in the first-level cache and the second-level cache (corresponding to at least one level of cache units in the above embodiments); in this way, when traversing to the target node during the ray tracing process, it can be first determined whether the node information of the target node exists in the first-level cache; if the node information of the target node does not exist in the first-level cache, it is determined whether the node information of the target node exists in the second-level cache; if the node information of the target node does not exist in the second-level cache, there is a cache miss, and the cache line corresponding to the node information of the target node is read from the main memory into the cache. In the above process, if the node information of the target node exists in the first-level cache, based on the node information of the target node, it is determined whether the ray intersects with the bounding box corresponding to the target node, and then the next traversed node is determined. It can be understood that Figure 5 the cache scenario in
[0134] is only one implementation scenario of this application. The embodiments of this application can also be applied to other multi-level cache scenarios such as three-level caches.
[0135] In some embodiments, the size threshold of the above BVH node, the cache prefetch size, and the cache level can all be made hardware-configurable to increase flexibility.
[0136] In other implementation scenarios, for the ray type, the module that initiates ray tracing needs to pass the ray type information to the downstream L1 cache. The L1 cache can determine the cache prefetch size and cache level based on the ray type information.
[0137] In some embodiments, the ray type may include: main ray, shadow ray, reflection ray and refraction ray, a total of 4 ray types, which can be encoded by 2 bits to obtain a 2-bit type code. The module that initiates ray tracing needs to pass the 2-bit type code to the downstream L1 cache. The L1 cache can determine the cache prefetch size and cache level based on the 2-bit type code.
[0138] In some other implementation scenarios, the cache hint provided by the user may be encoded to obtain a cache code, and then the hardware (cache) may execute corresponding hardware behavior based on the cache code.
[0139] In order to facilitate understanding of the above implementation scenarios, the following uses 4 types of cache prompt information, namely 2-bit encoding, as an example for explanation. Of course, the present application can also be applied to implementation scenarios of other numbers of cache prompt information. As shown in Table 6, different cache encodings correspond to different hardware behaviors.
[0140] Table 6
[0141]
[0142] Based on the above embodiments, in ray tracing, a lower hardware cost is used in exchange for higher cache space locality to improve cache performance. Through the node information of the BVH tree and the type of light, the cache module can adaptively adjust the pre-fetch size and the level of cache unit to be pre-fetched.
[0143] In an embodiment of the present application, the cache can adjust the size of memory access or prefetch to a certain level of cache according to node information of the BVH tree, such as the current level or the level from the leaf; the cache can also adjust the size of memory access or prefetch to a certain level of cache according to the type of light; the cache can also actively provide relevant information for prefetching based on the cache hint provided by the user; in addition, the size of memory access, prefetch to a certain level, and the corresponding BVH threshold can all be hardware configurable; there is one or more levels of cache in the ray tracing operation; the module that initiates ray tracing can be a CPU / GPU / dedicated hardware module.
[0144] Based on the foregoing embodiments, an embodiment of the present application provides a prefetching device. The device includes each unit included therein, as well as each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0145] Figure 6 It is a schematic structural diagram of a prefetching device provided by an embodiment of the present application. As Figure 6 shown, the prefetching device 600 includes: an acquisition module 610, a read prefetch module 620, where:
[0146] The acquisition module 610 is configured to, during the ray tracing process of the current ray based on the hierarchical bounding volume structure BVH tree, when the node information of the target node does not exist in the cache module, acquire prefetch parameters, and determine the cache prefetch size of the second data to be prefetched and / or the cache location of the second data in the cache module based on the prefetch parameters; the node information is used to determine whether the bounding box corresponding to the target node intersects the current ray.
[0147] The read prefetch module 620 is configured to read the first data including the node information from the main memory to the cache module, and when the second data exists, prefetch the second data from the main memory to the cache module based on the cache prefetch size of the second data and / or the cache location of the second data in the cache module; wherein, the prefetch parameters include at least one of the following: the hierarchical information of the target node in the BVH tree, the ray type of the current ray, and the cache encoding set by the user.
[0148] In some embodiments, the hierarchical information of the target node in the BVH tree is used to characterize the level of spatial locality of the node information of the target node; the spatial locality of the node information corresponding to different ray types is different; wherein, the cache prefetch size of the second data is positively correlated with the spatial locality; the storage distance is negatively correlated with the spatial locality, and the storage distance is the distance between the cache location and the processor that initiates the ray tracing.
[0149] In some embodiments, the cache module includes at least one level of cache units, and the distances between different levels of cache units and the processor are different; the hierarchical information and / or the ray type are used to determine a target cache unit for storing the second data in the at least one level of cache units, and the target cache unit is the cache location.
[0150] In some embodiments, the hierarchical information includes at least one of the following: first hierarchical information and second hierarchical information; the first hierarchical information is used to represent the number of levels between a target node and the root node of the BVH tree, and the second hierarchical information is used to represent the number of levels between the target node and the leaf node of the BVH tree; wherein, the number of levels included in the first hierarchical information is positively correlated with the spatial locality, and the number of levels included in the second hierarchical information is negatively correlated with the spatial locality.
[0151] In some embodiments, the ray type of the current ray includes a primary ray type, a shadow ray type, a reflected ray type, and a refracted ray type; wherein, when other prefetch parameters are the same, the cache prefetch size when the ray type of the current ray is the primary ray type or the shadow ray type is greater than the cache prefetch size when the ray type of the current ray is the reflected ray type or the refracted ray type; and / or, the distance between the cache location and the processor when the ray type of the current ray is the primary ray type or the shadow ray type is less than the distance between the cache location and the processor when the ray type of the current ray is the reflected ray type or the refracted ray type.
[0152] In some embodiments, the obtaining module 610 is further configured to: query the cache encoding set by the user for the ray tracing process; and determine the cache prefetch size and / or the cache location based on the cache encoding.
[0153] In some embodiments, the obtaining module 610 is further configured to: query the cache encoding set by the user for the ray tracing process; in the case where there is a cache encoding set by the user for the ray tracing process, determine the cache prefetch size and / or the cache location based on the cache encoding; in the case where there is no cache encoding set by the user for the ray tracing process, obtain the hierarchical information and / or the ray type, and determine the cache prefetch size and / or the cache location based on the hierarchical information and / or the ray type.
[0154] In some embodiments, the read prefetch module 620 is further configured to not perform prefetching when it is determined that the second data does not exist.
[0155] The description of the above device embodiments is similar to that of the above method embodiments, and has similar beneficial effects to the method embodiments. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0156] It should be noted that in the embodiments of the present application, if the above prefetching method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0157] The embodiments of the present application provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0158] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.
[0159] The embodiments of the present application provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.
[0160] An embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. This computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0161] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0162] Figure 7 The following is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. As Figure 7 shown, the hardware entity of the computer device 700 includes: a processor 701 and a memory 702. Among them, the memory 702 stores a computer program that can run on the processor 701. When the processor 701 executes the program, the steps in the method of any of the above embodiments are implemented.
[0163] The memory 702 stores a computer program that can run on the processor. The memory 702 is configured to store instructions and applications executable by the processor 701, and can also cache data to be processed or already processed by the processor 701 and each module in the computer device 700 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0164] When the processor 701 executes the program, the steps of the prefetch method of any of the above are implemented. The processor 701 generally controls the overall operation of the computer device 700.
[0165] An embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the prefetch method of any of the above embodiments.
[0166] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0167] The above-mentioned processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices for implementing the functions of the above-mentioned processor are also possible, and the embodiments of the present application do not make specific limitations.
[0168] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it can also be various terminals including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.
[0169] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above steps / processes do not mean the order of execution, and the order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0170] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0171] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0172] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0173] In addition, in each embodiment of the present application, each functional unit can be entirely integrated into one processing unit, or each unit can be separately regarded as one unit, or two or more units can be integrated into one unit; the above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical discs and other various media that can store program codes.
[0174] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the related art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical discs and other various media that can store program codes.
[0175] The above is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A pre-fetching method, characterized in that: The method comprises: In the process of executing ray tracing of the current ray based on the hierarchical bounding volume structure BVH tree, when the node information of the target node does not exist in the cache module, a prefetch parameter is obtained, and a cache prefetch size of second data to be prefetched and / or a cache position of the second data in the cache module are determined based on the prefetch parameter; the node information is used to determine whether the bounding box corresponding to the target node intersects with the current ray; Reading first data including the node information from the main memory to the cache module, and if the second data exists, pre-fetching second data from the main memory to the cache module based on the cache pre-fetch size of the second data and / or the cache position of the second data in the cache module; The pre-fetch parameters include at least one of the following: level information of the target node in the BVH tree, the ray type of the current ray, and a cache code set by a user.
2. The method according to claim 1, characterized in that The level information of the target node in the BVH tree is used to characterize the spatial locality of the node information of the target node; different ray types correspond to different spatial localities of the node information; Among them, the cache prefetch size of the second data is positively correlated with the spatial locality; the storage distance is negatively correlated with the spatial locality, and the storage distance is the distance between the cache location and the processor that initiates ray tracing.
3. The method according to claim 2, characterized in that The cache module includes at least one level of cache units, and the distances between cache units of different levels and the processor are different; the hierarchical information and / or the light type are used to determine a target cache unit for storing the second data in the at least one level of cache units, and the target cache unit is the cache location.
4. The method according to claim 2, characterized in that: The level information includes at least one of the following: first level information and second level information; the first level information is used to represent the number of levels between the target node and the root node of the BVH tree, and the second level information is used to represent the number of levels between the target node and the leaf node of the BVH tree; The number of levels included in the first-level information is positively correlated with the spatial locality, and the number of levels included in the second-level information is negatively correlated with the spatial locality.
5. The method according to claim 2, characterized in that: The ray type of the current ray includes a principal ray type, a shadow ray type, a reflection ray type and a refraction ray type; Wherein, when other pre-fetch parameters are the same, the cache pre-fetch size when the ray type of the current ray is the main ray type or the shadow line type is greater than the cache pre-fetch size when the ray type of the current ray is the reflection line type or the refraction line type; and / or, When the ray type of the current ray is the main ray type or the shadow line type, the distance between the cache position and the processor is less than the distance between the cache position and the processor when the ray type of the current ray is the reflection line type or the refraction line type.
6. The method according to any one of claims 1 to 5, characterized in that: The obtaining of the prefetch parameters, and determining the cache prefetch size of the second data to be prefetched and / or the cache position of the second data in the cache module based on the prefetch parameters, includes: Query the cache encoding set by the user for the ray tracing process; The cache prefetch size and / or cache location is determined based on the cache encoding.
7. The method according to any one of claims 1 to 5, characterized in that: The obtaining of the prefetch parameters, and determining the cache prefetch size of the second data to be prefetched and / or the cache position of the second data in the cache module based on the prefetch parameters, includes: Query the cache encoding set by the user for the ray tracing process; In the case where there is a cache code set by a user for the ray tracing process, determining the cache prefetch size and / or cache position based on the cache code; In the absence of cache encoding set by a user for the ray tracing process, the level information and / or the ray type are obtained, and the cache prefetch size and / or cache location are determined based on the level information and / or the ray type.
8. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: If it is determined that the second data does not exist, no pre-fetching is performed.
9. A pre-fetching device, characterized in that: The device comprises: An acquisition module is used to acquire prefetch parameters when executing ray tracing of the current ray based on the hierarchical bounding volume structure BVH tree, and determine the cache prefetch size of the second data to be prefetched and / or the cache position of the second data in the cache module based on the prefetch parameters when the node information of the target node does not exist in the cache module; the node information is used to determine whether the bounding box corresponding to the target node intersects with the current ray; a read pre-fetch module, configured to read the first data including the node information from the main memory to the cache module, and, when it is determined based on the cache pre-fetch size of the second data that the second data exists, pre-fetch the second data from the main memory to the cache module based on the cache pre-fetch size of the second data and / or the cache position of the second data in the cache module; The pre-fetch parameter is used to determine at least one of the following: the level information of the target node in the BVH tree, the ray type of the current ray, and the cache encoding set by the user.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps in the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.
12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Efficient caching of resource states for shared functions of graphics processing unit three-dimensional pipelines
CN115858411A
Data scheduling method and device based on ray tracing, equipment and storage medium
CN116049032A