Memory management device and method for multiple parallel channels, electronic equipment, storage medium and computer program product
By managing cache requests from multiple parallel channels through a shared target cache, the problem of wasted cache space and low initialization efficiency in the screen mapping data processing stage of the graphics processor is solved, achieving more efficient cache management and improved processing performance.
Patent Information
- Application Number
- CN202511406611.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-29
AI Technical Summary
In existing technologies, the graphics processor wastes cache space and has low initialization efficiency during the screen mapping data processing stage, resulting in limited processing performance.
A multi-parallel channel memory management device is adopted to manage cache requests from multiple parameter stream generators by sharing a target cache, thereby improving cache hit rate and reducing the number of data interactions.
It improves the efficiency of cache management, reduces redundant operations in cache initialization, reduces data interaction latency, and enhances the processing performance of the graphics processor.
Smart Images

Figure CN120876202A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of graphics processor technology, and more particularly to a multi-parallel channel memory management device, method, electronic device, storage medium, and computer program product. Background Technology
[0002] To improve the processing speed of the Graphics Processing Unit (GPU) during the screen mapping data processing stage, existing technologies typically design multi-threaded parallel pipelines (TPs) and a dedicated Tile Pointer Cache (TPC) for each TP to temporarily store data, thereby enabling parallel generation of the control stream. However, since the TPCs for each TP are independent, multiple initializations are required when initializing the cache line for each TPC, resulting in low initialization efficiency. Furthermore, when the effective screen size is small and the number of tiles is limited, the actual cache space required for data generation is also small, while the cache space for each TPC is fixed in size, leading to significant cache space waste. Summary of the Invention
[0003] In view of this, this disclosure proposes a technical solution for a multi-parallel channel memory management device, method, electronic device, storage medium, and computer program product.
[0004] According to one aspect of this disclosure, a multi-parallel channel memory management device is provided, comprising: a plurality of parameter stream generators, and a target cache shared by the plurality of parameter stream generators; any one parameter stream generator is configured to send a data read request to the target cache, wherein the data read request is configured to obtain reference data corresponding to the valid primitives that the parameter stream generator currently needs to process; the target cache is configured to respond to the data read request of any one parameter stream generator, determine the reference data corresponding to the valid primitives that the parameter stream generator currently needs to process, and return it to the parameter stream generator.
[0005] In one possible implementation, the target cache includes at least one cache line for caching reference data for N primitive groups in the screen to be processed; wherein N is a positive integer, and the value of N is determined based on one or more of the size of the screen to be processed, the number of valid primitives, and the positional distribution, and the number of primitives included in each primitive group is determined according to the number of parameter stream generators sharing the target cache.
[0006] In one possible implementation, the at least one cache line comprises N groups of primitives that are adjacent in the screen to be processed.
[0007] In one possible implementation, a data read request from any parameter stream generator includes the coordinates of the valid primitives that the parameter stream generator currently needs to process in the screen to be processed; the target cache is used to: determine the reference data corresponding to each valid primitive that the parameter stream generator currently needs to process based on the coordinates of the valid primitives that each parameter stream generator currently needs to process.
[0008] In one possible implementation, the target cache is used to: when it is determined that the coordinates of at least two valid primitives that the parameter stream generator currently needs to process belong to the same primitive group, perform a cache lookup based on the cache line identifier corresponding to the primitive group, and determine the cache lookup result.
[0009] In one possible implementation, the target cache is configured to: when the cache lookup result is a cache miss, read reference data of the primitive group from an external storage device according to the cache line identifier corresponding to the primitive group, and write it into a cache line.
[0010] In one possible implementation, the target cache is used to: determine, based on the coordinates of the valid primitives that at least two parameter stream generators currently need to process, the reference data corresponding to the valid primitives that each of the at least two parameter stream generators currently needs to process from the reference data of the primitive group, and return it to the corresponding parameter stream generator.
[0011] In one possible implementation, the target cache is used to: when it is determined that the coordinates of the valid primitives currently being processed by at least two parameter flow generators do not belong to the same primitive group, determine the cache line identifier corresponding to the primitive group to which each of the at least two parameter flow generators belongs, based on the coordinates of the valid primitives currently being processed by each of the at least two parameter flow generators.
[0012] In one possible implementation, the target cache is used to: perform a cache lookup based on the cache line identifier of the primitive group to which each of the at least two parameter stream generators currently needs to process the cached primitives, and determine the cache lookup result of each of the at least two parameter stream generators.
[0013] In one possible implementation, the target cache is used to: for any parameter stream generator that has a cache miss, read reference data of the graphite group to which the parameter stream generator currently needs to process belongs from an external storage device according to the cache line identifier of the graphite group to which the parameter stream generator currently needs to process belongs, and write it into a cache line; and determine the reference data corresponding to the valid graphite to be processed by the parameter stream generator in the reference data of the graphite group to which the valid graphite to be processed belongs, based on the graphite coordinates of the valid graphite to be processed by the parameter stream generator, and return it to the parameter stream generator.
[0014] In one possible implementation, the target cache is used to: determine the data reading order of reading reference data of the valid primitive group to be processed by each cache-missing parameter stream generator from the external storage device, based on the cache line identifier size corresponding to the primitive group to which each cache-missing parameter stream generator currently needs to process belongs, in the case of multiple cache-missing parameter stream generators.
[0015] In one possible implementation, any parameter stream generator is further configured to: compress the primitive mapping data of the valid primitives that the parameter stream generator needs to process, based on the reference data corresponding to the valid primitives that the parameter stream generator needs to process.
[0016] According to another aspect of this disclosure, a multi-parallel-channel memory management method is provided. The method is applied to a multi-parallel-channel memory management device, which includes multiple parameter stream generators and a target cache shared by the multiple parameter stream generators. The method includes: any one parameter stream generator sending a data read request to the target cache, wherein the data read request is used to obtain reference data corresponding to the valid primitives that the parameter stream generator currently needs to process; the target cache responding to the data read request from any one parameter stream generator determines the target reference data corresponding to the valid primitives that the parameter stream generator currently needs to process, and returns it to the parameter stream generator.
[0017] According to another aspect of this disclosure, an electronic device is provided, including the aforementioned multi-parallel channel memory management device.
[0018] According to another aspect of this disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described method.
[0019] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.
[0020] According to another aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0021] The multi-parallel-channel memory management device of this disclosure can send a data read request to the target cache to obtain reference data corresponding to the valid primitives that the parameter stream generator currently needs to process. By setting the target cache to be shared by all parameter stream generators, the target cache can respond to the data read request of any parameter stream generator, determine the reference data corresponding to the valid primitives that the parameter stream generator currently needs to process, and return it to the parameter stream generator. This achieves unified cache management, improves the overall hit rate of data read requests from all parameter stream generators, thereby reducing the number of data interactions between the memory management device and the external storage device, reducing latency in primitive mapping data processing, and improving the processing performance of the memory management device and each parameter stream generator.
[0022] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0023] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0024] Figure 1 A schematic diagram of a multi-parallel channel circuit in the prior art is shown.
[0025] Figure 2 This diagram illustrates the principle of a multi-parallel channel circuit in the prior art.
[0026] Figure 3 A schematic diagram of the structure of a multi-parallel channel memory management device according to an embodiment of the present disclosure is shown.
[0027] Figure 4 A schematic diagram of a multi-parallel channel memory management device according to an embodiment of the present disclosure is shown.
[0028] Figure 5 A flowchart illustrating a multi-parallel channel memory management method according to an embodiment of the present disclosure is shown.
[0029] Figure 6A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0030] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0031] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.
[0032] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.
[0033] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0034] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0035] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0036] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0037] In the screen mapping data processing stage of the Graphics Processing Unit (GPU), after the primitive block (PB) screen mapping is completed, all generated primitive mapping data (PB-Tiles) needs to be compressed to generate control flow data (Control Stream) for each tile and output it. During primitive mapping data processing, an on-chip cache is typically managed to temporarily store the intermediate data frequently read and written for generating the Control Stream for each tile. When the cache is full or a read / write request to the cache misses, the data stored in the cache needs to be flushed and loaded. This is done through data read / write interaction with external storage devices via the BIF (Bus Interface Function) to update the cache's stored data. Furthermore, when starting the mapping data processing task, the memory space of the external storage device also needs to be initialized before normal data processing can begin. A larger cache size can reduce the number of data interactions between the cache and off-chip storage, decreasing latency and thus improving GPU processing performance. However, an excessively large cache size sacrifices significant on-chip area. Therefore, cache design must balance area footprint and processing performance, increasing design complexity.
[0038] In the prior art, in order to improve the generation speed of Control Stream, multi-parallel pipelines (Tiling Pipe, TP) are usually designed, and a cache (Tile Pointer Cache, TPC) is designed for each TP for temporary data storage, thereby realizing the parallel generation of Control Stream.
[0039] Figure 1 A schematic diagram of a multi-parallel channel circuit in the prior art is shown. For example... Figure 1 As shown, circuit 100 includes four independent pipelines: pipeline 0, pipeline 1, pipeline 2, and pipeline 3. Each pipeline is connected to an arbiter (ARB), enabling access to external storage devices via the BIF bus. Each pipeline includes a parameter stream generator (PSG) and a buffer.
[0040] Figure 2 A schematic diagram of a multi-parallel channel circuit in the prior art is shown. For example... Figure 2 As shown in (a), the screen to be processed is divided into 8×4=32 primitives, where gray squares are valid primitives and white squares are invalid primitives. Each primitive has corresponding primitive coordinates, which indicate the position of each primitive in the screen to be processed. The multi-parallel channel circuit corresponding to the screen to be processed includes four pipelines, namely pipeline 0, pipeline 1, pipeline 2 and pipeline 3. The four pipelines perform primitive mapping data processing on the primitives in the screen to be processed based on preset processing rules, forming a mapping relationship between pipelines and primitives.
[0041] Specifically, the primitives corresponding to pipeline 0 include primitives (0,0), (2,0), and (0,2), etc.; the primitives corresponding to pipeline 1 include primitives (1,0), (3,0), and (1,2), etc.; the primitives corresponding to pipeline 2 include primitives (0,1), (2,1), and (0,3), etc.; and the primitives corresponding to pipeline 3 include primitives (1,1), (3,1), and (1,3), etc. Furthermore, the reference data for multiple primitives corresponding to any given pipeline is stored in the cache corresponding to that pipeline. For example... Figure 2 As shown in (b), each pipeline corresponds to a cache. Specifically, pipeline 0 corresponds to cache 0, pipeline 1 corresponds to cache 1, pipeline 2 corresponds to cache 2, and pipeline 3 corresponds to cache 3. The storage space of each cache is fixed and can store reference data corresponding to 8 primitives.
[0042] In any screen mapping process, each TP only needs to process primitive mapping data for its corresponding valid tile. Since each TP's corresponding TPC is independent, each TP's corresponding TPC needs to be allocated an independent cache space (memory cacheline). When the number of valid tiles on the screen is small and they are concentrated, this fixed allocation of memory cacheline will result in a large waste of storage resources.
[0043] Based on the above Figure 2 For example, Figure 2 As shown in (a), the screen to be processed has only four valid primitives: primitive (0,0), primitive (1,0), primitive (2,0), and primitive (3,0). In this case, since the reference data corresponding to the invalid primitives is not needed, the storage resources used to store the reference data corresponding to the invalid primitives are actually wasted.
[0044] On the other hand, when multiple TPs' corresponding PSGs miss data read / write requests to their TPCs, each of these TPs' corresponding TPCs needs to send data interaction requests to the external storage device to update the cached data for each TPC, resulting in significant processing latency. Because the data read / write requests between PSGs and TPCs are very frequent, when the screen size to be processed is large, each TPC needs to frequently interact with the external storage device, which limits the processing performance of each PSG and the entire GPU.
[0045] Based on the above Figure 1 For example, Figure 1 As shown, when each pipeline's corresponding cache needs to send data interaction requests to the external storage device, it needs to perform four data interactions through the BIF bus under the control of the arbitrator, resulting in significant processing latency and affecting the GPU's processing performance.
[0046] In view of this, the present disclosure provides a multi-parallel channel memory management device that can achieve unified cache management, improve the overall hit rate of data read requests from all parameter stream generators, thereby reducing the number of data interactions between the memory management device and the external storage device, reducing latency in primitive mapping data processing, and improving the processing performance of the memory management device and each parameter stream generator. The multi-parallel channel memory management device provided in this disclosure will be described in detail below.
[0047] Figure 3 A schematic diagram of a multi-parallel-channel memory management device according to an embodiment of the present disclosure is shown. Figure 3 As shown, the device 300 includes: multiple parameter stream generators 3010 to 301n, and a target buffer 302 shared by the multiple parameter stream generators. The different parameter stream generators belong to different pipelines.
[0048] Any parameter stream generator 301i (i can be any value from 0 to n) is used to send a data read request to the target cache 302. The data read request is used to obtain the reference data corresponding to the valid primitives that the parameter stream generator 301i needs to process.
[0049] The target cache 302 is used to respond to any data read request from the parameter stream generator 301i, determine the reference data corresponding to the valid primitives that the parameter stream generator 301i needs to process, and return it to the parameter stream generator 301i.
[0050] The screen to be processed can represent any type of display screen, and its specific form can be flexibly set according to actual usage requirements, such as liquid crystal display (LCD), light-emitting diode display (LED), organic light-emitting diode display (OLED), etc. This disclosure does not make any specific limitation in this regard.
[0051] The valid primitives corresponding to the screen to be processed can be defined as the primitives that require primitive mapping data compression during the primitive mapping data processing of the screen to be processed; the number of valid primitives must be less than or equal to the total number of primitives included in the screen to be processed. Correspondingly, primitives that do not require primitive mapping data compression during the primitive mapping data processing of the screen to be processed can be called invalid primitives. The primitive mapping data processing process can be defined as the process by which the GPU, after completing the PB screen mapping processing of the screen to be processed, compresses the generated primitive mapping data (PB-Tiles) based on the primitive partitioning in the screen to be processed. The specific method for determining the primitives of the screen to be processed can be found in related technical implementations, and this disclosure does not specifically limit it.
[0052] The specific form of any parameter stream generator can be found in the implementation methods in related technologies, and this disclosure does not impose any specific limitations on it.
[0053] During the processing of primitive mapping data for the screen to be processed, any parameter stream generator 301i can receive the mapping data of the currently valid primitives to be processed, and send a data read request to the target cache 302 according to the primitive coordinates corresponding to the currently valid primitives to be processed by the parameter stream generator 301i, so as to obtain the reference data corresponding to the currently valid primitives to be processed by the parameter stream generator 301i from the target cache 302. The specific form of the mapping data of the currently valid primitives to be processed received by any parameter stream generator 301i can be referred to in the implementation methods of related technologies, and this disclosure does not specifically limit it.
[0054] The specific form and content of the data read request sent by any parameter stream generator 301i to the target cache 302 can be flexibly set according to actual usage requirements, and this disclosure does not impose specific limitations on this.
[0055] The reference data corresponding to the valid primitives that any parameter flow generator 301i needs to process can represent the intermediate data that needs to be referenced when generating the control flow data corresponding to the valid primitives that need to be processed during the primitive mapping data processing of the screen to be processed. The specific form and content can be referred to the implementation methods in related technologies, and this disclosure does not make specific limitations in this regard.
[0056] The target cache 302 can receive data read requests from any parameter stream generator 301i, enabling unified cache management of all parameter stream generators. This improves the overall hit rate of data read requests from all parameter stream generators, thereby reducing the number of interactions between the target cache and the external storage device corresponding to the device 300, improving the overall processing performance of the device 300, and the data processing efficiency of each parameter stream generator. The specific form of the target cache 302 can be flexibly configured according to actual usage requirements; this disclosure does not impose specific limitations on it.
[0057] Specifically, in response to a data read request from any parameter stream generator 301i, the target cache 302 can determine the reference data corresponding to the valid primitives that the parameter stream generator 301i needs to process in its cache line or in the external storage device corresponding to the device 300, and return it to the parameter stream generator 301i.
[0058] The target cache 302 and its functions will be explained in detail later in conjunction with the possible implementation methods disclosed herein, and will not be repeated here.
[0059] The multi-parallel-channel memory management device of this disclosure can send a data read request to the target cache to obtain reference data corresponding to the valid primitives that the parameter stream generator currently needs to process. By setting the target cache to be shared by all parameter stream generators, the target cache can respond to the data read request of any parameter stream generator, determine the reference data corresponding to the valid primitives that the parameter stream generator currently needs to process, and return it to the parameter stream generator. This achieves unified cache management, improves the overall hit rate of data read requests from all parameter stream generators, thereby reducing the number of data interactions between the memory management device and the external storage device, reducing latency in primitive mapping data processing, and improving the processing performance of the memory management device and each parameter stream generator.
[0060] In one possible implementation, the target cache 302 includes at least one cache line, each cache line being used to cache reference data for N primitive groups in the screen to be processed, the number of primitives included in each primitive group being determined based on the number of parameter stream generators sharing the target cache; where N is a positive integer, and the value of N is determined based on one or more of the size of the screen to be processed, the number of valid primitives, and the positional distribution.
[0061] Specifically, before processing primitive mapping data, the target cache 302 needs to be initialized and data is written from the external storage device corresponding to the device 300 through burst mode to store the data in the target cache.
[0062] In existing multi-parallel channel memory management circuits, each TP corresponds to an independent TPC. Therefore, the initialization process for each TPC is also performed independently. When the number of valid primitives on the screen to be processed is small and their distribution is concentrated, the reference data corresponding to these valid primitives may be cached in a few TPC caches, while the caches of other TPCs do not store the reference data corresponding to valid primitives. However, based on the existing memory management methods, it is still necessary to initialize each TPC separately. Initializing the caches that do not store the reference data corresponding to valid primitives is a redundant operation, which affects the initialization efficiency.
[0063] Based on the above Figure 2 For example, Figure 2 As shown in (a), the only valid primitives corresponding to the screen to be processed are primitives (0,0), (1,0), (2,0), and (3,0). Figure 2 As shown in (b), the reference data corresponding to primitives (0,0) and (2,0) is cached in cache 1 corresponding to pipeline 1, and the reference data corresponding to primitives (1,0) and (3,0) is cached in cache 2 corresponding to pipeline 2. Although cache 3 corresponding to pipeline 3 and cache 4 corresponding to pipeline 4 do not cache any valid reference data corresponding to primitives, these two caches still need to be initialized.
[0064] In this embodiment, the target cache 302 can flexibly configure the number of cache lines and the number of reference data for each cache line of primitive groups based on its own storage performance (e.g., available storage size) and the number and location distribution of valid primitives included in the screen to be processed. Here, the primitive group is determined based on the number of parameter stream generators (the value of n) in the device 300 and preset processing rules. Its specific form can be flexibly set according to actual usage requirements, and this disclosure does not impose specific limitations on it. The number of primitives included in each primitive group is determined based on the number of parameter stream generators sharing the target cache.
[0065] When the number of valid graphic elements in the screen to be processed is small and their distribution is concentrated, the target cache 302 can determine only a small number of cache lines and centrally store the reference data corresponding to these valid graphic elements. Thus, only the reference data corresponding to the valid graphic elements can be initialized, which can effectively reduce the amount of data processed during initialization. Furthermore, since only one shared target cache is set, the reference data corresponding to the valid graphic elements can be written to the target cache with only one initialization, which improves the initialization efficiency.
[0066] Figure 4 This diagram illustrates the principle of a multi-parallel-channel memory management device according to an embodiment of the present disclosure. Figure 4 As shown in (a), the screen to be processed is divided into 8×4=32 primitives. Among them, gray squares are valid primitives and white squares are invalid primitives. Each primitive has corresponding primitive coordinates, which are used to indicate the position of each primitive in the screen to be processed. The valid primitives corresponding to the screen to be processed are only four: primitive (0,0), primitive (1,0), primitive (2,0), and primitive (3,0).
[0067] The screen to be processed has four pipelines: pipeline 0, pipeline 1, pipeline 2, and pipeline 3. Each pipeline includes a corresponding parameter stream generator 301i. The four pipelines process the primitives in the screen to be processed according to preset processing rules, forming a mapping relationship between pipelines and primitives. The primitives corresponding to pipeline 0 include primitives (0,0), primitives (2,0), primitives (4,0), etc. The primitives corresponding to pipeline 1 include primitives (1,0), primitives (3,0), primitives (1,2), etc. The primitives corresponding to pipeline 2 include primitives (0,1), primitives (2,1), primitives (0,3), etc. The primitives corresponding to pipeline 3 include primitives (1,1), primitives (3,1), primitives (1,3), etc.
[0068] Based on the mapping relationship between pipelines and primitives, the primitives in the screen to be processed can be divided into multiple primitive groups. Specifically, primitive (0,0), primitive (1,0), primitive (0,1) and primitive (1,1) form one primitive group, primitive (2,0), primitive (3,0), primitive (2,1) and primitive (3,1) form another primitive group, and so on.
[0069] Based on the size of the reference data corresponding to each graphic element, the total amount of reference data for the graphic element group including valid graphic elements can be determined. Combining this with the amount of data that a single cache line in the target cache 302 can cache, the number of cache lines required for the target cache 302 can be determined. If the total amount of reference data for the graphic element group including valid graphic elements is less than or equal to the amount of data that a single cache line can cache, the target cache 302 can use only one cache line. If the total amount of reference data for the graphic element group including valid graphic elements is greater than the amount of data that a single cache line can cache, the target cache 302 needs to use multiple cache lines.
[0070] In one possible implementation, at least one cache line comprises N groups of primitives that are adjacent in the screen to be processed.
[0071] Based on the above Figure 4 For example, Figure 4 As shown in (a), since the valid primitives (0,0) and (1,0) in the screen to be processed belong to primitive group 1, and the valid primitives (2,0) and (3,0) belong to primitive group 2, primitive group 1 and primitive group 2 are adjacent in the screen to be processed; and the total amount of reference data corresponding to primitive group 1 and primitive group 2 is less than or equal to the amount of data that can be cached by a cache line, the target cache 302 can set only one cache line to cache the reference data corresponding to primitive group 1 and primitive group 2.
[0072] like Figure 4 As shown in (b), the target cache 302 has one and only one cache line. The target cache 302 only needs to be initialized once to write the reference data corresponding to primitive 1 and primitive 2 from the external storage device corresponding to device 300 into the cache line of the target cache 302, thereby effectively reducing the number of initialization processes and improving initialization efficiency.
[0073] In one possible implementation, any data read request of a parameter stream generator 301i includes the coordinates of the valid primitives that the parameter stream generator 301i needs to process in the screen to be processed; the target cache 302 is used to: determine the reference data corresponding to the valid primitives that each parameter stream generator needs to process based on the coordinates of the valid primitives that each parameter stream generator needs to process.
[0074] Specifically, any data read request from a parameter stream generator 301i may include the coordinates of the valid primitives that the parameter stream generator 301i needs to process in the screen to be processed, so as to ensure that the reference data corresponding to the valid primitives that the parameter stream generator 301i needs to process can be accurately determined in the target cache 302.
[0075] Due to factors such as manufacturing precision, different parameter flow generators may have performance differences, resulting in differences in the speed at which different parameter flow generators process primitive mapping data. Consequently, the valid primitives that different parameter flow generators need to process may not belong to the same primitive group.
[0076] Therefore, to improve the efficiency of subsequent cache lookups, when the target cache 302 receives data read requests from at least two parameter stream generators simultaneously, it can determine whether the valid primitives currently being processed by these parameter stream generators belong to the same primitive group based on the primitive coordinates of these valid primitives in the screen to be processed. Since any cache line in the target cache 302 can cache reference data for N primitive groups, when the valid primitives currently being processed by these parameter stream generators belong to the same primitive group, the target cache 302 can perform cache lookups at the primitive group granularity, improving cache lookup efficiency, and determine the reference data corresponding to the valid primitives currently being processed by each parameter stream generator based on the primitive coordinates of each valid primitive.
[0077] If it is determined that the valid primitives to be processed by the at least two parameter stream generators belong to the same primitive group, the target cache 302 can directly use the group identifier corresponding to the primitive group as a basis to perform a cache lookup within the target cache 302 at the granularity of the primitive group, and determine the cache lookup result for the primitive group. The specific form and content of the group identifier can be flexibly set according to actual usage requirements, and this disclosure does not impose specific limitations on it.
[0078] Based on the above Figure 4 For example, Figure 4 As shown in (a), the primitives (0,0), (1,0), (0,1), and (1,1) in the screen to be processed form a primitive group. The cache line identifier corresponding to the primitive group can be set to cache line 1 according to the position of the primitive group in the screen to be processed. The primitives (2,0), (3,0), (2,1), and (3,1) form a primitive group. The cache line identifier corresponding to the primitive group can also be set to cache line 1 according to the position of the primitive group in the screen to be processed, so as to indicate that the two primitive groups can be stored in the same cache line.
[0079] The specific method for cache lookup can be flexibly configured according to actual usage requirements, and this disclosure does not impose specific limitations on it.
[0080] In one example, after the target cache 302 completes the initialization process or interacts with the external storage device corresponding to the device 300, and after storing the reference data of N primitive groups in each cache line, the mapping relationship between the cache address of the reference data of each primitive group in the target cache 302 and the cache line identifier corresponding to the primitive group can be determined. This allows the cache address of the reference data of the primitive group to be found in the target cache 302 based on the cache line identifier corresponding to the primitive group, thereby obtaining the reference data of the primitive group.
[0081] When the target cache 302 receives data read requests from at least two parameter stream generators at the same time, and determines that the valid primitives that the at least two parameter stream generators need to process belong to the same primitive group, it can quickly determine whether the target cache 302 has cached reference data for the primitive group based on the cache line identifier corresponding to the primitive group, thereby realizing cache lookup for the primitive group.
[0082] The cache lookup result of this tuple can be used to indicate whether the reference data of this tuple is cached in the target cache 302. Its specific form and content can be flexibly set according to actual usage requirements.
[0083] If the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the valid primitives that the at least two parameter stream generators need to process belong to the same primitive group, and the cache lookup determines that the target cache 302 has cached reference data for the primitive group, the target cache 302 can determine that the cache lookup result for the primitive group is a cache hit, that is, the target cache 302 has cached reference data for the primitive group, and thus the reference data for the primitive group can be directly obtained from the target cache 302.
[0084] If the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the valid primitives that the at least two parameter stream generators need to process belong to the same primitive group, and the cache lookup determines that the target cache 302 does not cache the reference data of the primitive group, the target cache 302 can determine that the cache lookup result for the primitive group is a cache miss. At this time, the target cache 302 needs to interact with the external storage device corresponding to the device 300 to obtain the reference data of the primitive group.
[0085] When the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the effective primitives currently being processed by the at least two parameter stream generators belong to the same primitive group, and the cache lookup result for the primitive group is a cache hit, the target cache 302 can directly determine the reference data of the primitive group. Based on the primitive coordinates of the effective primitives currently being processed by each of the at least two parameter stream generators, the target cache 302 determines the reference data corresponding to the effective primitives currently being processed by each of the at least two parameter stream generators from the reference data of the primitive group, and then returns the reference data corresponding to the effective primitives currently being processed by each of the at least two parameter stream generators to the corresponding parameter stream generator.
[0086] The specific method by which the target cache 302 returns the reference data corresponding to the valid primitives that the parameter stream generator 301i needs to process can be referred to the implementation methods in related technologies, and this disclosure does not make specific limitations on it.
[0087] Through the above process, the target cache 302 can perform cache lookup at the granularity of primitive groups. Compared with the prior art, which performs cache lookup separately on the reference data corresponding to each valid primitive that the parameter stream generator needs to process, the cache lookup efficiency can be improved, thereby improving the overall processing performance of the device 300.
[0088] In one possible implementation, the target cache 302 is used to: read the reference data of the primitive group from the external storage device and write it into a cache line according to the cache line identifier corresponding to the primitive group when the cache lookup result is a cache miss.
[0089] When the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the valid primitives currently being processed by the at least two parameter stream generators belong to the same primitive group, and the cache lookup result for the primitive group is a cache miss, the target cache 302 can read the reference data of the primitive group from the external storage device corresponding to the device 300 according to the cache line identifier corresponding to the primitive group, and write the reference data of the primitive group into a cache line. Compared with the existing data read / write method of writing the reference data corresponding to the valid primitives currently being processed by each parameter stream generator into the cache line corresponding to each parameter stream generator, this method can effectively reduce the number of data interactions between the target cache 302 and the external storage device corresponding to the device 300, reduce the latency caused by data interactions, and thus improve the overall processing performance of the device 300.
[0090] Based on the above Figure 2 For example, Figure 2As shown in (a), the only valid primitives corresponding to the screen to be processed are primitives (0,0), (1,0), (2,0), and (3,0). Figure 2 As shown in (b), each pipeline corresponds to a cache. Specifically, pipeline 0 corresponds to cache 0, pipeline 1 corresponds to cache 1, pipeline 2 corresponds to cache 2, and pipeline 3 corresponds to cache 3. The storage space of each cache is fixed and can store reference data corresponding to 8 primitives.
[0091] If, during a primitive mapping data processing step, the PSG corresponding to pipeline 0 needs to read the reference data primitive corresponding to primitive (2,0), and the PSG corresponding to pipeline 1 needs to read the reference data corresponding to primitive (3,0), but cache 0 does not cache the reference data corresponding to primitive (2,0), and cache 1 does not cache the reference data corresponding to primitive (3,0), then the TPCs corresponding to cache 0 and cache 1 need to interact with the external storage device separately. In other words, two data interactions are required to meet the usage requirements of the PSGs corresponding to pipeline 0 and pipeline 1, resulting in significant latency and limiting the processing performance of the PSGs corresponding to pipeline 0 and pipeline 1.
[0092] Based on the above Figure 4 For example, Figure 4 As shown in (a), the only valid primitives corresponding to the screen to be processed are primitives (0,0), (1,0), (2,0), and (3,0). Figure 4 As shown in (b), target cache 302 has one and only one cache line.
[0093] If, during a primitive mapping data processing step, the parameter stream generator 3010 corresponding to pipeline 0 needs to read the reference data primitive corresponding to primitive (2,0), and the parameter stream generator 3011 corresponding to pipeline 1 needs to read the reference data corresponding to primitive (3,0), but at this time the cache line of the target cache 302 does not cache the reference data corresponding to primitive (2,0) and primitive (3,0), since primitive (2,0) and primitive (3,0) both belong to primitive group 2, the target cache 302 only needs to perform one data interaction with the external storage device corresponding to device 300 to write the reference data corresponding to primitive group 2. This can simultaneously satisfy the usage requirements of the parameter stream generator 3010 corresponding to pipeline 0 and the parameter stream generator 3011 corresponding to pipeline 1, effectively reducing the latency of data interaction and improving the processing performance of the parameter stream generator 3010 and the parameter stream generator 3011.
[0094] In one possible implementation, the target cache 302 is used to: when it is determined that the valid primitives currently to be processed by at least two parameter flow generators do not belong to the same primitive group, determine the cache line identifier corresponding to the primitive group to which the valid primitives currently to be processed by each of the at least two parameter flow generators belong, based on the primitive coordinates of the valid primitives currently to be processed by each of the at least two parameter flow generators.
[0095] If the target cache 302 receives data read requests from at least two parameter stream generators at the same time, and determines that the effective primitives that the at least two parameter stream generators need to process do not belong to the same primitive group, then the target cache 302 needs to perform cache lookup on the reference data corresponding to the effective primitives that each parameter stream generator needs to process.
[0096] In this case, the target cache 302 can first determine the cache line identifier corresponding to the metamorphic group to which the effective metamorphic group to be processed by each of the at least two parameter stream generators belongs, based on the metamorphic coordinates of the effective metamorphic group to be processed by each of the at least two parameter stream generators. Then, it can perform a cache lookup at the granularity of the metamorphic group to determine the cache lookup result for each of the at least two parameter stream generators, thereby improving the efficiency of the cache lookup. It can also determine whether the target cache 302 caches the reference data corresponding to the effective metamorphic group to be processed by any one of the at least two parameter stream generators.
[0097] In one possible manner, the target cache 302 is used to: perform a cache lookup based on the cache line identifier of the primitive group to which each of the at least two parameter stream generators currently needs to process the valid primitives, and determine the cache lookup result of each of the at least two parameter stream generators respectively.
[0098] If the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the effective primitives currently being processed by the at least two parameter stream generators do not belong to the same primitive group, and for any one of the at least two parameter stream generators, a cache lookup determines that the target cache 302 has cached reference data for the primitive group to which the effective primitive currently being processed by the parameter stream generator belongs, then the target cache 302 can determine that the cache lookup result for the parameter stream generator is a cache hit, and directly determine the reference data corresponding to the effective primitive currently being processed by the parameter stream generator.
[0099] If the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the effective primitives currently being processed by the at least two parameter stream generators do not belong to the same primitive group, and for any one of the at least two parameter stream generators, a cache lookup determines that the target cache 302 does not cache the reference data of the primitive group to which the effective primitives currently being processed by the parameter stream generator belong, the target cache 302 can determine that the cache lookup result for the parameter stream generator is a cache miss. At this time, the target cache 302 needs to interact with the external storage device corresponding to the device 300 to obtain the reference data of the primitive group to which the effective primitives currently being processed by the parameter stream generator belong.
[0100] When the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the valid primitives currently being processed by the at least two parameter stream generators do not belong to the same primitive group, and the cache lookup result of any one of the at least two parameter stream generators is a cache hit, for the parameter stream generator with the cache hit, the target cache 302 can determine the reference data of the primitive group to which the valid primitives currently being processed by the cache hit parameter stream generator belong in the cache line of the target cache 302 according to the cache line identifier of the primitive group to which the valid primitives currently being processed by the cache hit parameter stream generator belong; and based on the primitive coordinates of the valid primitives currently being processed by the cache hit parameter stream generator, directly determine the reference data corresponding to the valid primitives currently being processed by the cache hit parameter stream generator, and return the reference data corresponding to the valid primitives currently being processed by the cache hit parameter stream generator to the parameter stream generator with the cache hit.
[0101] Correspondingly, if the target cache 302 receives only a data read request corresponding to a parameter stream generator 301i, the above process can be used to determine the reference data corresponding to the valid primitives that the parameter stream generator 301i needs to process. This will not be elaborated here.
[0102] In one possible implementation, the target cache 302 is used to: for any parameter stream generator that has a cache miss, read the reference data of the graphite group to which the parameter stream generator currently needs to process belongs from the external storage device according to the cache line identifier of the graphite group to which the parameter stream generator currently needs to process belongs, and write it into a cache line; and determine the reference data corresponding to the valid graphite to be processed by the parameter stream generator in the reference data of the graphite group to which the parameter stream generator currently needs to process belongs, based on the graphite coordinates of the valid graphite to be processed by the parameter stream generator, and return it to the parameter stream generator.
[0103] If the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the valid primitives currently being processed by the at least two parameter stream generators do not belong to the same primitive group, and the cache lookup result of any one of the at least two parameter stream generators is a cache miss, then for the parameter stream generator with the cache miss, the target cache 302 can read the reference data of the primitive group to which the valid primitives currently being processed by the parameter stream generator belongs from the external storage device corresponding to the device 300, based on the cache line identifier of the primitive group to which the valid primitives currently being processed by the parameter stream generator belongs, and write the reference data of the primitive group to which the valid primitives currently being processed by the parameter stream generator belongs into a cache line for subsequent data retrieval.
[0104] The target cache 302 can determine the reference data corresponding to the valid primitive that the parameter stream generator that the cache misses can directly determine from the reference data of the primitive group to which the valid primitive that the cache misses can currently process, based on the primitive coordinates of the valid primitive that the parameter stream generator that the cache misses can currently process, and return it to the parameter stream generator that the cache misses.
[0105] In one possible implementation, the target cache 302 is used to: determine the data reading order of the reference data of the valid primitive group to which each cache-missing parameter stream generator belongs, based on the cache line identifier size corresponding to the primitive group to which each cache-missing parameter stream generator belongs, in the case of multiple cache-missing parameter stream generators, from the external storage device.
[0106] When the target cache 302 simultaneously receives data read requests from at least two parameter stream generators, and determines that the effective primitives currently being processed by the at least two parameter stream generators do not belong to the same primitive group, and there are multiple parameter stream generators with cache misses among the at least two parameter stream generators, the target cache 302 can determine the data read order of reading the reference data of the primitive group to which the effective primitives currently being processed by each cache miss parameter stream generator belongs from the external storage device based on the cache line identifier size corresponding to the primitive group to which the effective primitives currently being processed by each cache miss parameter stream generator belongs. Based on the data read order, the target cache 302 sequentially writes the reference data of the primitive group to which the effective primitives currently being processed by each cache miss parameter stream generator 301i belongs from the external storage device corresponding to the device 300.
[0107] The specific form of the data reading order can be flexibly set according to actual usage needs. For example, it can be set to the order of group identifiers from smallest to largest, or it can be set to the order of group identifiers from largest to smallest. This disclosure does not make specific limitations on this.
[0108] In one possible implementation, any parameter stream generator 301i is further configured to: compress the valid primitives that the parameter stream generator 301i needs to process based on the reference data corresponding to the valid primitives that the parameter stream generator 301i needs to process.
[0109] Specifically, when any parameter flow generator 301i receives reference data corresponding to the valid primitives that need to be processed, it can compress the mapping data corresponding to the valid primitives that need to be processed by the parameter flow generator 301i based on the reference data, and determine the control flow data corresponding to the current mapping data processing of the valid primitives that need to be processed by the parameter flow generator 301i.
[0110] For specific methods of primitive mapping data compression, please refer to the implementation methods in related technologies. This disclosure does not make specific limitations on them.
[0111] After any parameter flow generator 301i generates the control flow data corresponding to the current primitive mapping data processing process for the valid primitives it needs to process, it can also cache the control flow data generated in this primitive mapping data processing, as well as the new intermediate data output, in the corresponding position within the target cache 302 for use in the next primitive mapping data processing process. The cache position of the new data in the target cache 302 can be referenced from implementation methods in related technologies and depends on the actual situation of the next primitive mapping processing process; this disclosure does not specifically limit it in this regard.
[0112] The multi-parallel channel memory management device of this disclosure can send a data read request to the target cache to obtain reference data corresponding to the valid primitives that the parameter stream generator currently needs to process. By setting the target cache to be shared by all parameter stream generators, the target cache can respond to the data read request of any parameter stream generator, determine the reference data corresponding to the valid primitives that the parameter stream generator currently needs to process, and return it to the parameter stream generator. This achieves unified cache management, improves the overall hit rate of data read requests from all parameter stream generators, thereby reducing the number of data interactions between the memory management device and the external storage device, reducing latency in primitive mapping data processing, and improving the processing performance of the memory management device and each parameter stream generator. Furthermore, the target cache can also flexibly allocate the size of the storage space according to the number and position of valid primitives in the screen to be processed, effectively reducing the waste of storage resources and improving the efficiency of initialization processing.
[0113] It should be noted that, although... Figure 1 The above example illustrates a multi-parallel-channel memory management device, but those skilled in the art will understand that this disclosure is not limited thereto. In fact, users can flexibly configure the component composition and structure of the multi-parallel-channel memory management device according to their personal preferences and / or actual application scenarios, as long as it achieves unified cache management based on the above principles, improving the processing performance of the memory management device and each parameter stream generator.
[0114] Furthermore, according to another aspect of this disclosure, a multi-parallel channel memory management method is provided, which is applied to a multi-parallel channel memory management device, the memory management device including multiple parameter stream generators and a target cache shared by the multiple parameter stream generators.
[0115] Figure 5 A flowchart illustrating a multi-parallel channel memory management method according to an embodiment of the present disclosure is shown. Figure 5 As shown, the method includes:
[0116] In step S501, any parameter stream generator sends a data read request to the target cache, wherein the data read request is used to obtain the reference data corresponding to the valid primitives that the parameter stream generator needs to process.
[0117] In step S502, the target cache responds to the data read request of any parameter stream generator, determines the target reference data corresponding to the valid primitive that the parameter stream generator needs to process, and returns it to the parameter stream generator.
[0118] In one possible implementation, the target cache includes at least one cache line, each cache line being used to cache reference data for N primitive groups in the screen to be processed, the number of primitives included in each primitive group being determined based on the number of parameter stream generators sharing the target cache; where N is a positive integer, and the value of N is determined based on one or more of the size of the screen to be processed, the number of valid primitives, and the positional distribution.
[0119] In one possible implementation, at least one cache line comprises N groups of primitives that are adjacent in the screen to be processed.
[0120] In one possible implementation, any data read request from a parameter stream generator includes the coordinates of the valid primitives that the parameter stream generator needs to process in the screen to be processed; the target cache determines the reference data corresponding to the valid primitives that each parameter stream generator needs to process based on the coordinates of the valid primitives that each parameter stream generator needs to process.
[0121] In one possible implementation, the method further includes: when it is determined that the coordinates of the valid primitives that the at least two parameter stream generators need to process belong to the same primitive group, performing a cache lookup based on the cache line identifier corresponding to the primitive group, and determining the cache lookup result.
[0122] In one possible implementation, the method further includes: if the cache lookup result is a cache miss, reading the reference data of the primitive group from the external storage device according to the cache line identifier corresponding to the primitive group, and writing it into a cache line.
[0123] In one possible implementation, the target cache determines the reference data corresponding to the valid primitives that each parameter stream generator needs to process based on the primitive coordinates of the valid primitives that each parameter stream generator needs to process. This includes: determining the reference data corresponding to the valid primitives that each of the at least two parameter stream generators needs to process based on the primitive coordinates of the valid primitives that at least two parameter stream generators needs to process from the reference data of the primitive group, and returning it to the corresponding parameter stream generator.
[0124] In one possible implementation, based on the cache lookup result, the reference data corresponding to the valid primitive that each of the at least two parameter stream generators needs to process is determined and returned to the corresponding parameter stream generator. This includes: for any one of the at least two parameter stream generators that has a cache hit, based on the primitive coordinates of the valid primitive that the cache hit parameter stream generator needs to process, determining the reference data corresponding to the valid primitive that the cache hit parameter stream generator needs to process from the reference data of the primitive group to which the valid primitive that the cache hit parameter stream generator belongs, and returning it to the cache hit parameter stream generator.
[0125] In one possible implementation, the target cache determines the reference data corresponding to the valid primitives that each parameter stream generator needs to process, based on the primitive coordinates of the valid primitives that each parameter stream generator needs to process: For any parameter stream generator that has a cache miss, the reference data of the primitive group to which the valid primitives to be processed belong, based on the cache line identifier of the primitive group to which the valid primitives to be processed belong, is read from the external storage device and written into a cache line; based on the primitive coordinates of the valid primitives to be processed, the reference data corresponding to the valid primitives to be processed by the parameter stream generator is determined from the reference data of the primitive group to which the valid primitives to be processed belong, and returned to the parameter stream generator.
[0126] In one possible implementation, for any parameter stream generator that has a cache miss, the reference data of the valid primitive group to which the parameter stream generator currently needs to process belongs is read from the external storage device according to the cache line identifier of the primitive group to which the parameter stream generator currently needs to process belongs, and written into a cache line. In the case of multiple parameter stream generators with cache misses, the data reading order of reading the reference data of the valid primitive group to which each parameter stream generator currently needs to process belongs is determined according to the size of the cache line identifier of the primitive group to which each parameter stream generator currently needs to process belongs.
[0127] In one possible implementation, the method further includes: compressing the primitive mapping data of the valid primitives that the parameter stream generator needs to process, based on the reference data corresponding to the valid primitives that the parameter stream generator needs to process.
[0128] This disclosure also provides an electronic device including the above-described multi-parallel channel memory management device.
[0129] This disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0130] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0131] This disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above method.
[0132] Figure 6 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. For example, device 1900 may be provided as a server or terminal device. (Refer to...) Figure 6 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0133] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0134] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.
[0135] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0136] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.
[0137] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.
[0138] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0139] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0140] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0142] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A memory management device with multiple parallel channels, characterized in that, The apparatus includes: a plurality of parameter stream generators, and a target cache shared by the plurality of parameter stream generators; Any parameter stream generator is used to send a data read request to the target cache, wherein the data read request is used to obtain the reference data corresponding to the valid primitives that the parameter stream generator needs to process at present; The target cache is used to respond to a data read request from any parameter stream generator, determine the reference data corresponding to the valid primitives that the parameter stream generator needs to process, and return it to the parameter stream generator.
2. The apparatus according to claim 1, characterized in that, The target cache includes at least one cache line, which is used to cache reference data for N primitive groups in the screen to be processed. The number of primitives included in each primitive group is determined according to the number of parameter stream generators sharing the target cache. Wherein, N is a positive integer, and the value of N is determined based on one or more of the size of the screen to be processed, the number of valid primitives, and the positional distribution.
3. The apparatus according to claim 2, characterized in that, The at least one cache line includes N groups of primitives that are adjacent in the screen to be processed.
4. The apparatus according to claim 2, characterized in that, A data read request from any parameter stream generator, including the coordinates of the valid primitives that the parameter stream generator currently needs to process in the screen to be processed; the target cache is used for: Based on the coordinates of the valid primitives that each parameter stream generator needs to process, the reference data corresponding to the valid primitives that each parameter stream generator needs to process is determined.
5. The apparatus according to claim 4, characterized in that, The target cache is used for: If it is determined that the coordinates of at least two valid primitives that the current parameter stream generator needs to process belong to the same primitive group, a cache lookup is performed based on the cache line identifier corresponding to the primitive group to determine the cache lookup result.
6. The apparatus according to claim 5, characterized in that, The target cache is used for: If the cache lookup result is a cache miss, the reference data of the primitive group is read from the external storage device according to the cache line identifier corresponding to the primitive group, and written into a cache line.
7. The apparatus according to any one of claims 4 to 6, characterized in that, The target cache is used for: Based on the coordinates of the valid primitives that at least two parameter flow generators need to process, the reference data corresponding to the valid primitives that each of the at least two parameter flow generators needs to process is determined from the reference data of the primitive group, and then returned to the corresponding parameter flow generator.
8. The apparatus according to claim 4, characterized in that, The target cache is used for: If it is determined that the coordinates of the valid primitives to be processed by at least two parameter stream generators do not belong to the same primitive group, the cache line identifier corresponding to the primitive group to which the valid primitives to be processed by each of the at least two parameter stream generators belong is determined according to the coordinates of the valid primitives to be processed by each of the at least two parameter stream generators.
9. The apparatus according to claim 8, characterized in that, The target cache is used for: Based on the cache line identifier of the valid primitive to be processed by each of the at least two parameter stream generators, a cache lookup is performed to determine the cache lookup result of each of the at least two parameter stream generators.
10. The apparatus according to claim 9, characterized in that, The target cache is used for: For any parameter stream generator that has a cache miss, the reference data of the valid primitive to be processed by the parameter stream generator is read from the external storage device according to the cache line identifier of the primitive group to which the parameter stream generator belongs, and written into a cache line. Based on the coordinates of the valid primitives that the parameter flow generator needs to process, the reference data corresponding to the valid primitives that the parameter flow generator needs to process is determined from the reference data of the primitive group to which the valid primitives that the parameter flow generator needs to process belong, and then returned to the parameter flow generator.
11. The apparatus according to claim 10, characterized in that, The target cache is used for: In the case of multiple cache-missing parameter stream generators, the data reading order for reading the reference data of the valid primitive group to which each cache-missing parameter stream generator belongs is determined based on the cache line identifier size corresponding to the primitive group to which the valid primitive to be processed currently needs to be processed by each cache-missing parameter stream generator.
12. The apparatus according to any one of claims 1 to 6, characterized in that, Any parameter stream generator can also be used for: Based on the reference data corresponding to the valid primitives that the parameter stream generator needs to process, primitive mapping data compression is performed on the valid primitives that the parameter stream generator needs to process.
13. A memory management method for multiple parallel channels, characterized in that, The method is applied to a multi-parallel channel memory management device, which includes multiple parameter stream generators and a target cache shared by the multiple parameter stream generators. The method includes: Any parameter stream generator sends a data read request to the target cache, wherein the data read request is used to obtain the reference data corresponding to the valid primitives that the parameter stream generator needs to process at present; The target cache responds to a data read request from any parameter stream generator by determining the target reference data corresponding to the valid primitives that the parameter stream generator currently needs to process, and returns it to the parameter stream generator.
14. An electronic device, characterized in that, The memory management device comprising any one of claims 1 to 12.
15. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 13.
16. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 13.
17. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 13.
Citation Information
Patent Citations
Processing unit and processing method thereof
CN106201980A
Vertex shader with primitive replication
CN110942415A
Rendering task scheduling method and device based on primitives and storage medium
CN112801855A
Graphics processing system and method based on bitmap primitives and GPU (Graphics Processing Unit)
CN115349136A
Picture block distribution control method, chip, device, controller, equipment and medium
CN116740248A