Chip system and method for reading data
By using the virtual address query partition mapping table in the chip system of the graphics processor and the memory manager, the problem of performance degradation of graphics processors when load is high is solved, and more efficient data access and lower power consumption are achieved.
Patent Information
- Application Number
- CN202311713420.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
How to improve the performance of graphics processors, especially when the load is high, the power consumption increases and performance decreases.
Design a chip system, including a graphics processor and a memory manager. When the graphics processor first requests to read data, the partition mapping table is queried based on the virtual address and reads data from memory. This system reduces the overhead of finding segmented area identifiers, increases the number of corresponding resources of segmented area identifiers, and reduces excessive mapping or remapping.
Improves the performance of the graphics processor, reduces power consumption, and enhances system compatibility, allowing it to support larger resolution scenarios.
Smart Images

Figure CN120147494A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of chip technology, and in particular, to a chip system and a method for reading data. Background Art
[0002] A graphics processing unit (GPU) is a processor used for graphics rendering generation. Currently, the graphics processing unit can also be applied to fields other than image rendering, such as accelerating artificial intelligence (AI), scientific computing, big data, bioinformatics, computer data, and financial models.
[0003] On the data center side, the graphics processing unit is deployed in the form of an independent computing card. In a system on chip (SOC) on the mobile terminal side, the cooperation of the graphics processing unit, central processing unit (CPU), image signal processor (ISP), encoder, decoder, and storage subsystem can improve the display quality of images on the mobile terminal side. At the same time, more and more artificial intelligence functions on the mobile terminal side rely on the graphics processing unit. When the load of the graphics processing unit is large, the power consumption of the graphics processing unit increases, and the performance of the graphics processing unit will be reduced. Therefore, how to improve the performance of the graphics processing unit has become an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of the present application provide a chip system and a method for reading data, which improve the performance of the graphics processing unit.
[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions.
[0006] In a first aspect, an embodiment of the present application provides a chip system, which includes a graphics processor and a memory manager. The graphics processor is configured to: in response to a graphics rendering task, output a first read instruction to the memory manager, where the first read instruction is used to indicate reading first data based on a first virtual address. The first virtual address belongs to a first virtual address storage space, and the first virtual address storage space includes a plurality of consecutive first virtual address areas. Each first virtual address area includes a plurality of first cache segment areas obtained by dividing based on a corresponding cache granularity. The first data is the data first read by the graphics processor based on the graphics rendering task, and the data volume of the first data is less than or equal to the cache granularity of the plurality of first virtual address areas. The memory manager is configured to: in response to the first read instruction, read the first data from the memory based on the first virtual address and a partition mapping table, where the partition mapping table includes a plurality of first cache segment areas and a plurality of first segment area identifiers. The plurality of first cache segment areas correspond to the plurality of first segment area identifiers one by one. Output the first data to the graphics processor.
[0007] Therefore, in the chip system provided by the embodiment of the present application, when the graphics processor first requests to read the first data, the first virtual address in the first read instruction can be used to query the partition mapping table, so as to read the first data from the memory. Among them, since the partition mapping table includes the correspondence between a plurality of first cache segment areas and a plurality of first segment area identifiers, and the cache granularity of each first cache segment area in the plurality of first cache segment areas is the same, the first segment area identifier corresponding to the first virtual address can be directly determined. Compared with the method of sequentially searching for the first segment area identifier, the chip system provided by the embodiment of the present application can reduce the overhead, and can increase the number of resources corresponding to the first segment area identifier, reduce excessive mapping or remapping, and improve the performance of the graphics processor.
[0008] In a possible design, the memory manager is specifically configured to: based on the first virtual address and the partition mapping table, determine a first target segment area identifier, where the first target segment area identifier is the first segment area identifier corresponding to the first cache segment area where the first data is located. Read the first data according to the first target segment area identifier. Therefore, since the cache granularity of each first cache segment area in the plurality of first cache segment areas in the partition mapping table is the same, the first target segment area identifier corresponding to the first virtual address can be determined through simple logical operations, reducing the overhead required to determine the first target segment area identifier.
[0009] In a possible design, the memory manager is specifically configured to: determine a first target segmentation area identifier based on a first virtual address, virtual partition start addresses of multiple first virtual address areas, and a cache granularity corresponding to a first virtual address area. Thus, based on the first virtual address and the virtual partition start addresses of the multiple first virtual address areas, the identifier of the first virtual address area can be determined, and further based on the cache granularity corresponding to the first virtual address area, the first target segmentation area identifier can be determined, reducing the overhead required to determine the first target segmentation area identifier.
[0010] In a possible design, the partition mapping table further includes multiple second cache segmentation areas and multiple second segmentation area identifiers. The multiple second cache segmentation areas form a continuous second virtual address storage space. The cache granularity of the second cache segmentation areas is greater than that of the multiple first cache segmentation areas. The multiple second cache segmentation areas are in one-to-one correspondence with the multiple second segmentation area identifiers. The graphics processing unit is further configured to: in response to a graphics rendering task, output a second read instruction to the memory manager, where the second read instruction is used to indicate reading second data based on a second virtual address, the second virtual address belongs to the second virtual address storage space, the data volume of the second data is greater than the cache granularity of the multiple first virtual address areas, and the second data is the data first read by the graphics processing unit based on the graphics rendering task. The memory manager is further configured to: in response to the second read instruction, read the second data from the memory based on the second virtual address, the partition mapping table, and the resource matching table. The resource matching table is used to indicate the mapping relationship between the second virtual address and the second segmentation area identifier. Output the second data to the graphics processing unit. Thus, when the second data required for the graphics processing unit to execute the graphics rendering task is located in the second virtual address storage space, the mapping relationship between the second virtual address and the second segmentation area identifier can be dynamically adjusted, that is, the texture resources corresponding to the second segmentation area identifier can be dynamically loaded or unloaded, which can improve the compatibility of the chip system and enable the chip system to support larger resolution scenarios.
[0011] In a possible design, the memory manager is specifically configured to: determine a second target segmentation area identifier based on the second virtual address and virtual partition start addresses of the multiple second cache segmentation areas. Read the second data based on the second target segmentation area identifier and the resource matching table. Thus, the chip system can sequentially search for the second cache segmentation area based on the second virtual address. When the second virtual address and the virtual partition start address of the current second cache segmentation area are the minimum values, the second target segmentation area identifier can be determined.
[0012] In a possible design, the memory manager further includes a cache for storing third data, where the third data is the data that the graphics processor requests to read again based on a graphics rendering task. The graphics processor is further configured to: in response to a graphics rendering task, output a third read instruction to the memory manager, where the third read instruction is used to indicate reading the third data. The memory manager is further configured to: in response to the third read instruction, read the third data from the cache and output the third data to the graphics processor. Thus, if the graphics processor requests to read the third data again, there is no need to read the data from the memory, and the third data in the cache can be directly output to the graphics processor, thereby improving the data processing speed.
[0013] In a possible design, the first read instruction further includes an identification bit for indicating the address path type corresponding to the first virtual address, and the chip system further includes: a snooping filter. When the identification bit is a first value, the snooping filter is configured to: output the first read instruction to the memory manager.
[0014] In a second aspect, an embodiment of the present application provides a method for reading data. The method is applied to a chip system including a graphics processor and a memory manager, and the method includes: The graphics processor outputs a first read instruction to the memory manager in response to a graphics rendering task, where the first read instruction is used to indicate reading first data based on a first virtual address. The first virtual address belongs to a first virtual address storage space, and the first virtual address storage space includes a plurality of consecutive first virtual address areas, and each first virtual address area includes a plurality of first cache segment areas obtained by dividing based on a corresponding cache granularity. The first data is the data that the graphics processor reads for the first time based on a graphics rendering task, and the data volume of the first data is less than or equal to the cache granularity of the plurality of first virtual address areas. The memory manager reads the first data from the memory based on the first virtual address and a partition mapping table in response to the first read instruction, where the partition mapping table includes a plurality of first cache segment areas and a plurality of first segment area identifiers. The plurality of first cache segment areas correspond to the plurality of first segment area identifiers one by one. Output the first data to the graphics processor.
[0015] In a possible design, the memory manager reads the first data from the memory based on the first virtual address and a partition mapping table in response to the first read instruction, including: The memory manager determines a first target segment area identifier based on the first virtual address and the partition mapping table, where the first target segment area identifier is the first segment area identifier corresponding to the first cache segment area where the first data is located. Read the first data from the memory according to the first target segment area identifier.
[0016] In a possible design, the memory manager determines a first target segment area identifier based on a first virtual address and a partition mapping table, including: The memory manager determines the first target segment area identifier based on the first virtual address, the virtual partition start addresses of multiple first virtual address areas, and the cache granularity corresponding to the first virtual address area.
[0017] In a possible design, the partition mapping table further includes multiple second cache segment areas and multiple second segment area identifiers. The multiple second cache segment areas form a continuous second virtual address storage space. The cache granularity of the second cache segment area is greater than that of the multiple first cache segment areas. The multiple second cache segment areas correspond one-to-one with the multiple second segment area identifiers. The method further includes: In response to a graphics rendering task, the graphics processor outputs a second read instruction to the memory manager, and the second read instruction is used to indicate reading second data based on a second virtual address, where the second virtual address belongs to the second virtual address storage space, the data volume of the second data is greater than the cache granularity of the multiple first virtual address areas, and the second data is the data first read by the graphics processor based on the graphics rendering task. In response to the second read instruction, the memory manager reads the second data from the memory based on the second virtual address, the partition mapping table, and a resource matching table. The resource matching table is used to indicate the mapping relationship between the second virtual address and the second segment area identifier. Output the second data to the graphics processor.
[0018] In a possible design, the memory manager reads the second data from the memory based on the second virtual address, the partition mapping table, and the resource matching table, including: The memory manager determines a second target segment area identifier based on the second virtual address and the virtual partition start addresses of the multiple second cache segment areas. Read the second data from the memory based on the second target segment area identifier and the resource matching table.
[0019] In a possible design, the memory manager further includes a cache, and the cache is used to store third data, where the third data is the data that the graphics processor requests to read again based on the graphics rendering task. The method further includes: In response to a graphics rendering task, the graphics processor outputs a third read instruction to the memory manager, and the third read instruction is used to indicate reading the third data. In response to the third read instruction, the memory manager reads the third data from the cache; Output the third data to the graphics processor.
[0020] In a possible design, the first read instruction further includes a flag bit, and the flag bit is used to indicate the address path type corresponding to the first virtual address. The chip system further includes: A snooping filter. The method further includes: When the flag bit is a first value, the snooping filter outputs the first read instruction to the memory manager.
[0021] In a third aspect, an embodiment of the present application further provides an electronic device, which includes one or more interface circuits and one or more chip systems according to the first aspect. The interface circuits and the chip systems are interconnected through lines.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including computer instructions, which when running on an electronic device, cause the electronic device to execute the method for reading data in any of the above aspects and any possible implementation manners.
[0023] In a fifth aspect, an embodiment of the present application provides a computer program product, which when running on a computer or a processor, causes the computer or the processor to execute the method for reading data in any of the above aspects and any possible implementation manners.
[0024] It can be understood that any of the above provided chip systems, electronic devices, computer-readable storage media, or computer program products can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, which will not be elaborated here.
[0025] These aspects or other aspects of the present application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of a multi-channel video playback scenario provided by an embodiment of the present application;
[0027] Figure 2 It is a processing flowchart of a second graphics processor executing tasks provided by an embodiment of the present application;
[0028] Figure 3 It is a schematic flowchart of a second graphics processor executing tasks in parallel provided by an embodiment of the present application;
[0029] Figure 4 It is a schematic flowchart of a second graphics processor executing tasks serially provided by an embodiment of the present application;
[0030] Figure 5 It is a schematic structural diagram of a chip system provided by an embodiment of the present application;
[0031] Figure 6 It is a mapping diagram of a first virtual address storage space provided by an embodiment of the present application;
[0032] Figure 7 It is another processing flowchart of a third graphics processor executing tasks provided by an embodiment of the present application;
[0033] Figure 8A mapping diagram of a second virtual address storage space provided by an embodiment of the present application;
[0034] Figure 9 A mapping diagram of a virtual address space of a third graphics processing unit provided by an embodiment of the present application;
[0035] Figure 10 A schematic structural diagram of another chip system provided by an embodiment of the present application;
[0036] Figure 11 A structural diagram of a decoder provided by an embodiment of the present application;
[0037] Figure 12 A flowchart of a method for reading data provided by an embodiment of the present application. Detailed implementation manners
[0038] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise stated, the meaning of "a plurality" is two or more than two.
[0039] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise stated, " / " means "or", for example, A / B may mean A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.
[0040] First, some basic concepts related to the embodiments of the present application are explained:
[0041] The slab mechanism is a memory allocation mechanism of the operating system that divides the memory of the operating system kernel into block memories of unequal sizes for management. Specifically, a slab can be represented as a set of consecutive physical address pages for storing objects of a specific type. Among them, when the operating system allocates slabs, each slab is pre-allocated, and each slab can accommodate a certain number of objects of a fixed size.
[0042] In the interaction scenario between a graphics processing unit and an external media processing unit, for example, in the case of a heavy multi-channel load (the number of conference call access channels is greater than 64), the chip system provided by the embodiments of the present application can assist the graphics processing unit in performing efficient and low-latency resource management on multi-channel resources (video memory) to achieve scenarios such as multi-channel video conferencing, calls, and multi-channel video editing. As Figure 1 shown, Figure 1 is a schematic diagram of a multi-channel video playback scenario provided by the embodiments of the present application. Among them, Figure 1 only 16 video channels are shown, namely video channel 1, video channel 2,..., video channel 16. It can be understood that the graphics processing unit can also support the simultaneous playback of videos on more video channels.
[0043] Among them, the chip system provided by the embodiments of the present application greatly simplifies the access cost and processing cost of the graphics processing unit to external resources. The chip system provided by the embodiments of the present application can be applied not only to the scenario where the graphics processing unit is connected to a video decoder, but also to the scenario where the graphics processing unit is connected to other media channels, such as an image processor, a video encoder, etc. In addition, the chip system provided by the embodiments of the present application can also be applied to scenarios such as artificial intelligence, video compression, and general computing. The embodiments of the present application do not limit the application scenarios.
[0044] In some embodiments, the chip system provided by the embodiments of the present application can be a system-on-a-chip (SoC) or a server chip, etc. The device to which the chip system provided by the embodiments of the present application is applied can be an execution device, and the execution device can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) device, a virtual reality (VR) device, and a vehicle-mounted terminal, etc., or a server cluster, etc.
[0045] In some possible embodiments, a first graphics processing unit is proposed. The first graphics processing unit copies external resources into the memory and then processes the external resources in the memory. However, this way of copying external resources will bring additional memory occupancy and data copying overhead.
[0046] In some possible embodiments, in order to reduce additional memory occupancy and data replication overhead, a second graphics processor is proposed. The second graphics processor adopts a zero-copy scheme, that is, the direct memory access buffer (DMA_BUF) mechanism of the kernel. In the DMA_BUF mechanism, a virtual address space is shared between the second graphics processor and external resources. The virtual address space can also be referred to as virtual memory, thereby avoiding the replication of external resources. Among them, each external resource buffer can be directly mapped to the virtual address space of the second graphics processor. The second graphics processor needs to perform a mapping operation. Specifically, an application or an upper-layer application (such as a decoder application) will send a file handle to the driver of the second graphics processor, but the driver still needs to perform the work of page table remapping so that the second graphics processor can access this segment of the virtual address space. The DMA_BUF mechanism usually allocates dynamic virtual address mappings on demand in units of the buffer of each DMA_BUF. Under the constraint of a certain address range, a virtual address space is dynamically obtained and synchronously written into the page table (such as an address mapping table) of the second graphics processor. Compared with the method of replicating external resources, the DMA_BUF mechanism reduces memory occupancy and data replication overhead, optimizes the data pipeline, and greatly improves the efficiency and performance of the second graphics processor in accessing external resources. Due to the dynamic virtual address mapping, the utilization rate of the virtual address space of the second graphics processor is improved, and at the same time, the second graphics processor can support a wider range of complex scenarios.
[0047] However, the number of resources in the virtual address space that the second graphics processor can support simultaneously is limited. When the number of external resources exceeds the limit of the number of resources in the virtual address space, the system needs to release the mapped virtual address space and then remap it. Among them, the rotation and remapping operations are similar to the swap mechanism of the storage system. The remapping operation will bring the following disadvantages: (1) The performance of the second graphics processor decreases. Among them, repeated mapping or unmapping will increase the overhead, and it is necessary to flush the translation lookaside buffer (TLB), which will increase the latency of the second graphics processor accessing memory. (2) It increases the bandwidth pressure on the second graphics processor. Among them, repeated mapping or unmapping requires reading and writing unused memory pages through the memory bus, occupying the bandwidth of the memory bus. (3) It increases the power consumption of the second graphics processor. Among them, additional memory access and data transmission will increase the power consumption. (4) It increases the programming complexity of the second graphics processor. Among them, the second graphics processor needs to consider the mapping limit and manage the mapped memory and unmapped memory in a timely manner. (5) Repeated mapping or unmapping fragments the virtual address space of the second graphics processor.
[0048] The processing flow of the second graphics processor executing the task is as follows: Figure 2 As shown, Figure 2 The processing flow of the application program interface (API) is shown in the figure. The processing flow may include: S201, create an image (create image). Specifically, S201 may include: allocate a virtual address (VA), that is, allocate a virtual address in the virtual address space of the second graphics processor, and mapping, that is, implement the mapping of the virtual address to the physical address through the memory management unit (MMU). S202, swap memory (swap buffers). Specifically, S202 may include: load a texture identity (ID), start a rendering task (renderjob), end a rendering task, and unload a texture identity. S203, destroy an image (destroy image). Specifically, S203 may include: demapping, that is, unmapping the virtual address and the physical address through the memory management unit, and releasing the virtual address.
[0049] During the drawing process of the second graphics processor, the texture ID needs to be loaded. Due to hardware resource limitations, there is an upper limit on the number of texture ID hardware resources in an address space (AS), while there is no upper limit on the number of texture ID software resources. Therefore, in order to make the limited texture ID hardware resources support as many texture ID software resources as possible, dynamic loading and unloading methods can be used. Specifically, at the end of the task, the second graphics processor releases idle texture ID hardware resources and time-multiplexes the limited texture ID hardware resources to improve the utilization of texture ID hardware resources.
[0050] In some possible embodiments, when there are enough texture ID hardware resources, the second graphics processor can submit tasks in parallel. Figure 3 As shown, Figure 3 A schematic diagram of a process flow of a second graphics processor executing tasks in parallel provided in an embodiment of the present application. The task chain may include task [i-1], task [i], task [i+1], and task [i+2]. Task [i-1] is a task in the queue, task [i] and task [i+1] are running tasks, and task [i+2] is a completed task. Figure 3The texture ID hardware resources in two address spaces (i.e., AS0 and AS1) are shown. Each address space includes 8 texture ID hardware resources, namely tex0, tex1, ……, tex7. Assume that task [i + 1] also needs to call 3 texture ID software resources in AS0, i.e., tex0, tex1, and tex2. Then, tex0, tex1, and tex2 in the texture ID software resources can be respectively matched one-to-one with tex4, tex5, and tex6 in AS0 of the second graphics processor. Continuing to assume that task [i] needs to call 3 texture ID software resources in AS0, i.e., tex0, tex1, and tex2. Then, tex0, tex1, and tex2 in the texture ID software resources can be respectively matched one-to-one with tex0, tex1, and tex2 in AS0 of the second graphics processor. Thus, task [i] and task [i + 1] can run in parallel.
[0051] In some possible embodiments, when there are not enough texture ID hardware resources, the second graphics processor can submit tasks serially. As Figure 4 shown, Figure 4 is a schematic flowchart of the second graphics processor serially executing tasks provided by an embodiment of the present application. Among them, task [i + 1] is a running task, and task [i] is a task waiting to run. Assume that task [i + 1] also needs to call 3 texture ID software resources in AS0, i.e., tex0, tex1, and tex2. Then, tex0, tex1, and tex2 in the texture ID software resources can be respectively matched one-to-one with tex5, tex6, and tex7 in AS0 of the second graphics processor. Continuing to assume that task [i] needs to call 6 texture ID software resources in AS0, i.e., tex0, tex1, ……, tex5. Then, tex1, tex2, ……, tex5 in the texture ID software resources can be respectively matched one-to-one with tex0, tex1, ……, tex4 in AS0 of the second graphics processor. Among them, tex0 in the texture ID software resources needs to wait for task [i + 1] to release the texture ID hardware resources before matching with the texture ID hardware resources. That is, when there are not enough texture ID hardware resources, if no idle texture ID hardware resources can be found, the task needs to queue up and wait for the previous task to release the texture ID hardware resources before continuing to submit.
[0052] In the scenario of resource conflict blocking described above, the second graphics processor needs to perform concurrent control of resource conflicts, such as submitting tasks serially. However, if there are more texture ID software resources, resource conflict blocking occurs more frequently, resulting in certain performance losses. To reduce resource conflict blocking, the number of texture ID software resources needs to be restricted. That is, the limited number of texture ID hardware resources affects the upper limit of the number of texture ID software resources, thereby reducing the performance of the second graphics processor.
[0053] Therefore, an embodiment of the present application provides a chip system, which includes a third graphics processor and a memory manager. Among them, when the third graphics processor requests to read the first data for the first time, the partition mapping table can be queried based on the first virtual address in the first read instruction, so as to read the first data from the memory. Since the partition mapping table includes the corresponding relationship between multiple first cache segment areas and multiple first segment area identifiers, and the cache granularity of each first cache segment area in the multiple first cache segment areas is the same, the first segment area identifier corresponding to the first virtual address can be directly determined. Compared with the method of sequentially searching for the first segment area identifier, the chip system provided by the embodiment of the present application can reduce overhead, and can increase the number of resources corresponding to the first segment area identifier, reduce excessive mapping or remapping, and improve the performance of the third graphics processor.
[0054] The chip system provided by the embodiment of the present application will be introduced below.
[0055] As Figure 5 shown, Figure 5 FIG. 13 is a schematic structural diagram of a chip system provided by an embodiment of the present application. The chip system 50 may include a third graphics processor 51 and a memory manager 52.
[0056] Among them, the third graphics processor 51 is configured to: in response to a graphics rendering task, output a first read instruction to the memory manager 52. The first read instruction is used to indicate reading the first data based on the first virtual address, and the first virtual address belongs to the first virtual address storage space. The first virtual address storage space includes a plurality of consecutive first virtual address areas, and each first virtual address area includes a plurality of first cache segment areas obtained based on the corresponding cache division. The first data is the data that the third graphics processor reads for the first time based on the graphics rendering task, and the data volume of the first data is less than or equal to the cache granularity of the plurality of first virtual address areas.
[0057] Exemplarily, the first data may be a texture resource, which may also be referred to as descriptor information. When the third graphics processor 51 executes a graphics rendering task, it needs to obtain the texture resource to achieve a certain surface effect in combination with relevant algorithms (such as shader algorithms). Each texture resource corresponds to a unique texture ID, and the texture ID can ensure the accuracy and integrity of the texture resource.
[0058] In the embodiment of the present application, the first data is the data first read by the third graphics processor 51, and at this time, the first data is stored in the memory. The memory may be a high bandwidth memory (HBM), or the memory may also be a double data rate synchronous dynamic random access memory (DDR SDRAM), where DDR SDRAM may be abbreviated as "DDR" for short.
[0059] Exemplarily, the first virtual address storage space may be a partial address space in the virtual address space of the third graphics processor 51, that is, the first virtual address storage space may be a continuous address space starting from the starting address of 0. The chip system 50 may divide the first virtual storage space into a plurality of continuous first virtual address areas, where the storage spaces of the plurality of first virtual address areas may be different, or the storage spaces of the plurality of first virtual address areas may also be the same. Each first virtual address area may include a plurality of first cache segment areas divided based on the corresponding cache granularity, that is, the storage space sizes of each first cache segment area in the same first virtual address area are the same.
[0060] Exemplarily, the cache granularity may be a multiple of a cache line. In one example, the cache granularity may be 512 KB, 1 MB, 2 MB, or 4 MB, etc. That is, the storage space sizes of the plurality of first cache segment areas may be 512 KB, 1 MB, 2 MB, or 4 MB, etc.
[0061] In a possible example, if the virtual partition start address of the first virtual address area is 0 and the cache granularity is N, then the virtual partition start address of the first first cache segment area among the plurality of first cache segment areas is 0, the virtual partition start address of the second first cache segment area is N, and so on. The virtual partition start address of the i-th first cache segment area is (i - 1)N. Thus, even if there are a plurality of first cache segment areas, the virtual partition start address of each first cache segment area can be determined through simple logical operations.
[0062] Among them, the memory manager 52 is used to read the first data from the memory based on the first virtual address and the partition mapping table in response to the first read instruction. The partition mapping table includes a plurality of first cache segment areas and a plurality of first segment area identifiers, and the plurality of first cache segment areas correspond to the plurality of first segment area identifiers one by one. The memory manager 52 is further used to output the first data to the third graphics processor.
[0063] Exemplarily, each first cache segment area corresponds to a first segment area identifier. Since each first cache segment area stores only one texture resource, the first segment area identifier can also be understood as the identifier of the texture resource, that is, the texture ID mentioned above.
[0064] Specifically, taking the starting address of the first virtual address storage space as 0, the size of the first virtual address storage space as 512 MB, and the first virtual address storage space including three first virtual address areas (such as slab0, slab1, and slab2) as an example, the partition mapping table is shown in Table 1. Table 1 shows the first virtual address area, the first segment area identifier, the number of texture resources, the cache granularity, and the virtual partition start address. It can be understood that Table 1 is only an example of the partition mapping table.
[0065] Table 1
[0066] First virtual address area First segment area identifier Number of texture resources Cache granularity Virtual partition start address slab0 0-63 64 512KB 0MB slab1 64-175 112 2MB 32MB slab2 176-239 64 4MB 256MB
[0067] In some possible embodiments, for the first virtual address storage space, when allocating a single texture resource, if the size of the texture resource ≤ 512 KB, the virtual address and texture ID can be allocated from slab0; if 512 KB < the size of the texture resource ≤ 2 MB, the virtual address and texture ID can be allocated from slab1; if 2 MB < the size of the texture resource ≤ 4 MB, the virtual address and texture ID can be allocated from slab2.
[0068] Specifically, slab0 includes 64 texture IDs and occupies 32 MB of address space. slab1 includes 112 texture IDs and occupies 224 MB of address space. slab2 includes 64 texture IDs and occupies 256 MB of address space.
[0069] Among them, the memory manager 52 is specifically used to: determine the first target segment area identifier based on the first virtual address and the partition mapping table, where the first target segment area identifier is the first segment area identifier corresponding to the first cache segment area where the first data is located. Read the first data from the memory according to the first target segment area identifier.
[0070] Exemplarily, the memory manager 52 may query the partition mapping table based on the first virtual address to determine the first target segment area identifier. Figure 6 As shown, Figure 6 A mapping diagram of a first virtual address storage space provided in an embodiment of the present application. Figure 6 , M texture ID software resources in the third graphics processor are shown, for example, tex[0], tex[1], tex[2], ..., tex[M], Figure 6 Also shown are N texture ID hardware resources, such as tex[0], tex[1], tex[2], ..., tex[N]. Among them, the texture ID software resources and the texture ID hardware resources correspond one-to-one, that is, tex[1] in the texture ID software resources corresponds to tex[1] in the texture ID hardware resources, and tex[2] in the texture ID software resources corresponds to tex[2] in the texture ID hardware resources, etc. Therefore, when the software submits a task, since the texture ID software resources and the texture ID hardware resources of the third graphics processor correspond one-to-one, and the cache granularity of each first cache segment area is the same, a simple operation can be performed through the first virtual address to determine the location of the texture ID hardware resource, that is, the first target segment area identifier. Since the overhead of simple operations is small, the number of texture ID hardware resources does not need to be constrained, so that the third graphics processor can support more texture ID software resources.
[0071] Specifically, Figure 7 As shown, Figure 7 Another processing flow chart of a third graphics processor performing a task provided for the implementation of the present application. Wherein, when the third graphics processor performs a graphics rendering task, it reads the first data in the first virtual address storage space, and the processing flow may include: S701, creating an image. Specifically, S701 may include: allocating a virtual address, that is, allocating a virtual address in the first virtual address storage space of the third graphics processor, and mapping, that is, implementing the mapping of the virtual address to the physical address through the memory management unit. S702, swapping memory. Specifically, S702 may include starting a drawing task and ending a drawing task. S703, destroying the image. Specifically, S703 may include demapping, that is, unmapping the virtual address and the physical address through the memory management unit, and releasing the virtual address.
[0072] That is, compared to Figure 2In the flowchart shown in , the flowchart of the third graphics processor provided by the embodiments of the present application reduces the operations of loading texture IDs and unloading texture IDs, reduces the overhead required for mapping and unmapping, and can improve the performance of the third graphics processor.
[0073] Among them, the memory manager 52 is specifically configured to: determine a first target segmentation area identifier based on a first virtual address, virtual partition start addresses of multiple first virtual address areas, and a cache granularity corresponding to the first virtual address area.
[0074] Exemplarily, the memory manager 52 may first query the partition mapping table based on the first virtual address to determine the identifier of the first virtual address area (such as slab0, slab1, and slab2), and further determine the first target segmentation area identifier based on the virtual partition start address of the first virtual address area and the cache granularity corresponding to the first virtual address area.
[0075] Continuing to take the partition mapping table as Table 1 as an example, assuming the first virtual address is 261 MB, then the identifier of the first virtual address area corresponding to the first virtual address can be determined as slab2. Since the virtual partition start address of slab2 is 256 MB and the cache granularity of slab2 is 4 MB, the first virtual address corresponds to the second first cache segmentation area in slab2. Among them, the first segmentation area identifier of slab2 is tex
[176] to tex
[239] . Thus, the first target segmentation area identifier corresponding to the first virtual address is tex
[177] .
[0076] Optionally, the partition mapping table further includes multiple second cache segmentation areas and multiple second segmentation area identifiers. The multiple second cache segmentation areas form a continuous second virtual address storage space. The cache granularity of the second cache segmentation area is greater than the cache granularity of the multiple first cache segmentation areas. The multiple second cache segmentation areas are in one-to-one correspondence with the multiple second segmentation area identifiers.
[0077] Referring to the above example, when the maximum cache granularity of the first cache segmentation area is 4 MB, the cache granularity of the second cache segmentation area is greater than 4 MB, that is, if the size of the texture resource > 4 MB, virtual addresses and texture IDs can be allocated from the second virtual address storage space. In a possible example, the second virtual address storage space may form a continuous storage space with the first virtual address storage space, or the second virtual address storage space may also form a discontinuous storage space with the first virtual address storage space. In addition, the cache granularity of each second cache segmentation area may be the same, or the cache granularity of each second cache segmentation area may also be different.
[0078] Among them, the third graphics processor 51 is further configured to: in response to a graphics rendering task, output a second read instruction to the memory manager 52, where the second read instruction is used to indicate reading second data based on a second virtual address, the second virtual address belongs to a second virtual address storage space, the data volume of the second data is greater than the cache granularity of a plurality of first virtual address areas, and the second data is the data first read by the third graphics processor 51 based on the graphics rendering task.
[0079] Exemplarily, the second data may also be a texture resource. The second data in the embodiments of the present application is the data first read by the third graphics processor.
[0080] Among them, the memory manager 52 is further configured to: in response to the second read instruction, read the second data from the memory based on the second virtual address, the partition mapping table, and the resource mapping table. The resource mapping table is used to indicate the mapping relationship between the second virtual address and the second segment area identifier. The memory manager 52 is further configured to output the second data to the third graphics processor 51.
[0081] Exemplarily, due to limited texture ID hardware resources, texture ID software resources need to be dynamically loaded and unloaded. Among them, the resource mapping table stores the mapping relationship between texture ID software resources and texture ID hardware resources. Continuing to refer to Figure 3 , when the third graphics processor executes tasks serially, when executing task [i + 1], the resource matching table stores the mapping relationship between tex0 in the texture ID software resource of task [i + 1] and tex5 in the texture ID hardware resource. When starting to execute task [i] after task [i + 1] ends, the resource matching table stores the updated mapping relationship between tex0 in the texture ID software resource of task [i] and tex5 in the texture ID hardware resource, that is, a new mapping is performed on tex5 in the texture ID hardware resource, so that tex5 in the texture ID hardware resource can be time-division multiplexed.
[0082] That is to say, if the third graphics processor wants to access the second data in the second virtual address storage space, it does not depend on a pre-allocated address space, and can allocate and release the address space in real time, providing greater flexibility and scalability for the chip system.
[0083] Optionally, the memory manager 52 is specifically configured to: determine a second target segment area identifier based on the second virtual address and the virtual partition start addresses of a plurality of second cache segment areas, and read the second data from the memory based on the second target segment area identifier and the resource matching table.
[0084] Exemplarily, the memory manager 52 may query the partition mapping table and the address matching table based on the second virtual address to determine the second target segment area identifier. As Figure 8 shown, Figure 8 FIG. 4 is a mapping diagram of a second virtual address storage space provided by an embodiment of the present application. Figure 8 The M texture ID software resources in the third graphics processor are shown, such as tex[0], tex[1], tex[2], ……, tex[M], Figure 8 and the N texture ID hardware resources are also shown, such as tex[0], tex[1], tex[2], ……, tex[N]. Among them, the mapping relationship between the texture ID software resources and the texture ID hardware resources is stored in the resource matching table. For example, tex[1] of the texture ID software resources corresponds to tex[4] of the texture ID hardware resources, tex[2] of the texture ID software resources corresponds to tex[1] of the texture ID hardware resources, and tex[3] of the texture ID software resources corresponds to tex[2] of the texture ID hardware resources. The texture ID software resources and the texture ID hardware resources are dynamically bound. When the third graphics processor executes a graphics rendering task, the memory manager determines the location of the texture ID hardware resources, that is, the second target segment area identifier, by sequential search according to the second virtual address.
[0085] Thus, in the second virtual address storage space, although texture IDs are dynamically loaded and unloaded, certain compatibility can be provided for the chip system, enabling the chip system to support larger resolution scenarios.
[0086] Among them, slab0, slab1, and slab2 in the partition mapping table can be understood as fixed partitions, and slab3 can be understood as a flexible partition.
[0087] Thus, the partition mapping table of the virtual address space in the third graphics processor is shown in Table 2. Table 2 shows the texture IDs in each virtual address area of multiple address spaces (such as AS0, AS1, ……, AS7). It can be understood that Table 2 is only an example of the virtual address area of the virtual address space of the third graphics processor, and the embodiments of the present application do not specifically limit the number of address spaces and the number of texture IDs.
[0088] Table 2
[0089]
[0090]
[0091] Among them, as Figure 9 shown, Figure 9 This is a mapping diagram of the virtual address space of a third graphics processor provided by an embodiment of the present application. Figure 9 Three first virtual address areas (slab0, slab1, and slab2) and a second virtual address area (slab3) are shown in Figure 9 . Among them, slab0 includes 64 first cache segment areas with a cache granularity of 512 KB, occupying a total address space of 32 MB, slab1 includes 112 first cache segment areas with a cache granularity of 2 MB, occupying a total address space of 224 MB, slab2 includes 64 first cache segment areas with a cache granularity of 4 MB, occupying a total address space of 256 MB, and slab3 includes 8 second cache segment areas with a cache granularity of any size, occupying a total address space of 512 MB.
[0092] Among them, the texture ID software resources and texture ID hardware resources in each first virtual address area correspond one by one. During the process of the third graphics processor executing a graphics rendering task, there is no need to perform dynamic loading and unloading of texture IDs, which solves the limitation of the number of texture ID hardware resources, can support more texture ID software resources, and reduces blocking, improving the performance of the chip system. The texture ID software resources and texture ID hardware resources in the second virtual address area are dynamically bound. During the process of the third graphics processor executing a graphics rendering task, texture ID hardware resources are dynamically loaded and unloaded, and the limited texture ID hardware resources are multiplexed in a time-sharing manner.
[0093] Optionally, the memory manager 52 further includes a cache for storing third data, where the third data is data that the third graphics processor 51 requests to read again based on a graphics rendering task. The third graphics processor 51 is further configured to output a third read instruction to the memory manager 52 in response to a graphics rendering task, where the third read instruction is used to indicate reading the third data. The memory manager 52 is further configured to read the third data from the cache in response to the third read instruction and output the third data to the third graphics processor 51.
[0094] Exemplarily, when the third graphics processor 51 requests to read the third data for the first time, the memory manager 52 reads the third data from the memory, outputs the third data to the third graphics processor 51, and stores the third data in the cache. When the third graphics processor 51 requests to read the third data again, the memory manager 52 parses the third read instruction. Since the virtual address carried by the third read instruction matches the previous access record, it indicates that the third data is stored in the cache. The memory manager 52 reads the third data from the cache and outputs the third data to the third graphics processor 51.
[0095] Thus, when the third graphics processor 51 outputs a read instruction, the memory manager 52 responds to the read instruction. If the data to be read by the read instruction is the data requested to be read for the first time, it means that the memory manager 52 needs to update the data. The memory manager 52 reads the data to be read by the read instruction from the memory and stores the data in the cache. If the data to be read by the read instruction is the data requested to be read again and the status of the data is valid, that is, the data is the latest, the memory manager 52 can directly read the data to be read by the read instruction from the cache without reloading the data, thereby improving the data processing speed.
[0096] Optionally, the first read instruction further includes an identification bit, and the identification bit is used to indicate the address path type corresponding to the first virtual address. The chip system may further include a snooping filter 53. When the identification bit is a first value, the snooping filter 53 is used to: output the first read instruction to the memory manager 52.
[0097] Exemplarily, the address path type may include a normal path, a standard compression path, and a secure path. Each address path has an identification bit to ensure data isolation and security. In a possible example, the identification bit may be represented by 2 bits (bit, b). If the identification bit is "00", the address path type corresponding to the first virtual address is a normal path. If the identification bit is "01", the address path type corresponding to the first virtual address is a standard compression path. If the identification bit is "10", the address path type corresponding to the first virtual address is a secure path.
[0098] Among them, the snooping filter 53 can snoop on the read operation instructions of the third graphics processor 51, etc., effectively identify and select the address access of the third graphics processor 51, and ensure that the memory manager 52 only processes critical address traffic.
[0099] Exemplarily, the first value may be "01", that is, the address path type corresponding to the first virtual address is the standard compression path. When the snooping filter 53 identifies that the first read request is of the address path type of the standard compression path, it can output the first read request to the memory manager 52 for subsequent processing.
[0100] Among them, as Figure 10 shown, Figure 10 is a schematic structural diagram of another chip system provided by an embodiment of the present application. The memory manager 52 may include a partition manager 521, a decoder 522, and a direct memory accessor 523.
[0101] Specifically, the snooping filter 53 is used to identify the read instructions output by the third graphics processor 51. The read instructions may include a virtual address and an identification bit. If the identification bit is the first value, that is, the address path type corresponding to the virtual address is the standard compression path, the snooping filter 53 sends the first read instruction to the partition manager 521.
[0102] If the virtual address is within the range of the first virtual address storage space, the partition manager 521 queries the partition mapping table based on the virtual address to obtain the identification of the first virtual address area and the first segment area identification. If the data corresponding to the read instruction is the data that the third graphics processor 51 requests to read for the first time, the partition manager 521 sends the virtual address and the first segment area identification to the direct memory accessor 523, and the direct memory accessor 523 reads the data from the memory. If the data corresponding to the read instruction is the data that the third graphics processor 51 requests to read again, and the data in the cache is in a valid state, the partition management 521 reads the data from the cache.
[0103] If the virtual address is within the range of the second virtual address storage space, the partition manager 521 outputs the read instruction to the decoder 522. If the data corresponding to the read instruction is the data that the third graphics processor 51 requests to read again, the decoder 522 queries the partition mapping table based on the virtual address to obtain the second segment area identification. If the data corresponding to the read instruction is the data that the third graphics processor 51 requests to read for the first time, the decoder 522 sends the virtual address and the second segment area identification to the direct memory accessor 523, and the direct memory accessor 523 reads the data from the memory. If the data corresponding to the read instruction is the data that the third graphics processor 51 requests to read again, and the data in the cache is in a valid state, the decoder 521 reads the data from the cache.
[0104] Specifically, the structural diagram of the decoder 522 is as Figure 11 shown. The decoder may include a determination unit and multiple comparators. Figure 11Eight comparators are shown, namely comparator 0, comparator 1, comparator 2, …, comparator 7. Among them, it is assumed that the second virtual address storage space includes eight second cache segment areas, and each second cache segment area corresponds to a virtual partition start address, that is, the eight second cache segment areas correspond to eight virtual partition start addresses, which are address 0, address 1, address 2, …, address 7 respectively. The inputs of each comparator are the second virtual address and the corresponding virtual partition start address respectively. Among them, the comparator is used to compare the sizes of the second virtual address and the corresponding virtual partition start address and output a comparison result. The determination unit is used to determine that the second segment area identifier corresponding to the smallest comparison result in the comparison results is the second target segment area identifier.
[0105] Applied to the above chip system, a method for reading data provided by an embodiment of the present application will be introduced below.
[0106] As Figure 12 shown, Figure 12 is a flowchart of a method for reading data provided by an embodiment of the present application. The method includes the following processes.
[0107] S1201. In response to a graphics rendering task, the third graphics processor outputs a first read instruction to the memory manager.
[0108] Among them, the first read instruction is used to indicate reading the first data based on the first virtual address. The first virtual address belongs to the first virtual address storage space. The first virtual address storage space includes a plurality of consecutive first virtual address areas, and each first virtual address area includes a plurality of first cache segment areas obtained by dividing based on the corresponding cache granularity. The first data is the data first read by the third graphics processor based on the graphics rendering task, and the data volume of the first data is less than or equal to the cache granularity of the plurality of first virtual address areas.
[0109] S1202. In response to the first read instruction, the memory manager reads the first data from the memory based on the first virtual address and the partition mapping table and outputs the first data to the third graphics processor.
[0110] Among them, the partition mapping table includes a plurality of first cache segment areas and a plurality of first segment area identifiers, and the plurality of first cache segment areas correspond to the plurality of first segment area identifiers one by one.
[0111] Exemplarily, when the third graphics processor requests to read the first data for the first time, the partition mapping table can be queried based on the first virtual address in the first read instruction, so as to read the first data from the memory. Among them, since the partition mapping table includes the corresponding relationship between multiple first cache segment areas and multiple first segment area identifiers, and the cache granularity of each first cache segment area in the multiple first cache segment areas is the same, the first segment area identifier corresponding to the first virtual address can be directly determined. Compared with the method of sequentially searching for the first segment area identifier, the chip system provided by the embodiments of the present application can reduce the overhead, and can increase the number of resources corresponding to the first segment area identifier, reduce excessive mapping or remapping, and improve the performance of the third graphics processor.
[0112] Optionally, S1202 may include: The memory manager determines the first target segment area identifier based on the first virtual address and the partition mapping table, and the first target segment area identifier is the first segment area identifier corresponding to the first cache segment area where the first data is located. Read the first data from the memory according to the first target segment area identifier.
[0113] Optionally, S1202 may include: The memory manager determines the first target segment area identifier based on the first virtual address, the virtual partition start address of multiple first virtual address areas, and the cache granularity corresponding to the first virtual address area.
[0114] Optionally, the method may further include: The third graphics processor outputs a second read instruction to the memory manager in response to a graphics rendering task. The second read instruction is used to indicate reading the second data based on the second virtual address, the second virtual address belongs to the second virtual address storage space, the data volume of the second data is greater than the cache granularity of multiple first virtual address areas, and the second data is the data read by the third graphics processor for the first time based on the graphics rendering task. The memory manager reads the second data from the memory in response to the second read instruction based on the second virtual address, the partition mapping table, and the resource matching table. The resource matching table is used to indicate the mapping relationship between the second virtual address and the second segment area identifier. Output the second data to the third graphics processor.
[0115] Optionally, the method may further include: The memory manager determines the second target segment area identifier based on the second virtual address and the virtual partition start addresses of multiple second cache segment areas. Read the second data from the memory based on the second target segment area identifier and the resource matching table.
[0116] Optionally, the method may further include: The third graphics processor outputs a third read instruction to the memory manager in response to a graphics rendering task, and the third read instruction is used to indicate reading the third data. The memory manager reads the third data from the cache in response to the third read instruction; Output the third data to the third graphics processor.
[0117] Optionally, the method may further include: when the identification bit is the first value, the snooping filter outputs a first read instruction to the memory manager.
[0118] Optionally, an embodiment of the present application further provides an electronic device, which includes one or more interface circuits and one or more chip systems, and the interface circuits and the chip systems are interconnected through lines.
[0119] It can be understood that, in order to implement the above functions, the electronic device includes corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.
[0120] An embodiment of the present application further provides a computer storage medium, in which computer instructions are stored. When the computer instructions run on the electronic device, the electronic device is enabled to execute the above-related method steps to implement the method for reading data in the above embodiments.
[0121] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above-related steps to implement the method for reading data executed by the electronic device in the above embodiments.
[0122] In addition, an embodiment of the present application further provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other; wherein, the memory is used to store computer execution instructions. When the device runs, the processor may execute the computer execution instructions stored in the memory so that the chip executes the method for reading data executed by the electronic device in each of the above method embodiments.
[0123] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0124] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0125] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0126] The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0128] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.
[0129] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A chip system, characterized in that, it includes: a graphics processor and a memory manager; The graphics processor is configured to: in response to a graphics rendering task, output a first read instruction to the memory manager, the first read instruction being used to indicate reading first data based on a first virtual address; the first virtual address belongs to a first virtual address storage space, the first virtual address storage space includes a plurality of consecutive first virtual address areas, and each of the first virtual address areas includes a plurality of first cache segment areas obtained by dividing based on a corresponding cache granularity; the first data is the data first read by the graphics processor based on the graphics rendering task, and the data volume of the first data is less than or equal to the cache granularity of the plurality of first virtual address areas; The memory manager is configured to: in response to the first read instruction, read the first data from the memory based on the first virtual address and a partition mapping table, the partition mapping table including the plurality of first cache segment areas and a plurality of first segment identifiers; the plurality of first cache segment areas and the plurality of first segment identifiers are in one-to-one correspondence; output the first data to the graphics processor.
2. The chip system according to claim 1, characterized in that, the memory manager is specifically configured to: determine a first target segment identifier based on the first virtual address and the partition mapping table, the first target segment identifier being the first segment identifier corresponding to the first cache segment area where the first data is located; read the first data from the memory according to the first target segment identifier.
3. The chip system according to claim 2, characterized in that, the memory manager is specifically configured to: determine the first target segment identifier based on the first virtual address, the virtual partition start address of the plurality of first virtual address areas, and the cache granularity corresponding to the first virtual address area.
4. The chip system according to any one of claims 1-3, characterized in that, the partition mapping table further includes a plurality of second cache segment areas and a plurality of second segment identifiers; the plurality of second cache segment areas form a continuous second virtual address storage space; the cache granularity of the second cache segment areas is greater than the cache granularity of the plurality of first cache segment areas; the plurality of second cache segment areas and the plurality of second segment identifiers are in one-to-one correspondence; the graphics processor is further configured to: in response to the graphics rendering task, output a second read instruction to the memory manager, the second read instruction being used to indicate reading second data based on a second virtual address, the second virtual address belongs to the second virtual address storage space, the data volume of the second data is greater than the cache granularity of the plurality of first virtual address areas, and the second data is the data first read by the graphics processor based on the graphics rendering task; The memory manager is further configured to: In response to the second read instruction, read the second data from the memory based on the second virtual address, the partition mapping table, and the resource matching table; the resource matching table is used to indicate the mapping relationship between the second virtual address and the second segment area identifier; Output the second data to the graphics processor.
5. The chip system according to claim 4, wherein, the memory manager is specifically configured to: determine a second target segment area identifier based on the second virtual address and the virtual partition start address of the multiple second cache segment areas; read the second data from the memory based on the second target segment area identifier and the resource matching table.
6. The chip system according to any one of claims 1-5, wherein, the memory manager further includes a cache; the cache is used to store third data, and the third data is data that the graphics processor requests to read again based on the graphics rendering task; the graphics processor is further configured to: in response to the graphics rendering task, output a third read instruction to the memory manager, and the third read instruction is used to indicate reading the third data; the memory manager is further configured to: in response to the third read instruction, read the third data from the cache; output the third data to the graphics processor.
7. The chip system according to any one of claims 1-6, wherein, the first read instruction further includes a flag bit, and the flag bit is used to indicate the address path type corresponding to the first virtual address; the chip system further includes: a snooping filter; when the flag bit is a first value, the snooping filter is configured to: output the first read instruction to the memory manager.
8. A method for reading data, wherein, the method is applied to a chip system, and the chip system includes a graphics processor and a memory manager; the method includes: the graphics processor, in response to a graphics rendering task, outputs a first read instruction to the memory manager, and the first read instruction is used to indicate reading first data based on a first virtual address; the first virtual address belongs to a first virtual address storage space, and the first virtual address storage space includes a plurality of consecutive first virtual address areas, and each of the first virtual address areas includes a plurality of first cache segment areas obtained by dividing based on a corresponding cache granularity; the first data is data that the graphics processor reads for the first time based on the graphics rendering task, and the data volume of the first data is less than or equal to the cache granularity of the plurality of first virtual address areas; the memory manager, in response to the first read instruction, reads the first data from the memory based on the first virtual address and the partition mapping table, and the partition mapping table includes the plurality of first cache segment areas and a plurality of first segment area identifiers; the plurality of first cache segment areas correspond to the plurality of first segment area identifiers one by one; output the first data to the graphics processor.
9. The method according to claim 8, wherein, The memory manager reads the first data from the memory based on the first virtual address and the partition mapping table, including: The memory manager determines a first target segment area identifier based on the first virtual address and the partition mapping table, where the first target segment area identifier is the first segment area identifier corresponding to the first cache segment area where the first data is located; Read the first data from the memory according to the first target segment area identifier.
10. The method according to claim 9, wherein, The memory manager determines a first target segment area identifier based on the first virtual address and the partition mapping table, including: The memory manager determines the first target segment area identifier based on the first virtual address, the virtual partition start address of the multiple first virtual address areas, and the cache granularity corresponding to the first virtual address area.
11. The method according to any one of claims 8-10, wherein, The partition mapping table further includes a plurality of second cache segment areas and a plurality of second segment area identifiers; the plurality of second cache segment areas form a continuous second virtual address storage space; the cache granularity of the second cache segment area is greater than the cache granularity of the plurality of first cache segment areas; The plurality of second cache segment areas correspond to the plurality of second segment area identifiers one by one; the method further includes: In response to the graphics rendering task, the graphics processor outputs a second read instruction to the memory manager, where the second read instruction is used to indicate reading second data based on a second virtual address, the second virtual address belongs to the second virtual address storage space, the data volume of the second data is greater than the cache granularity of the plurality of first virtual address areas, and the second data is the data first read by the graphics processor based on the graphics rendering task; In response to the second read instruction, the memory manager reads the second data from the memory based on the second virtual address, the partition mapping table, and the resource matching table; the resource matching table is used to indicate the mapping relationship between the second virtual address and the second segment area identifier; and outputs the second data to the graphics processor.
12. The method according to claim 11, wherein, The memory manager reads the second data from the memory based on the second virtual address, the partition mapping table, and the resource matching table, including: The memory manager determines a second target segment area identifier based on the second virtual address and the virtual partition start address of the plurality of second cache segment areas; Read the second data from the memory based on the second target segment area identifier and the resource matching table.
13. The method according to any one of claims 8-12, wherein, The memory manager further includes a cache; the cache is used to store third data, and the third data is the data requested to be read again by the graphics processor based on the graphics rendering task; the method further includes: In response to the graphics rendering task, the graphics processor outputs a third read instruction to the memory manager, and the third read instruction is used to indicate reading the third data; In response to the third read instruction, the memory manager reads the third data from the cache and outputs the third data to the graphics processor.
14. The method according to any one of claims 8-15, wherein, the first read instruction further includes an identification bit, and the identification bit is used to indicate the address path type corresponding to the first virtual address; the chip system further includes: a snooping filter; and the method further includes: when the identification bit is a first value, the snooping filter outputs the first read instruction to the memory manager.
15. An electronic device, wherein, comprising one or more interface circuits, and one or more chip systems according to any one of claims 1-7, and the interface circuits and the chip systems are interconnected by lines.
16. A computer-readable storage medium, wherein, comprising computer instructions, when the computer instructions run on an electronic device, enabling the electronic device to execute the method according to any one of claims 8-14 above.
Citation Information
Cited By
Data processing method and device, graphics processor, electronic equipment and storage medium
CN120634835A