Geometric processing method, graphics processor and computer equipment
By introducing a method in which multiple sub-processing units are used in parallel to process data and write it to the storage space in the graphics processor, the serial writing problem in the stream output stage is solved, the data writing efficiency is improved, the waiting time is reduced, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202511279855.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
AI Technical Summary
In the stream output stage of the graphics processor, although the sub-processing units can process tasks in parallel in the geometry shader or vertex shader stage, the data is written to the storage space in serial, resulting in a waste of time.
By introducing multiple sub-processing units into the graphics processor, the first sub-processing unit is used to obtain the data amount and the write start address, and the write start address of the second sub-processing unit is predetermined, so that the second sub-processing unit can execute tasks in parallel and write to the storage space, reducing waiting time.
The efficiency of the sub-processing unit in writing output data into the storage space is improved, the waiting time is reduced, parallel data writing is achieved, and the overall processing efficiency is improved.
Smart Images

Figure CN120765448A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to, but is not limited to, the field of graphics processing technology, and in particular to a geometry processing method, an image processor, and a computer device. Background Art
[0002] In the processing flow of a graphics processing unit (GPU), the main purpose of the streamout stage is to continuously output data processed by the geometry shader stage or vertex shader stage to one or more buffers in the storage space. In related technologies, each sub-processing unit can process corresponding tasks in parallel during the geometry shader stage or vertex shader stage. However, during the streamout stage, the output data of each sub-processing unit must be written serially to the storage space. Therefore, even if each sub-processing unit processes the corresponding task in parallel, the data is written serially, which does not save time. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure at least provide a geometry processing method, a graphics processor, and a computer device.
[0004] The technical solution of the embodiment of the present disclosure is implemented as follows: An embodiment of the present disclosure provides a geometry processing method applied to a graphics processor, wherein the graphics processor includes at least two sub-processing units, each of which is configured to process a corresponding data segment in geometric data to be processed. The method includes: In response to receiving a first task, the first sub-processing unit obtains a data volume and a first write start address corresponding to first output data to be output by the first task; the first task is the last task of the first data segment corresponding to the first sub-processing unit in the target processing stage; The second sub-processing unit executes a second task to generate second output data, and writes the second output data into the storage space based on a second write start address; the second write start address is determined based on the data amount and the first write start address, the second task is the first task of the second data segment corresponding to the second sub-processing unit in the target processing stage, and the second data segment is the next data segment of the first data segment.
[0005] An embodiment of the present disclosure provides a graphics processor, comprising at least two sub-processing units, each of which is configured to process a corresponding data segment in geometric data to be processed, wherein: The first sub-processing unit is configured to, in response to receiving a first task, obtain a data volume and a first write start address corresponding to first output data to be output by the first task; the first task being the last task of the first data segment corresponding to the first sub-processing unit in the target processing stage; The second sub-processing unit is used to execute a second task to generate second output data, and write the second output data into the storage space based on a second write start address; the second write start address is determined based on the data amount and the first write start address, the second task is the first task of the second data segment corresponding to the second sub-processing unit in the target processing stage, and the second data segment is the next data segment of the first data segment.
[0006] An embodiment of the present disclosure provides a computer device including the aforementioned graphics processor.
[0007] In the embodiment of the present disclosure, the first task is the last task of the first data segment corresponding to the first sub-processing unit in the target processing stage. The first sub-processing unit executes the first task and obtains the data amount and the first write start address corresponding to the first output data to be output by the first task; the second write start address is determined based on the data amount and the first write start address. The second task is the first task of the second data segment corresponding to the second sub-processing unit in the target processing stage. The second data segment is the next data segment of the first data segment. The second sub-processing unit executes the second task to generate the second output data and writes the second output data into the storage space based on the second write start address. In this way, after the first sub-processing unit obtains the data amount and the first write start address, it can determine the second write start address in advance based on the data amount and the first write start address. The second sub-processing unit can then start executing the second task based on the second write start address without having to wait for the first sub-processing unit to complete the first task. This reduces the time spent waiting for the first sub-processing unit to complete the first task, thereby improving the efficiency of each sub-processing unit in writing the corresponding output data into the storage space.
[0008] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0010] Figure 1 A schematic diagram of the implementation process of a geometric processing method provided in an embodiment of the present disclosure Figure One ; Figure 2A schematic diagram of the structure of a graphics processor provided in an embodiment of the present disclosure Figure One ; Figure 3 A schematic diagram of the structure of a graphics processor provided in an embodiment of the present disclosure Figure Two ; Figure 4 A schematic diagram of the implementation process of a geometric processing method provided in an embodiment of the present disclosure Figure Two ; Figure 5 A schematic diagram of writing data into a storage space provided by an embodiment of the present disclosure; Figure 6 A schematic diagram of a method for processing data in parallel by multiple sub-processing units provided in an embodiment of the present disclosure Figure One ; Figure 7 A schematic diagram of sequentially restoring output data corresponding to multiple sub-processing units provided by an embodiment of the present disclosure; Figure 8 A schematic diagram of the implementation process of a geometric processing method provided in an embodiment of the present disclosure Figure Three ; Figure 9 A schematic diagram of a method for processing data in parallel by multiple sub-processing units provided in an embodiment of the present disclosure Figure Two ; Figure 10 A schematic diagram of the implementation process of a geometric processing method provided in an embodiment of the present disclosure Figure Four . DETAILED DESCRIPTION
[0011] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the technical solutions of the present disclosure are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting the present disclosure. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0012] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0013] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure.
[0015] The embodiments of the present disclosure provide a geometry processing method, applied to a graphics processor, the graphics processor comprising at least two sub-processing units, each of the sub-processing units being configured to process a corresponding data segment in the geometry data to be processed, Figure 1 The embodiments of the present disclosure provide a geometry processing method Figure One As shown in the figure, the method comprises the following steps S101 and S102: Figure 1 The step S101 comprises the following steps: The step S101 comprises the following steps: Here, the first task is the last task to be processed by the first sub-processing unit in the target processing stage; the first output data to be output by the first task is written into the storage space from the first write start address.
[0016] Here, the data amount corresponding to the first output data can be the vertex number of the vertex data to be output by the first sub-processing unit in the target processing stage.
[0017] Here, the first data segment is the data segment corresponding to the first sub-processing unit in the geometry data.
[0018] In some embodiments, the geometry data to be processed can be divided into multiple data segments, and each data segment is allocated to a corresponding sub-processing unit for processing, wherein the first data segment is the data segment allocated to the first sub-processing unit for processing.
[0019] In some embodiments, since the data amount corresponding to the first output data is the data amount corresponding to the first output data to be output by the last task executed by the first sub-processing unit, the data amount corresponding to the first output data can be determined according to the output data amount corresponding to the output data of all tasks executed by the first sub-processing unit and the total data amount to be output by the first sub-processing unit.
[0020] In some embodiments, the first sub-processing unit calculates in advance the data amount corresponding to the first output data to be output by the first task in response to the first task, so that the second write start address corresponding to the second task executed by the second sub-processing unit can be determined according to the data amount corresponding to the first output data and the corresponding first write start address.
[0021] In some implementations, the second write start address may be a write end address corresponding to the first output data.
[0022] In some implementations, during geometry processing, a graphics processor includes multiple sub-processing units, each of which can execute tasks in parallel.
[0023] In some embodiments, the geometric data to be processed may include vertex data, primitive data, texture data, and lighting data.
[0024] In some embodiments, the target processing stage includes a vertex shader stage and a geometry shader stage, and the first task may be the last task of the vertex shader stage or the last task of the geometry shader stage.
[0025] In some implementations, vertex data is processed in a vertex shader stage, and primitive data is processed in a geometry shader stage. The primitive data may include vertices, line segments, triangles, and the like.
[0026] In some embodiments, during the vertex shader stage, vertex data, including vertex positions, normals, texture coordinates, etc., is input. When the vertex shader is executed, the input vertex data is transformed from model space to world space or view space, and then a projection transformation is applied to transform the vertex coordinates from 3D space to 2D space. The vertex shader can also perform transformations such as translation, rotation, and scaling on the vertex data, and output the transformed vertex data to the next stage.
[0027] In some embodiments, the geometry shader stage can receive input data from the vertex shader or other stages and transform, merge, or split the input primitives consisting of one or more vertices to generate new primitives. This can transform the vertex data into completely different primitives, generating more vertices than the original. Each geometry shader stage generates a different amount of output data, making parallel processing in the geometry shader more complex than in other stages.
[0028] In some embodiments, the first sub-processing unit executes the first task to generate first output data, and obtains a data volume and a first write start address corresponding to the first output data.
[0029] In some implementations, the first write start address is an end address for writing output data of a previous task of the first task into the storage space.
[0030] In some embodiments, when starting to execute the first task, the first sub-processing unit obtains the data amount in advance, and determines the write end address corresponding to the first output data according to the data amount and the first write start address.
[0031] Step S102: The second sub-processing unit executes a second task to generate second output data, and writes the second output data into the storage space based on a second write start address; the second write start address is determined based on the data amount and the first write start address, the second task is a first task of a second data segment corresponding to the second sub-processing unit in the target processing stage, and the second data segment is a next data segment of the first data segment.
[0032] Here, the second data segment is a data segment corresponding to the second sub-processing unit in the geometry data.
[0033] The second sub-processing unit is a next sub-processing unit of the first sub-processing unit, that is, the second data segment corresponding to the second sub-processing unit is a next data segment of the first data segment corresponding to the first sub-processing unit.
[0034] The second task is a first task to be processed by the second sub-processing unit in the target processing stage; and the second output data to be output by the second task is written into the storage space from the second write start address.
[0035] In some embodiments, the geometry data to be processed can be divided into multiple data segments, and each data segment is allocated to a corresponding sub-processing unit for processing, wherein the first data segment is a data segment allocated to the first sub-processing unit for processing, the second data segment is a data segment allocated to the second sub-processing unit for processing, and the second data segment is a next data segment connected to the first data segment in the geometry data to be processed, that is, the start address of the second data segment is the end address of the first data segment.
[0036] In some embodiments, after the first sub-processing unit determines the write end address corresponding to the first output data, the second sub-processing unit can start executing the second task without waiting for the first sub-processing unit to finish executing the first task, and the write end address corresponding to the first output data is taken as the second write start address corresponding to the second output data.
[0037] In some embodiments, the storage space can be a global memory (Global Memory), which can include a dynamic random-access memory (DRAM, Dynamic Random-Access Memory) on a GPU motherboard.
[0038] In some embodiments, the second output data is written into the storage space in a stream output stage. The stream output stage is located after the vertex shader and the geometry shader and before the rasterization stage, and the main purpose is to continuously output vertex data in an inactive state in the vertex shader and the geometry shader stage to the storage space, instead of continuing to forward to the next stage (i.e., the rasterization stage).
[0039] In the embodiment of the present disclosure, the first task is the last task of the first data segment corresponding to the first sub-processing unit in the target processing stage. The first sub-processing unit executes the first task and obtains the data amount and the first write start address corresponding to the first output data to be output by the first task; the second write start address is determined based on the data amount and the first write start address. The second task is the first task of the second data segment corresponding to the second sub-processing unit in the target processing stage. The second data segment is the next data segment of the first data segment. The second sub-processing unit executes the second task to generate second output data and writes the second output data into the storage space based on the second write start address. In this way, after the first sub-processing unit obtains the data amount and the first write start address, it can determine the second write start address in advance based on the data amount and the first write start address. The second sub-processing unit can then start executing the second task based on the second write start address without having to wait for the first sub-processing unit to complete the first task. This reduces the time it takes to wait for the first sub-processing unit to complete the first task, thereby improving the efficiency of each sub-processing unit writing the corresponding output data into the storage space.
[0040] In some embodiments, the second write start address is determined based on the data volume and the first write start address before the first output data is written into the storage space; the above method may further include the following step S111: Step S111: the first sub-processing unit executes the first task to generate the first output data, and writes the first output data into a storage space based on the first write start address.
[0041] In some embodiments, the first sub-processing unit performs the first task, generates first output data, and sends the first write start address and the first output data to the data write module; the data write module uses the data write port corresponding to the first sub-processing unit to write the first output data into the storage space according to the first write start address.
[0042] In some embodiments, before writing the first output data into the storage space, the second write start address is determined based on the data amount and the first write start address.
[0043] In some embodiments, the second sub-processing unit executes the second task to generate second output data, and writes the second output data into the storage space based on the second write start address. At the same time, the first sub-processing unit executes the first task in parallel. In this way, the first task and the second task can be executed in parallel, and the first output data corresponding to the first task and the second output data corresponding to the second task can be written into the storage space in parallel, thereby improving the efficiency of each sub-processing unit in writing the corresponding output data into the storage space.
[0044] In the disclosed embodiment, before writing the first output data into the storage space, a second write start address is determined based on the data volume and the first write start address. The first sub-processing unit executes the first task to generate the first output data, and writes the first output data into the storage space based on the first write start address. Thus, before writing the first output data into the storage space, the second write start address can be determined in advance based on the first write start address and the data volume, thereby enabling data to be written in advance based on the second write start address, thereby improving the efficiency of writing the output data into the storage space.
[0045] In some embodiments, before the first sub-processing unit executes the first task to generate the first output data, the method may further include the following step S121: Step S121: The first sub-processing unit determines the second write start address based on the data amount and the first write start address, and sends the second write start address to the second sub-processing unit.
[0046] In some embodiments, after obtaining the first write start address and data amount, the first sub-processing unit determines a second write start address based on the first write start address and data amount, and sends the second start address to the second sub-processing unit.
[0047] In the disclosed embodiment, before the first sub-processing unit executes the first task and generates the first output data, the first sub-processing unit determines a second write start address based on the data volume and the first write start address, and sends the second write start address to the second sub-processing unit. In this way, the second sub-processing unit does not need to pay attention to the first write start address and data volume of the first sub-processing unit and can directly receive the second write start address sent by the first sub-processing unit, thereby reducing the coupling between the first sub-processing unit and the second sub-processing unit.
[0048] In some embodiments, before the first sub-processing unit executes the first task to generate the first output data, the method may further include the following steps S131 and S132: Step S131: the first sub-processing unit sends the data amount and the first write start address to the second sub-processing unit; Step S132: the second sub-processing unit determines the second write start address according to the received data volume and the first write start address.
[0049] In some embodiments, after obtaining the first write start address and data amount, the first sub-processing unit sends the first write start address and data amount to the second sub-processing unit, and the second sub-processing unit determines the second write start address based on the first write start address and data amount.
[0050] In some embodiments, before the first output data is written to the storage space, the first sub-processing unit can send the data amount and the first write start address to the data write module, and the data write module determines the second write start address based on the data amount and the first write start address. After the second sub-processing unit sends the second output data generated after the second task is executed to the data write module, the data write module writes the second output data to the storage space according to the second write start address.
[0051] Step S132: the second sub-processing unit determines the second write start address according to the received data volume and the first write start address.
[0052] In the disclosed embodiment, before the first sub-processing unit executes the first task to generate the first output data, the first sub-processing unit sends the data amount and the first write start address to the second sub-processing unit; the second sub-processing unit determines the second write start address based on the received data amount and the first write start address. In this way, after obtaining the data amount and the first write start address, the first sub-processing unit immediately sends the data amount and the first write start address to the second sub-processing unit, allowing the first sub-processing unit to continue executing the first task. This allows the second sub-processing unit to more quickly determine the second write start address based on the data amount and the first write start address, without having to wait for the first sub-processing unit to complete the first task before determining the second start address, thereby improving the efficiency of data output by each sub-processing unit.
[0053] In some embodiments, the above step S102 may include the following step S171: Step S171: In the process of the first sub-processing unit executing the first task to generate the first output data, and writing the first output data into the storage space based on the first write start address, the second sub-processing unit executes the second task to generate the second output data, and writes the second output data into the storage space based on the second write start address.
[0054] In some embodiments, after determining the second write start address, while the first sub-processing unit is executing the first task to generate the first output data, the second sub-processing unit is executing the second task in parallel to generate the second output data. Since the second write start address is determined in advance, the second output data can be written to the storage space based on the second write start address without waiting for the first sub-processing unit to complete the first task.
[0055] In the embodiment of the present disclosure, while the first sub-processing unit is executing a first task to generate first output data and writing the first output data to a storage space based on a first write start address, the second sub-processing unit is executing a second task to generate second output data and writing the second output data to the storage space based on a second write start address. Thus, while the first sub-processing unit is executing the first task, the second sub-processing unit can execute the second task in parallel and write the second output data to the storage space. This improves the efficiency of each sub-processing unit writing data to the storage space compared to the related art where each sub-processing unit serially writes the corresponding output data to the storage space.
[0056] In some embodiments, the graphics processor further includes a data writing module; the second sub-processing unit in the above step S102 executes the second task to generate second output data, and writes the second output data into the storage space based on the second write start address, and may further include the following steps S141 to S143: Step S141: the second sub-processing unit executes the second task to generate second output data; In some implementations, after determining the second write start address, the second sub-processing unit starts executing the second task and generates second output data.
[0057] Step S142: the second sub-processing unit sends the second write start address and the second output data to the data write module; Step S143: the data writing module uses the data writing port corresponding to the second sub-processing unit to write the second output data into the storage space according to the second write start address.
[0058] In some embodiments, there are multiple data writing ports corresponding to each sub-processing unit between the data writing module and the storage space, and the output data corresponding to each sub-processing unit can be written into the storage space in parallel using the multiple data writing ports.
[0059] In some embodiments, each sub-processing unit has a corresponding data write port. The second sub-processing unit sends the second write start address and the second output data to the data write module. The data write module uses the data write port corresponding to the second sub-processing unit to write the second output data into the storage space starting from the second write start address. There is no need to wait until multiple sub-processing units are processed as in the related art. After the data write module assembles the output data corresponding to each sub-processing unit into a complete output data in sequence, the assembled output data is written to the storage space through a data write port. Therefore, the solution of the embodiment of the present disclosure saves time for writing data into the storage space.
[0060] In some embodiments, the data writing module may be as follows: Figure 8 The recovery module 841 in the related art is shown.
[0061] In the disclosed embodiment, the second sub-processing unit performs a second task and generates second output data. The second sub-processing unit sends the second write start address and the second output data to the data write module. The data write module uses the data write port corresponding to the second sub-processing unit to write the second output data into the storage space according to the second write start address. Thus, data write ports corresponding to multiple sub-processing units exist between the data write module and the storage space, enabling the output data of multiple sub-processing units to be written into the storage space in parallel, thereby improving overall data write efficiency.
[0062] In some embodiments, the target processing stage includes a vertex shader stage; the above method may further include the following steps S151 and S152: Step S151: In the vertex shader stage, the first sub-processing unit receives a first to-be-executed task of the first data segment, where the first to-be-executed task has a first task attribute; Here, the first task attribute represents whether the first task to be executed is the last task in the vertex shader stage.
[0063] In some implementations, when tasks to be executed are assigned to each sub-processing unit, each task to be executed has a task attribute.
[0064] In some implementations, the first task to be executed has a first task attribute, and it can be identified based on the first task attribute whether the first task to be executed is the last task in the vertex shader stage.
[0065] In some implementations, the first task attribute may be a task identifier, or may be a result of counting tasks.
[0066] In some implementations, whether the first task to be executed is the first task may be determined directly based on the first task attribute.
[0067] Step S152: when the first task attribute indicates that the first to-be-executed task is the last task of the first data segment in the vertex shader stage, the first sub-processing unit determines that the first to-be-executed task is the first task.
[0068] In some embodiments, when the first task attribute is a task identifier, and when the task identifier indicates that the first task to be executed is the last task of the vertex shader stage, the first task to be executed is determined to be the first task. In one example, the first sub-processing unit has a total of five tasks to be executed, and a task identifier is set for each task to be executed in the order in which the tasks are to be executed. For example, the task identifiers of the five tasks to be executed are 1, 2, 3, 4, and 5, respectively. When the first sub-processing unit determines that the task identifier of the first task to be executed is 5, it determines that the first task to be executed is the first task.
[0069] In some embodiments, when the first task attribute is the result of counting tasks, the first to-be-executed task is determined to be the first task if the task count reaches the total number of tasks when the first to-be-executed task is received. In one example, it is known that the total number of tasks executed by the first sub-processing unit is 10. The first sub-processing unit increments the count by 1 each time it receives a task. When the first sub-processing unit receives the first to-be-executed task and the count result after incrementing by 1 is 10, the first to-be-executed task is determined to be the first task.
[0070] In the disclosed embodiment, during the vertex shader stage, a first sub-processing unit receives a first task to be executed for a first data segment, where the first task to be executed has a first task attribute. The first sub-processing unit determines that the first task to be executed is the first task if the first task attribute indicates that the first task to be executed is the last task in the vertex shader stage for the first data segment. In this way, by determining whether the first task to be executed is the first task based on the first task attribute, the first write start address and data size corresponding to the first task can be obtained in advance.
[0071] In some embodiments, the target processing stage includes a geometry shader stage; the above method may further include the following steps S161 and S162: Step S161: In the geometry shader stage, the first sub-processing unit receives at least one second to-be-executed task of the second data segment, and stores the second to-be-executed task in a to-be-executed task queue; In some embodiments, in the geometry shader stage, since the amount of output data of each task to be executed by each sub-processing unit is uncertain when performing the geometry shader stage, it is necessary to first collect all tasks to be processed by the first sub-processing unit and cache each task in a queue, and utilize the first-in-first-out feature of the queue to determine the last task of the first sub-processing unit in the geometry shader stage.
[0072] Step S162: the first sub-processing unit sequentially reads the second to-be-executed tasks from the to-be-executed task queue, and when the to-be-executed task queue is empty after reading the current second to-be-executed task, determines the current second to-be-executed task as the first task.
[0073] In some implementations, after the second task to be executed is read from the task queue to be executed, if there is no task in the task queue to be executed, the second task to be executed is determined to be the first task in the geometry shader stage.
[0074] In the disclosed embodiment, during the geometry shader stage, the first sub-processing unit receives at least one second task to be executed for the second data segment and stores the second task to be executed in a queue of tasks to be executed. The first sub-processing unit sequentially reads the second tasks to be executed from the queue of tasks to be executed, and if the queue of tasks to be executed is empty after the current second task to be executed is read, the current second task to be executed is determined as the first task. In this way, by utilizing the first-in, first-out feature of the queue, the last task in the queue, i.e., the first task, is determined by sequentially reading the second tasks to be executed. The first write start address and data size corresponding to the first task can be obtained in advance, allowing the second sub-processing unit to obtain the second write start address in advance, thereby improving the efficiency of data writing.
[0075] The present disclosure provides a graphics processor. Figure 2 A schematic diagram of the structure of a graphics processor provided in an embodiment of the present disclosure Figure One ,like Figure 2 As shown, the graphics processor 200 includes at least two sub-processing units, each of which is used to process a corresponding data segment in the geometric data to be processed, wherein: The first sub-processing unit 201 is configured to, in response to receiving a first task, obtain a data volume and a first write start address corresponding to first output data to be output by the first task; the first task is the last task of the first data segment corresponding to the first sub-processing unit 201 in the target processing stage; The second sub-processing unit 202 is used to execute a second task to generate second output data, and write the second output data into the storage space based on a second write start address; the second write start address is determined based on the data amount and the first write start address, the second task is the first task of the second data segment corresponding to the second sub-processing unit 202 in the target processing stage, and the second data segment is the next data segment of the first data segment.
[0076] In some implementations, the first sub-processing unit 201 determines the second write start address according to the first write start address and the first output data obtained by executing the first task, and sends the second write start address to the second sub-processing unit 202 .
[0077] In some embodiments, the first sub-processing unit 201 sends the first write start address and data amount obtained by executing the first task to the second sub-processing unit, and the second sub-processing unit 202 determines the second write start address based on the first write start address and the first output data.
[0078] In some embodiments, while the second sub-processing unit 202 starts executing the second task and writes the generated second output data into the storage space according to the second write start address, the first sub-processing unit 201 processes the first task and writes the generated first output data into the storage space.
[0079] In some embodiments, the first sub-processing unit 201 sends the first output data and the first write start address to the data writing module, and the data writing module writes the first output data into the storage space using the data writing port corresponding to the first sub-processing unit 201 .
[0080] In some embodiments, the second sub-processing unit 202 sends the second output data and the second write start address to the data writing module, and the data writing module writes the second output data into the storage space using the data writing port corresponding to the second sub-processing unit 202 .
[0081] In some embodiments, the target processing stage may include a vertex shader stage and a geometry shader stage, and the first sub-processing unit 201 and the second sub-processing unit 202 perform the vertex shader stage and the geometry shader stage respectively.
[0082] In some embodiments, the first task is the last task in the target processing stage, that is, the last task in the vertex shader stage or the geometry shader stage. Due to the differences in the processing processes of the vertex shader stage and the geometry shader stage, the method of determining the last task is also different.
[0083] In some embodiments, at the vertex shader stage, whether the first task to be executed is the last task is determined based on the task attribute of each task to be executed. That is, if the first task attribute indicates that the first task to be executed is the last task for the first data segment at the vertex shader stage, the first sub-processing unit 201 determines that the first task to be executed is the first task.
[0084] In some embodiments, in the vertex shader stage, the first sub-processing unit 201 stores a task to be executed into a task queue to be executed, and in a case that the task queue to be executed is empty after reading a current second task to be executed, the current second task to be executed is determined as the first task.
[0085] In the embodiments of the present disclosure, the first task is a last task of a first data segment corresponding to the first sub-processing unit in the target processing stage, the first sub-processing unit executes the first task, and obtains a data amount of first output data to be output by the first task and a first write start address; the second write start address is determined based on the data amount and the first write start address, the second task is a first task of a second data segment corresponding to the second sub-processing unit in the target processing stage, the second data segment is a next data segment of the first data segment, the second sub-processing unit executes the second task to generate second output data, and writes the second output data into the storage space based on the second write start address. In this way, after the first sub-processing unit obtains the data amount and the first write start address, the second write start address can be determined in advance according to the data amount and the first write start address, and the second sub-processing unit can start executing the second task based on the second write start address, without waiting for the first sub-processing unit to finish executing the first task, thereby reducing the time for waiting for the first sub-processing unit to finish executing the first task, and improving the efficiency of writing the corresponding output data into the storage space by each sub-processing unit.
[0086] In some embodiments, the second write start address is determined based on the data amount and the first write start address before the first output data is written into the storage space. The first sub-processing unit 201 is further configured to execute the first task to generate the first output data, and write the first output data into the storage space based on the first write start address.
[0087] In the embodiments of the present disclosure, the second write start address is determined according to the data amount and the first write start address before the first output data is written into the storage space, the first sub-processing unit executes the first task to generate the first output data, and writes the first output data into the storage space based on the first write start address. In this way, the second write start address can be determined in advance according to the first write start address and the data amount, so that data writing can be performed in advance based on the second write start address, and the efficiency of writing the output data into the storage space is improved.
[0088] In some embodiments, the first sub-processing unit 201 is further configured to, before the first sub-processing unit 201 executes the first task to generate the first output data, determine the second write start address based on the data amount and the first write start address, and send the second write start address to the second sub-processing unit 202.
[0089] In some embodiments, the second data segment executed by the second sub-processing unit is a next data segment of the first data segment executed by the first sub-processing unit.
[0090] In the disclosed embodiment, before the first sub-processing unit executes the first task and generates the first output data, the first sub-processing unit determines a second write start address based on the data volume and the first write start address, and sends the second write start address to the second sub-processing unit. In this way, the second sub-processing unit does not need to pay attention to the first write start address and data volume of the first sub-processing unit and can directly receive the second write start address sent by the first sub-processing unit, thereby reducing the coupling between the first sub-processing unit and the second sub-processing unit.
[0091] In some embodiments, the first sub-processing unit is further configured to send the data amount and the first write start address to the second sub-processing unit before the first sub-processing unit executes the first task to generate the first output data; The second sub-processing unit is further configured to determine the second write start address according to the received data volume and the first write start address.
[0092] In the disclosed embodiment, before the first sub-processing unit executes the first task and generates the first output data, the first sub-processing unit sends the data amount and the first write start address to the second sub-processing unit; the second sub-processing unit then determines the second write start address based on the received data amount and the first write start address. In this way, after receiving the data amount and the first write start address, the first sub-processing unit immediately sends them to the second sub-processing unit, allowing the first sub-processing unit to continue executing the first task while the second sub-processing unit can more quickly determine the second write start address and begin executing the second task, thereby improving the efficiency of both the first and second sub-processing units in executing tasks.
[0093] In some embodiments, the second sub-processing unit is further used to: in the process of the first sub-processing unit executing the first task to generate the first output data, and writing the first output data to the storage space based on the first write start address, execute the second task to generate the second output data, and write the second output data to the storage space based on the second write start address.
[0094] In the embodiments of the present disclosure, in a process in which the first sub-processing unit executes the first task to generate first output data and writes the first output data into the storage space based on a first write start address, the second sub-processing unit executes the second task to generate second output data and writes the second output data into the storage space based on a second write start address. In this way, when the first sub-processing unit executes the first task, the second sub-processing unit can execute the second task in parallel and write the second output data into the storage space, which improves the efficiency of writing data into the storage space by each sub-processing unit compared with writing corresponding output data into the storage space by each sub-processing unit in a serial manner in the related art.
[0095] In some embodiments, as shown in FIG. 2, the graphics processor 200 further includes a data writing module 203, wherein: Figure 3 The second sub-processing unit 202 is further configured to execute the second task to generate second output data. The second sub-processing unit 202 is further configured to send the second write start address and the second output data to the data writing module. The data writing module 203 is configured to write the second output data into the storage space according to the second write start address by using a data writing port corresponding to the second sub-processing unit 202.
[0096] In the embodiments of the present disclosure, the second sub-processing unit executes the second task to generate second output data, the second sub-processing unit sends the second write start address and the second output data to the data writing module, and the data writing module writes the second output data into the storage space according to the second write start address by using a data writing port corresponding to the second sub-processing unit. In this way, there are data writing ports corresponding to multiple sub-processing units between the data writing module and the storage space, and the output data of the multiple sub-processing units can be written into the storage space in parallel, which improves the efficiency of overall data writing.
[0097] In some embodiments, the target processing stage includes a vertex shader stage. In the vertex shader stage, the first sub-processing unit is further configured to receive a first to-be-executed task of the first data segment, the first to-be-executed task having a first task attribute. The first sub-processing unit is further configured to determine that the first to-be-executed task is the first task in a case where the first task attribute indicates that the first to-be-executed task is the last task of the first data segment in the vertex shader stage.
[0098] In the disclosed embodiment, during the vertex shader stage, a first sub-processing unit receives a first task to be executed for a first data segment, where the first task to be executed has a first task attribute. The first sub-processing unit determines that the first task to be executed is the first task if the first task attribute indicates that the first task to be executed is the last task in the vertex shader stage for the first data segment. In this way, by determining whether the first task to be executed is the first task based on the first task attribute, the first write start address and data size corresponding to the first task can be obtained in advance.
[0099] In some embodiments, the target processing stage includes a geometry shader stage; In the geometry shader stage, the first sub-processing unit is further configured to receive at least one second to-be-executed task of the second data segment, and store the second to-be-executed task in a to-be-executed task queue; The first sub-processing unit is further configured to sequentially read second tasks to be executed from the task queue to be executed, and when the task queue to be executed is empty after the current second task to be executed is read, determine the current second task to be executed as the first task.
[0100] In the disclosed embodiment, during the geometry shader stage, the first sub-processing unit receives at least one second task to be executed for the second data segment and stores the second task to be executed in a queue of tasks to be executed. The first sub-processing unit sequentially reads the second tasks to be executed from the queue of tasks to be executed, and if the queue of tasks to be executed is empty after the current second task to be executed is read, the current second task to be executed is determined as the first task. In this way, by utilizing the first-in, first-out feature of the queue, the last task in the queue, i.e., the first task, is determined by sequentially reading the second tasks to be executed. The first write start address and data size corresponding to the first task can be obtained in advance, allowing the second sub-processing unit to obtain the second write start address in advance, thereby improving the efficiency of data writing.
[0101] An embodiment of the present disclosure provides a computer device including the aforementioned graphics processor.
[0102] The following describes the application of the embodiments of the present disclosure in actual scenarios.
[0103] like Figure 4As shown, the processing flow of the GPU may include an input assembler step 401, a vertex shader stage 402, a geometry shader stage 403, a stream output stage 404, a rasterization stage 405, a pixel shader stage 406, and an output merge stage 407. The main purpose of the stream output stage 404 is to continuously output inactive vertex data during the geometry shader stage 403 or the vertex shader stage 402 to one or more buffers in a storage space 408. After the vertex data is transferred to the storage space 408 through the stream output stage 404, it can be read back into the GPU's graphics processing pipeline for further processing during subsequent rendering, or the data can be copied to the storage space 408 for easy reading by the central processing unit (CPU).
[0104] In the related art, when the GPU is in the stream output stage 404, only one port is connected to the storage space 408 at a time. Figure 5 As shown, data will be continuously written to the starting address 501 allocated to the storage space 408 until the data writing is completed, and the end address 502 of the data writing is obtained. This data output method is inefficient and single.
[0105] In the GPU geometry pipeline, such as Figure 6 As shown, the input data can be processed in segments and in parallel, that is, multiple sub-process units (SPUs) process the input data in segments in parallel, that is, the total data to be processed is divided into data processed by spu0, data processed by spu1, data processed by spu2 and data processed by spu3, and the data processed by spu0, data processed by spu1, data processed by spu2 and data processed by spu3 can be processed in parallel.
[0106] like Figure 7 As shown, after each sub-processing unit completes data processing, it sends the output data to the data recovery module, and the data recovery module assembles the output data of each sub-processing unit in sequence and then outputs it to the storage space.
[0107] It can be seen from the above implementation that if Figure 8As shown, spu0 completes the execution of the first vertex shader stage 801 and the first geometry shader stage 802, and enters the first stream output stage 803, which sends the data to be output to the recovery module 841; spu1 completes the execution of the second vertex shader stage 811 and the second geometry shader stage 812, and enters the second stream output stage 813, which sends the data to be output to the recovery module 841; spu2 completes the execution of the third vertex shader stage 821 and the third geometry shader stage 822, and enters the third stream output stage 8 23, the third stream output stage 823 sends the data to be output to the recovery module 841; after SPU3 executes the fourth vertex shader stage 831 and the fourth geometry shader stage 832, it enters the fourth stream output stage 833, which sends the data to be output to the recovery module 841; the recovery module 841 assembles the output data of the first stream output stage 803, the second stream output stage 813, the third stream output stage 823, and the fourth stream output stage 833 in sequence, and writes the assembled data to the storage space 408 using a data write port. Therefore, even if the data is processed in parallel, it must still be written serially to the storage space in sequence. This kind of parallelism is "false" parallelism and does not save time.
[0108] Based on the above description, an embodiment of the present disclosure provides a method for improving stream output performance. This method changes a purely serial data output method to a partially parallel data output method, thereby improving the efficiency of the GPU in the stream output stage.
[0109] The specific implementation of this method is as follows: In the geometry processing stage of the GPU, the input data can be processed in parallel in segments, that is, multiple sub-processing units process a segment of data in parallel.
[0110] When the vertex shader stage, such as Figure 9 As shown, since the total number of vertices processed by each task is predictable, even if the total number of tasks to be processed by the sub-processing unit is unknown, it can be known that this task is the last task when the last task is executed. At this time, the subsequent recovery module is notified in advance, the total number of vertices of this task is obtained in advance, and the end address of the data write of the current sub-processing unit is determined. The end address of the data of the last task of the current sub-processing unit can be written into the storage space as the starting address of the data write of the next sub-processing unit. The next sub-processing unit can process the data first, and the current sub-processing unit then processes the data of the last task in the vertex shader in parallel. From Figure 9 It can be seen from the comparison between the data writing of each sub-processing unit of the solution provided by the present disclosure and the solution in the related art that the time saved when writing data is found.
[0111] When the geometry shader stage, since the geometry shader will be uncertain output vertex number, need to collect all the tasks first, in the execution of the last task, according to its attribute information contained in advance, calculate the number of vertices, so as to inform the next sub processing unit, flow output processing.
[0112] Based on the above two optimization information, as shown in Figure 10 The storage space 408 is optimized from a single port to multiple ports for data parallel writing. That is, there is a data writing port corresponding to each sub-processing unit between the recovery module 841 and the storage space 408, and the output data of each sub-processing unit is written into the storage space 408 by using the data writing port corresponding to each sub-processing unit.
[0113] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the disclosure. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the disclosure, the size of the sequence number of each step / process does not mean the order of execution, and the execution order of each step / process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the disclosure. The sequence number of the above embodiments of the disclosure is only for description, not representing the advantages and disadvantages of the embodiments.
[0114] It should be noted that in this document, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0115] In the several embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0116] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0117] In addition, the functional units in the various embodiments of the present disclosure may all be integrated into one processing unit, or each unit may be independently used as a unit, or two or more units may be integrated into one unit; the aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units. A person skilled in the art will understand that all or part of the steps of the aforementioned method embodiments may be completed by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the aforementioned method embodiments; and the aforementioned storage medium includes various media that can store program code, such as a mobile storage device, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0118] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0119] The above is only an embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or replacements that can be easily conceived by any technician familiar with this technical field within the technical scope disclosed in this disclosure should be covered by the protection scope of the present disclosure.
Claims
1. A geometric processing method, characterized in that: Applied to a graphics processor, the graphics processor includes at least two sub-processing units, each of the sub-processing units is used to process a corresponding data segment in the geometric data to be processed, the method includes: In response to receiving a first task, the first sub-processing unit obtains a data volume and a first write start address corresponding to first output data to be output by the first task; the first task is the last task of the first data segment corresponding to the first sub-processing unit in the target processing stage; The second sub-processing unit executes a second task to generate second output data, and writes the second output data into a storage space based on a second write start address; the second write start address is determined based on the data volume and the first write start address, the second task is the first task of the second data segment corresponding to the second sub-processing unit in the target processing stage, and the second data segment is the next data segment of the first data segment.
2. The geometric processing method according to claim 1, characterized in that: The second write start address is determined based on the data amount and the first write start address before the first output data is written into the storage space; The method further comprises: The first sub-processing unit executes the first task to generate the first output data, and writes the first output data into a storage space based on the first write start address.
3. The geometric processing method according to claim 2, characterized in that: Before the first sub-processing unit executes the first task to generate the first output data, the method further includes: The first sub-processing unit determines the second write start address based on the data amount and the first write start address, and sends the second write start address to the second sub-processing unit.
4. The geometric processing method according to claim 2, characterized in that: Before the first sub-processing unit executes the first task to generate the first output data, the method further includes: The first sub-processing unit sends the data amount and the first write start address to the second sub-processing unit; The second sub-processing unit determines the second write start address according to the received data amount and the first write start address.
5. The geometric processing method according to claim 2, characterized in that: The second sub-processing unit executes the second task to generate second output data, and writes the second output data into the storage space based on a second write start address, including: In the process of the first sub-processing unit executing the first task to generate the first output data and writing the first output data into the storage space based on the first write start address, the second sub-processing unit executes the second task to generate the second output data and writes the second output data into the storage space based on the second write start address.
6. The geometric processing method according to any one of claims 1 to 5, characterized in that: The graphics processor further includes a data writing module; and the method further includes: The second sub-processing unit executes the second task to generate second output data, and writes the second output data into the storage space based on a second write start address, including: The second sub-processing unit performs the second task and generates second output data; The second sub-processing unit sends the second write start address and the second output data to the data writing module; The data writing module uses the data writing port corresponding to the second sub-processing unit to write the second output data into the storage space according to the second writing start address.
7. The geometric processing method according to any one of claims 1 to 5, characterized in that: The target processing stage includes a vertex shader stage; the method further includes: In the vertex shader stage, the first sub-processing unit receives a first to-be-executed task of the first data segment, where the first to-be-executed task has a first task attribute; When the first task attribute indicates that the first to-be-executed task is the last task of the first data segment in the vertex shader stage, the first sub-processing unit determines that the first to-be-executed task is the first task.
8. The geometric processing method according to any one of claims 1 to 5, characterized in that: The target processing stage includes a geometry shader stage; the method further includes: In the geometry shader stage, the first sub-processing unit receives at least one second to-be-executed task of the second data segment, and stores the second to-be-executed task in a to-be-executed task queue; The first sub-processing unit sequentially reads second tasks to be executed from the task queue to be executed, and when the task queue to be executed is empty after the current second task to be executed is read, determines the current second task to be executed as the first task.
9. A graphics processor, characterized in that: The graphics processor includes at least two sub-processing units, each of which is used to process a corresponding data segment in the geometric data to be processed, wherein: The first sub-processing unit is configured to, in response to receiving a first task, obtain a data volume and a first write start address corresponding to first output data to be output by the first task; the first task being the last task of the first data segment corresponding to the first sub-processing unit in the target processing stage; The second sub-processing unit is used to execute a second task to generate second output data, and write the second output data into a storage space based on a second write start address; the second write start address is determined based on the data amount and the first write start address, the second task is the first task of the second data segment corresponding to the second sub-processing unit in the target processing stage, and the second data segment is the next data segment of the first data segment.
10. The graphics processor according to claim 9, wherein: The second write start address is determined based on the data amount and the first write start address before the first output data is written into the storage space; The first sub-processing unit is further configured to: execute the first task to generate the first output data, and write the first output data into a storage space based on the first write start address.
11. The graphics processor according to claim 10, wherein: The first sub-processing unit is further configured to: determine the second write start address based on the data volume and the first write start address, and send the second write start address to the second sub-processing unit before the first sub-processing unit executes the first task to generate the first output data.
12. The graphics processor according to claim 10, wherein: The first sub-processing unit is further configured to send the data amount and the first write start address to the second sub-processing unit before the first sub-processing unit executes the first task to generate the first output data; The second sub-processing unit is further configured to determine the second write start address according to the received data volume and the first write start address.
13. The graphics processor according to claim 10, wherein: The second sub-processing unit is also used to: in the process of the first sub-processing unit executing the first task to generate the first output data, and writing the first output data into the storage space based on the first write start address, execute the second task to generate the second output data, and write the second output data into the storage space based on the second write start address.
14. The graphics processor according to any one of claims 9 to 13, characterized in that: The graphics processor further includes a data writing module, wherein: The second sub-processing unit is further configured to execute the second task and generate second output data; The second sub-processing unit is further configured to send the second write start address and the second output data to the data write module; The data writing module is configured to write the second output data into the storage space according to the second write start address by using the data writing port corresponding to the second sub-processing unit.
15. The graphics processor according to any one of claims 9 to 13, characterized in that: The target processing stage includes a vertex shader stage; In the vertex shader stage, the first sub-processing unit is further configured to receive a first to-be-executed task of the first data segment, where the first to-be-executed task has a first task attribute; The first sub-processing unit is further configured to determine that the first to-be-executed task is the first task when the first task attribute indicates that the first to-be-executed task is the last task of the first data segment in the vertex shader stage.
16. The graphics processor according to any one of claims 9 to 13, characterized in that: The target processing stage includes a geometry shader stage; In the geometry shader stage, the first sub-processing unit is further configured to receive at least one second to-be-executed task of the second data segment, and store the second to-be-executed task in a to-be-executed task queue; The first sub-processing unit is further configured to sequentially read second tasks to be executed from the task queue to be executed, and when the task queue to be executed is empty after the current second task to be executed is read, determine the current second task to be executed as the first task.
17. A computer device, characterized in that: Comprising a graphics processor as claimed in any one of claims 9 to 16.
Citation Information
Patent Citations
Writing method and device of programs
CN106802811A
Geometry processing method and device, equipment and storage medium
CN117853309A
Data storage method and device, electronic equipment and readable storage medium
CN118864225A
Data processing method and related equipment
CN119988005A
Graphics processing
US20210295584A1