Geometric processing method, graphic processing chip and computer equipment

By broadcasting the write end address and target processor identifier in the graphics processing chip, the problem of complex and time-consuming data writing in multi-core mode is solved, and an efficient data writing process is achieved.

CN120807267AActive Publication Date: 2025-10-17MOORE THREADS TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511284697.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-17
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

In graphics processing chips, multiple processor cores need to switch from multi-core mode to single-core mode when writing to storage space, making the data writing process complex and time-consuming.

Method used

By broadcasting the write end address and target processor identifier of the processor core, each processor core can directly write to the storage space in multi-core mode, avoiding mode switching and ensuring data continuity by using the write end address.

Benefits of technology

The data writing process of multiple processor cores is simplified, time consumption is reduced, and data writing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807267A_ABST
    Figure CN120807267A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a geometric processing method, a graphic processing chip and computer equipment, the method is applied to the graphic processing chip, and each processor core in the graphic processing chip is used for processing a corresponding data segment in geometric data to be processed. The method comprises the steps that a first processor core obtains a first write end address of first output data in a storage space; the first processor core broadcasts a first write end address and a first target processor identifier corresponding to a second data segment, and the second data segment is the next data segment of the first data segment; and the second processor core receives the first write end address and the first target processor identifier, and writes the second output data into the storage space based on the first write end address under the condition that the first target processor identifier is matched with the processor identifier of the second processor core. According to the embodiment of the invention, the data writing process of the multiprocessor core is simplified in a broadcasting manner, and the time consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to, but is not limited to, the technical field of graphics processing, and particularly relates to a geometry processing method, a graphics processing chip and a computer device. BACKGROUND

[0002] In a graphics processing chip, each processor core continuously outputs data processed by a geometry shader stage or a vertex shader stage to one or more buffers in a storage space in a stream out stage. In the related art, when multiple processor cores write corresponding output data to the storage space, data writing needs to be performed in a single-core mode switched from a multi-core mode, which leads to a complex process and a long time consumption when each processor core writes data. SUMMARY

[0003] Therefore, the embodiments of the present disclosure provide at least a geometry processing method, a graphics processing chip and a computer device.

[0004] The technical scheme of the embodiments of the present disclosure is implemented as follows: The embodiments of the present disclosure provide a geometry processing method applied to a graphics processing chip, wherein the graphics processing chip comprises a plurality of processor cores, each of which is configured to process a corresponding data segment in to-be-processed geometry data, and the method comprises the following steps: A first processor core acquires a first write end address of first output data in a storage space, wherein the first output data is obtained by performing geometry processing on a corresponding first data segment by the first processor core and is written into the storage space by the first processor core; The first processor core broadcasts the first write end address and a first target processor identifier corresponding to a second data segment, wherein the second data segment is a next data segment of the first data segment; A second processor core receives the first write end address and the first target processor identifier, and in a case where the first target processor identifier matches a processor identifier of the second processor core, writes second output data into the storage space based on the first write end address, wherein the second output data is obtained by performing geometry processing on the corresponding second data segment by the second processor core.

[0005] The embodiments of the present disclosure provide a graphics processing chip, comprising a plurality of processor cores, each of which is configured to process a corresponding data segment in to-be-processed geometry data, and wherein: The first processor core is configured to: obtain a first write end address of the first output data in the storage space; the first output data is obtained by performing geometric processing on a corresponding first data segment by the first processor core and is written into the storage space by the first processor core; The first processor core is further configured to: broadcast the first write end address and a first target processor identifier corresponding to the second data segment, the second data segment being a next data segment of the first data segment; The second processor core is configured to: receive the first write end address and the first target processor identifier, and in a case where the first target processor identifier matches a processor identifier of the second processor core, write second output data into the storage space based on the first write end address; the second output data is obtained by performing geometric processing on the corresponding second data segment by the second processor core.

[0006] The present disclosure provides a computer device, comprising the above-mentioned graphics processing chip.

[0007] In the present disclosure, the first processor core performs geometric processing on the first data segment to obtain the first output data, the first processor core writes the first output data into the storage space, and obtains the first write end address of the first output data in the storage space; the second data segment is a next data segment of the first data segment, the first processor core broadcasts the first write end address and the first target processor identifier corresponding to the second data segment; the second processor core receives the first write end address and the first target processor identifier, and in a case where the first target processor identifier matches a processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment to obtain the second output data, and writes the second output data into the storage space according to the first write end address. In this way, when writing the output data of multiple processor cores into the storage space, the output data of each processor core can be written into the storage space by broadcasting the write end address corresponding to the current processor core and the processor identifier corresponding to the processor core that performs the next data writing, without switching from the multi-core mode to the single-core mode, thereby simplifying the data writing process of the multiple processor cores and reducing the time consumption.

[0008] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the technical solutions of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0009] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the technical solutions of the present disclosure.

[0010] Figure 1An implementation flowchart of a geometry processing method provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 2 An assembly structure diagram of a graphics processing chip provided by an embodiment of the present disclosure is shown in FIG. 3. Figure 3 An assembly structure diagram of a processor core provided by an embodiment of the present disclosure is shown in FIG. 4. Figure 4 An assembly structure diagram of a computer device provided by an embodiment of the present disclosure is shown in FIG. 5. Figure 5 An implementation flowchart of a geometry processing method in the related art is shown in FIG. 1. Figure 6 A data output diagram of a multi-processor core in the related art is shown in FIG. 2. Figure 7 A diagram of data writing into a storage space in the related art is shown in FIG. 3. Figure 8 An implementation flowchart of a stream output one-by-one output scheme between multi-cores provided by an embodiment of the present disclosure is shown in FIG. 4. Figure 9 A specific flowchart of stream output synchronization provided by an embodiment of the present disclosure is shown in FIG. 5. Figure 10 A synchronization processing diagram between multi-processor cores provided by an embodiment of the present disclosure is shown in FIG. 6. DETAILED DESCRIPTION

[0011] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the technical solutions of the present disclosure are further described in detail below with reference to the drawings and embodiments, and the described embodiments should not be regarded as limiting the present disclosure, and all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present disclosure.

[0012] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0013] The terms "first / second / third" involved only distinguish similar objects, and do not represent a specific order of the objects, and it can be understood that "first / second / third" can interchange specific order or sequence as allowed, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure.

[0015] Embodiments of the present disclosure provide a geometry processing method, applied to a graphics processing chip, the graphics processing chip comprising a plurality of processor cores, each of the processor cores being configured to process a corresponding data segment in to-be-processed geometry data, Figure 1 An implementation flowchart of the geometry processing method provided by the embodiments of the present disclosure is shown in FIG. 1, which comprises the following steps S101-S103: Figure 1 Step S101: A first processor core acquires a first write end address of first output data in a storage space; the first output data is obtained by the first processor core performing geometry processing on a corresponding first data segment and is written into the storage space by the first processor core. Here, the first processor core is a processor core in the graphics processing chip that processes the first data segment in the geometry data.

[0016] The first write end address is an end address of the first processor core writing the first output data into the storage space.

[0017] In some embodiments, the to-be-processed geometry data is divided into a plurality of data segments, each processor core processes a corresponding data segment, for example, the first processor core processes the first data segment and the second processor core processes the second data segment.

[0018] In some embodiments, the to-be-processed geometry data can include, but is not limited to, at least one of vertex data, primitive data, texture data, lighting data, etc.

[0019] In some embodiments, the geometry processing performed by the first processor core on the first data segment can include, but is not limited to, at least one of an input assembler step, a vertex shader stage, a geometry shader stage, a stream output stage, a rasterization stage, a pixel shader stage, and an output merger stage, etc.

[0020] In some embodiments, the data amount and the write start address corresponding to the first output sub-data can be acquired when the first processor core receives the last task of the first data segment; the first write end address is determined by the first processor core according to the data amount and the write start address corresponding to the first output sub-data.

[0021] In some embodiments, the first processor core writes the first output data into the storage space in the stream output stage.

[0022] ​In some embodiments, the stream output stage is to continuously output the vertex data outputted from the vertex shader stage and the geometry shader stage into a storage space. Here, the continuously outputting into the storage space means that the output data of each processor core is continuous in the storage space after the corresponding output data is written into the storage space by the processor core.

[0023] In some embodiments, the storage space can be a global memory, which mainly refers to a dynamic random-access memory (DRAM) on a GPU motherboard.

[0024] Step S102: The first processor core broadcasts the first write end address and a first target processor identifier corresponding to a second data segment, the second data segment being a next data segment of the first data segment. Here, the first target processor identifier is a processor identifier of a processor core used for processing the second data segment, and the second data segment is a next data segment of the first data segment.

[0025] In some embodiments, each data segment is assigned a corresponding processor core in advance. After obtaining the first write end address of the first output data in the storage space, the first processor core determines the processor core used for processing the second data segment, and determines the processor identifier corresponding to the processor core used for processing the second data segment as the first target processor identifier. The first processor core broadcasts the first write end address and the first target processor identifier.

[0026] For example, the processor core used for processing the second data segment is a second processor core, and the first target processor identifier can be determined as the processor identifier corresponding to the second processor core.

[0027] In some embodiments, each processor core processes each data segment in the to-be-processed geometry data in order according to the respectively corresponding processor identifier. For example, the processor identifiers of the processor cores can be 1, 2 and 3, the processor core with the processor identifier of 1 processes the first data segment, the processor core with the processor identifier of 2 processes the second data segment, and the processor core with the processor identifier of 3 processes the third data segment, wherein the second data segment is a next data segment of the first data segment, and the third data segment is a next data segment of the second data segment. Therefore, the first target processor identifier is a next processor identifier of the processor identifier of the first processor core. For example, when the processor identifier of the first processor core is 1, the first target processor identifier is 2.

[0028] In some embodiments, the first processor core can include a synchronization module for broadcasting the first write end address and the first target processor identifier.

[0029] In some embodiments, the synchronization module can also be used to receive broadcast information sent by other processor cores.

[0030] Step S103: The second processor core receives the first write end address and the first target processor identifier, and in the case that the first target processor identifier matches the processor identifier of the second processor core, writes second output data into the storage space based on the first write end address; the second output data is obtained by the second processor core performing geometric processing on the corresponding second data segment.

[0031] In some embodiments, in the case that the first target processor identifier is equal to the processor identifier of the second processor core, it is determined that the first target processor identifier matches the processor identifier of the second processor core.

[0032] In some embodiments, the first write end address can be used as the write start address of the second output data in the storage space.

[0033] In some embodiments, the first write end address and the first target processor identifier are received in the case that the second processor core enters the stream output stage.

[0034] In some embodiments, in the case that the second processor core receives the first write end address and the first target processor identifier, it is possible that the geometric processing of the second data segment has been completed, and it is also possible that the geometric processing of the second data segment has not been completed; in the case that the geometric processing of the second data segment has not been completed by the first processor core, the geometric processing of the second data segment is continued, and after the geometric processing of the second data segment by the second processor core is completed, the second processor core enters the stream output stage and writes the second output data into the storage space.

[0035] In implementation, after the second processor core receives the first write end address and the first target processor identifier, the processor identifier of the second processor core is compared with the first target processor identifier; in the case that it is determined that the first target processor identifier matches the processor identifier of the second processor core, it is determined that the second processor core is the processor core processing the second data segment, and therefore, the second output data can be written into the storage space according to the first write end address; in this way, each processor core writes the corresponding output data into the storage space according to the corresponding write end address, so that the output data in the storage space is continuous.

[0036] In some embodiments, the first processor core broadcasts the first write end address and the first target processor identifier, and the second processor core can take the first write end address as a write start address of the second processor core, so that the multiple processor cores write the corresponding output data into the storage space, and the output data in the storage space is continuous, and the multiple processor cores have unique processor identifiers, so that only one processor core has a processor identifier matching the first target processor identifier, and only the second processor core can write the corresponding second output data into the storage space in the case that the processor identifier of the second processor core matches the first target processor identifier, thereby avoiding write conflicts of multiple processor cores, and thus the output data of each processor core can be written into the storage space without switching the multi-core mode to the single-core mode.

[0037] In the embodiments of the present disclosure, the first processor core performs geometric processing on the first data segment to obtain first output data, writes the first output data into the storage space, and obtains a first write end address of the first output data in the storage space; the second data segment is a next data segment of the first data segment, and the first processor core broadcasts the first write end address and a first target processor identifier corresponding to the second data segment; the second processor core receives the first write end address and the first target processor identifier, and in the case that the first target processor identifier matches a processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment to obtain second output data, and writes the second output data into the storage space according to the first write end address. In this way, when writing the output data of multiple processor cores into the storage space, the output data of each processor core can be written into the storage space by broadcasting the write end address corresponding to the current processor core and the processor identifier corresponding to the processor core that writes the next data, the write end address is used to make the output data in the storage space continuous after each processor core writes the corresponding output data into the storage space, and the first target processor identifier is used to make only one processor core currently write data, thereby avoiding data write conflicts, so that the multi-core mode does not need to be switched to the single-core mode, thereby simplifying the data write process of the multiple processor cores and reducing time consumption.

[0038] In some embodiments, the above method further includes the following step S111: Step S111: After the first processor core determines the first write end address, the first processor core broadcasts the first write end address and the first target processor identifier; the first write end address is determined based on a first data amount corresponding to first output sub-data to be output by the first task and a first write start address after the first processor core receives the first task; and the first task is the last task of the first data segment.

[0039] Here, the first output sub-data is the first output sub-data obtained after the first processor core processes the first task.

[0040] In some embodiments, when the first processor core receives the first task, the first data amount corresponding to the first output sub-data to be output by the first task and the first write start address can be obtained in advance.

[0041] In some embodiments, the first processor core can determine the first data amount according to the total number of vertices corresponding to the first task.

[0042] In some embodiments, after the first processor core determines the first write end address, the first processor core can broadcast the first write end address and the first target processor identifier without waiting until the first processor core finishes executing the first task and writing the first output sub-data into the storage space.

[0043] In some embodiments, the first processor core can process the first task and broadcast the first write end address and the first target processor identifier in parallel.

[0044] In some embodiments, after the first processor core determines the first write end address, the first processor core executes the first task and writes the first output sub-data corresponding to the first task into the storage space.

[0045] In some embodiments, the first processor core can first execute the first task and write the first output sub-data corresponding to the first task into the storage space, and then determine the first write end address according to the first data amount and the first write start address.

[0046] In some embodiments, the first sub-processing unit in the first processor core executes the first task to generate the first output sub-data, and writes the first output sub-data into the storage space based on the first write start address.

[0047] In the embodiments of the present disclosure, after the first processor core determines the first write end address, the first processor core broadcasts the first write end address and the first target processor identifier; the first write end address is determined based on the first data amount corresponding to the first output sub-data to be output by the first task and the first write start address after the first processor core receives the first task; and the first task is the last task of the first data segment. In this way, when the first processor core receives the first task, the first processor core can obtain the first data amount and the first write start address, and determine the first write end address in advance according to the first data amount and the first write start address, so that the first processor core can broadcast the first write end address and the first target processor identifier without waiting until the first task is executed, thereby improving the efficiency of writing the corresponding output data into the storage space by the plurality of processor cores.

[0048] In some embodiments, the first processor core comprises at least two sub-processing units, each of the sub-processing units being configured to process a corresponding sub-data segment in the first data segment to be processed. The method further comprises the following step S131: Step S131: In response to receiving the first task, the first sub-processing unit acquires a first data amount and a first write start address corresponding to first output sub-data to be output by the first task; the first task is the last task of a first sub-data segment corresponding to the first sub-processing unit; and the first sub-data segment is the last data segment in the first data segment.

[0049] Here, the first sub-processing unit is the sub-processing unit processing the last sub-data segment in the first data segment.

[0050] In some embodiments, upon receiving the first task, the first sub-processing unit can acquire the first data amount and the first write start address corresponding to the first output sub-data to be output in advance, and according to the first data amount and the first write start address, the first sub-processing unit can determine the first write end address. Therefore, the first sub-processing unit can determine the first write end address without waiting for the first output sub-data to be written into the storage space.

[0051] In some embodiments, the first sub-processing unit respectively acquires the first data amount and the first write start address corresponding to the first output sub-data to be output by the first task corresponding to the vertex shader stage and the geometry shader stage.

[0052] In some embodiments, the first sub-processing unit respectively determines the first write end address corresponding to the vertex shader stage and the geometry shader stage.

[0053] In the embodiments of the present disclosure, the first processor core comprises at least two sub-processing units, each of the sub-processing units being configured to process a corresponding sub-data segment in the first data segment to be processed, the first task is the last task of a first sub-data segment corresponding to the first sub-processing unit, and the first sub-data segment is the last data segment in the first data segment. Upon receiving the first task, the first sub-processing unit acquires a first data amount and a first write start address corresponding to first output sub-data to be output by the first task in advance. In this way, the first sub-processing unit can determine the first write end address according to the first data amount and the first write start address when executing the first task, without waiting for the first sub-processing unit to complete execution of the first task, thereby reducing the dependence on whether the first sub-processing unit has completed execution of the first task, and thus the dependence between the sub-processing units in the processor core in the process of executing tasks and writing data can be reduced.

[0054] In some embodiments, the first processor core obtaining the first write end address of the first output data in the storage space according to the step S101 can include the following step S141: Step S141: After the first processor core writes the first output data into the storage space, the first write end address is obtained.

[0055] In some embodiments, the first processor core can also broadcast the first write end address and the first target processor identifier corresponding to the second data segment after the first processor core writes the first output data into the storage space and obtains the first write end address.

[0056] In the embodiments of the present disclosure, the first processor core obtains the first write end address after writing the first output data into the storage space. In this way, the first processor core can directly obtain the first write end address after writing the first output data into the storage space, and this method is simple to implement and can be directly obtained without calculation.

[0057] In some embodiments, the method further includes the following step S151: Step S151: After the second processor core obtains the second write end address of the second output data in the storage space, the second processor core broadcasts the second write end address of the second output data in the storage space and the second target processor identifier corresponding to the third data segment, and the third data segment is the next data segment of the second data segment.

[0058] Here, the second target processor identifier is the processor identifier of the processor core used to process the third data segment, and the third data segment is the next data segment of the second data segment.

[0059] In some embodiments, the processor core corresponding to each data segment is allocated in advance, and after the second processor core obtains the second write end address of the second output data in the storage space, the processor core used to process the third data segment is determined, and the processor identifier corresponding to the processor core used to process the third data segment is determined as the second target processor identifier; the second processor core broadcasts the second write end address and the second target processor identifier.

[0060] In some embodiments, each processor core processes each data segment in the geometry data to be processed in sequence according to a respective corresponding processor identifier, for example, the processor identifier of each processor core can be 1, 2, 3, the processor core with the processor identifier of 1 processes the first data segment, the processor core with the processor identifier of 2 processes the second data segment, and the processor core with the processor identifier of 3 processes the third data segment, wherein the second data segment is the next data segment of the first data segment, and the third data segment is the next data segment of the second data segment. Therefore, the second target processor identifier is the next processor identifier of the processor identifier of the second processor core. For example, in the case where the processor identifier of the second processor core is 2, the second target processor identifier is 3.

[0061] In some embodiments, when each processor core broadcasts the target processor identifier, the broadcasting is performed in the order of the processor core processing the corresponding data segment in the geometry data, for example, after determining the corresponding write end address, the first processor core broadcasts the first target processor identifier matching the processor identifier of the second processor core processing the second data segment and the corresponding write end address; after determining the corresponding write end address, the second processor core broadcasts the second target processor identifier matching the processor identifier of the third processor core processing the third data segment and the corresponding write end address.

[0062] In the case where the first target processor identifier matches the processor identifier of the second processor core, the second processor core starts processing the second data segment, and the second processor core obtains the second write end address of the second output data in the storage space; the second processor core broadcasts the second write end address and the second target processor identifier corresponding to the third data segment. In this way, the output data of each processor core in the graphics processing chip can be written into the storage space.

[0063] In the embodiment of the present disclosure, the third data segment is the next data segment of the second data segment, after the second processor core obtains the second write end address of the second output data in the storage space, the second processor core broadcasts the second write end address of the second output data in the storage space and the second target processor identifier corresponding to the third data segment. In this way, the output data of each processor core in the graphics processing chip can be written into the storage space by sending broadcast information, so that the corresponding output data of each processor core in the storage space is continuous.

[0064] In some embodiments, the method further comprises the following step S161: Step S151: In the case where the first target processor identifier does not match the processor identifier of the second processor core, the second processor core continues to wait for writing the second output data into the storage space.

[0065] In some embodiments, in the case that the first target processor identifier does not match the processor identifier of the second processor core, the second processor core continues to wait and does not output data until, in the case that a target processor identifier matching the processor identifier of the second processor core is received, the second processor core writes the corresponding second output data to the storage space.

[0066] In the embodiments of the present disclosure, in the case that the first target processor identifier does not match the processor identifier of the second processor core, the second processor core continues to wait. In this way, each processor core continues to wait without receiving a data output command, so that each processor core can receive a write end address corresponding to the output data of the processor core, and thus can write the output data to the storage space according to the corresponding write end address, so that the data in the storage space is continuous after each processor core writes the respective output data to the storage space.

[0067] In some embodiments, the method further includes the following step S121: Step S121: In the case that the first target processor identifier does not match the processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment.

[0068] In some embodiments, in the case that the first target processor identifier does not match the processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment while waiting to write the second output data to the storage space.

[0069] In some embodiments, when each processor core broadcasts the corresponding write end address and the target processor identifier, other processors can process the corresponding data segment in parallel.

[0070] In some embodiments, in the case that the first target processor identifier does not match the processor identifier of the second processor core and the second processor core has not completed the geometric processing on the second data segment, the second processor core performs geometric processing on the second data segment.

[0071] In the embodiments of the present disclosure, in the case that the first target processor identifier does not match the processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment, so that the second processor core has completed the geometric processing on the second data segment and obtains the second output data when a target processor identifier matching the processor identifier of the second processor core is received, and thus the second output data can be directly written to the storage space, improving the efficiency of writing the output data to the storage space.

[0072] In some embodiments, the method further comprises steps S171 and S172: Step S171: the second processor core sends response information to the first processor core in the case that the first target processor identifier matches the processor identifier of the second processor core. Step S172: the first processor core stops broadcasting the first write end address and the first target processor identifier in response to receiving the response information.

[0073] In the embodiments of the present disclosure, in the case that the first target processor identifier matches the processor identifier of the second processor core, the second processor core sends response information to the first processor core, and the first processor core stops broadcasting the first write end address and the first target processor identifier after receiving the response information. In this way, after receiving the response information sent by the second processor core, the first processor core stops broadcasting, thereby reducing the resource consumption of the first processor core.

[0074] The embodiments of the present disclosure provide a graphics processing chip, Figure 2 A schematic diagram of the composition structure of a graphics processing chip provided by the embodiments of the present disclosure is shown in Figure 2 As shown in the figure, the graphics processing chip 200 comprises a plurality of processor cores, each of which is configured to process a corresponding data segment in the to-be-processed geometric data, wherein: The first processor core 201 is configured to obtain a first write end address of first output data in a storage space; the first output data is obtained by performing geometric processing on a corresponding first data segment by the first processor core 201, and is written into the storage space by the first processor core 201; The first processor core 201 is further configured to broadcast the first write end address and a first target processor identifier corresponding to a second data segment, the second data segment being a next data segment of the first data segment. The second processor core 202 is configured to receive the first write end address and the first target processor identifier, and in the case that the first target processor identifier matches the processor identifier of the second processor core 202, write second output data into the storage space based on the first write end address; the second output data is obtained by performing geometric processing on the corresponding second data segment by the second processor core 202.

[0075] In some embodiments, after the plurality of processor cores in the graphics processing chip 200 enter the flow output stage, the output data corresponding to each processor core is written into the storage space.

[0076] In some embodiments, the first processor core 201 acquires, in response to receiving the first task, a first data amount and a first write start address corresponding to a first output sub-data to be output by the first task, the first task being the last task of the first data segment.

[0077] In some embodiments, the first processor core 201 determines a first write end address according to the first data amount and the first write start address.

[0078] In some embodiments, the first processor core 201 executes the first task to generate the first output sub-data and writes the first output sub-data into the storage space based on the first write start address after determining the first write end address.

[0079] In some embodiments, the first processor core 201 obtains the first write end address after writing the first output data into the storage space.

[0080] In some embodiments, in a case where the first target processor identifier matches the processor identifier of the second processor core 202, it indicates that the second processor core 202 is the processor core currently to be performed data output, and the second processor core 202 writes the second output data into the storage space based on the first write end address.

[0081] In the embodiments of the present disclosure, the first processor core performs geometric processing on the first data segment to obtain first output data, writes the first output data into the storage space, and acquires a first write end address of the first output data in the storage space; the second data segment is a next data segment of the first data segment, the first processor core broadcasts the first write end address and a first target processor identifier corresponding to the second data segment; the second processor core receives the first write end address and the first target processor identifier, and in a case where the first target processor identifier matches the processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment to obtain second output data, and writes the second output data into the storage space according to the first write end address. In this way, when writing the output data of multiple processor cores into the storage space, the output data of each processor core can be written into the storage space by broadcasting the write end address corresponding to the current processor core and the processor identifier corresponding to the processor core performing data writing next, without switching from the multi-core mode to the single-core mode, thereby simplifying the data writing process of the multiple processor cores and reducing time consumption.

[0082] In some embodiments, the first processor core is further used to: broadcast the first write end address and the first target processor identifier after the first processor core determines the first write end address; the first write end address is determined after the first processor core receives the first task based on the first data volume and the first write start address corresponding to the first output sub-data to be output by the first task; the first task is the last task of the first data segment.

[0083] In some embodiments, after receiving the first task, the first processor core can obtain in advance the first data size and first write start address corresponding to the first output sub-data to be output by the first task, and determine the first write end address based on the first write start address and the first data size. In this way, the first processor core can execute the first task and broadcast the first write end address and the first target processor identifier corresponding to the second data segment in parallel.

[0084] In some embodiments, the first processor core does not need to wait until the first output sub-data to be output by the first task is written into the storage space before determining the first data amount and the first write start address, thereby reducing the dependence on whether the first processor core has completed the first task, thereby reducing the dependence between the processor chips in the graphics processing chip during the task execution and data writing process.

[0085] In the disclosed embodiment, after the first processor core determines the first write end address, the first processor core broadcasts the first write end address and the first target processor identifier. The first write end address is determined based on the first data amount and the first write start address corresponding to the first output sub-data to be output by the first task after the first processor core receives the first task. The first task is the last task in the first data segment. In this way, upon receiving the first task, the first processor core can obtain the first data amount and the first write start address, and determine the first write end address in advance based on the first data amount and the first write start address. The first write end address and the first target processor identifier can be broadcast without waiting for the first task to be completed, thereby improving the efficiency of multiple processor cores in writing corresponding output data into the storage space.

[0086] In some embodiments, as Figure 3 As shown, the first processor core 201 includes at least two sub-processing units, each of which is used to process a corresponding sub-data segment in the first data segment to be processed, wherein: The first sub-processing unit 211 is configured to: in response to receiving the first task, acquire a first data amount and a first write start address corresponding to first output sub-data to be output by the first task; the first task is a last task of a first sub-data segment corresponding to the first sub-processing unit; and the first sub-data segment is a last data segment in the first data segment.

[0087] In some embodiments, the first processor core 201 includes at least two sub-processing units, such as the first sub-processing unit 211 and the second sub-processing unit 212, configured to process corresponding sub-data segments in the first data segment to be processed.

[0088] In some embodiments, after receiving the first task, the first sub-processing unit 211 acquires a total number of vertices corresponding to the first task, and determines the first data amount corresponding to the first output sub-data to be output by the first task according to the total number of vertices.

[0089] In some embodiments, after receiving the first task, the first sub-processing unit 211 can acquire the first data amount and the first write start address corresponding to the first output sub-data to be output by the first task in advance, and determine the first write end address according to the first write start address and the first data amount.

[0090] In the embodiments of the present disclosure, the first processor core includes at least two sub-processing units, each of which is configured to process corresponding sub-data segments in the first data segment to be processed; the first task is a last task of a first sub-data segment corresponding to the first sub-processing unit; the first sub-data segment is a last data segment in the first data segment; and the first sub-processing unit acquires the first data amount and the first write start address corresponding to the first output sub-data to be output by the first task in advance after receiving the first task. In this way, the first sub-processing unit can determine the first write end address according to the first data amount and the first write start address when executing the first task, without waiting for the first sub-processing unit to complete the execution of the first task, thereby reducing the dependence between the sub-processing units in the processor core in the process of executing tasks and writing data.

[0091] In some embodiments, the first processor core is further configured to: obtain the first write end address after the first output data is written into the storage space.

[0092] In the embodiments of the present disclosure, the first processor core obtains the first write end address after the first output data is written into the storage space. In this way, the first processor core can directly obtain the first write end address after the first output data is written into the storage space, and this method is simple to implement and can be directly obtained without calculation.

[0093] In some embodiments, the second processor core is further configured to: after the second processor core acquires the second write end address of the second output data in the storage space, broadcast the second write end address of the second output data in the storage space and a second target processor identifier corresponding to a third data segment, the third data segment being a next data segment of the second data segment.

[0094] In the embodiments of the present disclosure, the third data segment is a next data segment of the second data segment, and after the second processor core acquires the second write end address of the second output data in the storage space, the second processor core broadcasts the second write end address of the second output data in the storage space and a second target processor identifier corresponding to the third data segment. In this way, the output data of each processor core in the graphics processing chip can be sequentially written into the storage space by sending broadcast information.

[0095] In some embodiments, the second processor core is further configured to: in a case where the first target processor identifier does not match the processor identifier of the second processor core, continue to wait for writing the second output data into the storage space.

[0096] In the embodiments of the present disclosure, in a case where the first target processor identifier does not match the processor identifier of the second processor core, the second processor core continues to wait. In this way, each processor core continues to wait without receiving a data output command, so that each processor core can receive a write end address corresponding to the output data of the processor core, and thus can write the output data into the storage space according to the corresponding write end address, so that the data in the storage space is continuous after each processor core writes the respective output data into the storage space.

[0097] In some embodiments, the second processor core is further configured to, in a case where the first target processor identifier does not match the processor identifier of the second processor core, perform geometric processing on the second data segment.

[0098] In the embodiments of the present disclosure, in a case where the first target processor identifier does not match the processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment, so that when the second processor core receives a target processor identifier matching the processor identifier of the second processor core, the geometric processing on the second data segment is completed, and the second output data is obtained, so that the second output data can be directly written into the storage space, and the efficiency of writing the output data into the storage space is improved.

[0099] In some embodiments, the second processor core is further configured to: send response information to the first processor core if the first target processor identifier matches the processor identifier of the second processor core; The first processor core is further configured to: in response to receiving the response information, stop broadcasting the first write end address and the first target processor identifier.

[0100] In the disclosed embodiment, if the first target processor identifier matches the processor identifier of the second processor core, the second processor core sends a response message to the first processor core. After receiving the response message, the first processor core stops broadcasting the first write end address and the first target processor identifier. Thus, after receiving the response message from the second processor core, the first processor core stops broadcasting, thereby reducing resource consumption of the first processor core.

[0101] The present disclosure provides a computer device, such as Figure 4 As shown, the computer device 400 includes the aforementioned graphics processing chip 200 .

[0102] The following describes the application of the embodiments of the present disclosure in actual scenarios.

[0103] like Figure 5 As shown, the processing flow of a graphics processing unit (GPU) in the related art may include an input assembler step 501, a vertex shader stage 502, a geometry shader stage 503, a stream output stage 504, a rasterization stage 505, a pixel shader stage 506, and an output merge stage 507. The main purpose of the stream output stage 504 is to continuously output inactive vertex data during the geometry shader stage 503 or the vertex shader stage 502 to one or more buffers in a storage space 508. After the data is transferred to the storage space 508 through the stream output stage 504, it can be read back into the GPU's graphics processing pipeline for further processing during subsequent rendering, or the data can be copied to the storage space 508 for easy reading by the central processing unit (CPU).

[0104] When the GPU in the related art processes the stream out, if there are multiple processor cores (cores), such as Figure 6As shown, the output of each processor core needs to be serially streamed, and the processor cores can include a first core 601, a second core 602, a third core 603, and a fourth core 604. In the streaming output stage 504, the output data of the first core 601 is written into the storage space 508, and then the output data of the second core 602, the third core 603, and the fourth core 604 are sequentially written into the storage space 508. Such a processing efficiency is low. Even if the input data is parallelly distributed to multiple processor cores for processing in the graphics processing stage, the processing of the streaming output stage is still serially single-processor-core output processing.

[0105] As shown in FIG. 4, the GPU 400 includes a graphics processing stage 401, a streaming output stage 402, and a storage space 403. Figure 7 As shown, the data is continuously written from the start address 701 allocated to the storage space 403 until the data writing is completed, and then the end address 702 of the data writing is obtained.

[0106] Currently, if the GPU works in a multi-core mode, the GPU needs to switch back to a single-core mode for streaming output after the system issues a streaming output command.

[0107] Based on the above description, the embodiments of the present disclosure provide a stream out output scheme between multiple cores, which can improve the efficiency of the GPU in the streaming output stage and increase the use rate of the processor cores. The original streaming output processing needs to switch to a single-core mode, and the optimized scheme can perform streaming output processing on each processor core without mode switching.

[0108] As shown in FIG. 5, the GPU 500 includes a graphics processing stage 501, a streaming output stage 502, and a storage space 503. Figure 8As shown, the embodiment of the present disclosure uses a synchronization module 801 to broadcast the stream output information of the processor core among various processor cores. When the current processor core is polled, the processor core enables the stream output module to enter the stream output stage for data output. When the data stream output of the current processor core is completed, the synchronization module broadcasts data, mainly including the end address of the current processor core in the storage space 508, which will be used as the start address of the next processor core for writing the storage space 508. For example, the first core 601 receives the broadcast information, confirms that the first core 601 performs data output, and the first core 601 writes the output data into the storage space 508 through the first stream output stage 802. After the first core 601 finishes writing the data, the synchronization broadcast information is sent, and the broadcast information includes the end address of the data written by the first core 601 and the identifier corresponding to the second core 602. The second core 602 receives the broadcast information, confirms that the second core 602 performs data output, and the second core 602 writes the output data into the storage space 508 through the second stream output stage 803. After the second core 602 finishes writing the data, the synchronization broadcast information is sent, and the broadcast information includes the end address of the data written by the second core 602 and the identifier corresponding to the third core 603. The third core 603 receives the broadcast information, confirms that the third core 603 performs data output, and the third core 603 writes the output data into the storage space 508 through the third stream output stage 804. After the third core 603 finishes writing the data, the synchronization broadcast information is sent, and the broadcast information includes the end address of the data written by the third core 603 and the identifier corresponding to the fourth core 604. The fourth core 604 receives the broadcast information, confirms that the fourth core 604 performs data output, and the fourth core 604 writes the output data into the storage space 508 through the fourth stream output stage 805.

[0109] In the stream output in the single core mode, because only one processor core outputs data, the data in the storage space is continuous, and there is no need to record the start address.

[0110] The specific flow of the stream output synchronization of one processor core of the embodiment of the present disclosure is as shown in Figure 9 The specific flow of the stream output synchronization of one processor core of the embodiment of the present disclosure is as shown in Step S901: whether to perform the stream output flow process; if yes, go to step S902; if no, go to step S906; If there is a processor core performing stream output, the synchronization module broadcast information needs to be prevented; if there is no processor core performing stream output, it is needed to determine whether the current processor core needs to perform stream output.

[0111] Step S902: prevent the synchronization module from broadcasting information, and continue to process the data that has not been output. Step S903: whether the existing data output is completed; if yes, go to step S904; if no, go to step S902; Step S904: add 1 to the current processor core identification and broadcast the end address of the storage space; After the output data of the current processor core is written into the storage space, add 1 to the current processor core identification as the processor core identification of the next processor core for stream output, and take the end address in the storage space as the start address of the next processor core for stream output.

[0112] Step S905: whether there is other processor core feedback receiving broadcast data; if yes, end the stream output process of the current processor core; if no, go to step S904; Step S906: whether the broadcast identification is equal to the current processor core identification; if yes, go to step S907; if no, continue to wait and go to step S901; Step S907: save the end address in the storage space in the broadcast information, and start the stream output process.

[0113] Take the end address in the storage space as the start address of the current processor core for stream output.

[0114] The synchronization process among the multiple processor cores is as shown in Figure 10 After the current processor core is processed, add 1 to the processor core identification of this time and broadcast it. When the new broadcast identification is equal to the processor core identification of the current processor core, it means that this processor core is selected. For example, when the first core 601 determines that the first identification 11 is equal to the first broadcast identification 21, the first core 601 performs stream output; the first core 601 adds 1 to the first broadcast identification 21 to obtain the second broadcast identification 22, and broadcasts the second broadcast identification 22; when the second core 602 determines that the second identification 12 is equal to the second broadcast identification 22, the second core 602 performs stream output; the second core 602 adds 1 to the second broadcast identification 22 to obtain the third broadcast identification 23, and broadcasts the third broadcast identification 23; when the third core 603 determines that the third identification 13 is equal to the third broadcast identification 23, the third core 603 performs stream output; the third core 603 adds 1 to the third broadcast identification 23 to obtain the fourth broadcast identification 24, and broadcasts the fourth broadcast identification 24; when the fourth core 604 determines that the fourth identification 14 is equal to the fourth broadcast identification 24, the fourth core 604 performs stream output.

[0115] In the embodiments of the present disclosure, when the stream output is performed in the case of multiple processor cores, the mult core does not need to be switched to the single core mode for the stream output, and the stream output of each processor core can be realized by polling each processor core through the manner of sending broadcast information. In this way, in the case of multiple processor cores, the time for switching the mode during the stream output is reduced, and the efficiency of the stream output is improved.

[0116] It should be understood that every technical feature mentioned in the specification refers to a specific feature related to an embodiment, which is included in at least one embodiment of the present disclosure. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that the size of the sequence number of each step / process in various embodiments of the present disclosure does not mean the order of execution, and the execution order of each step / process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure. The sequence number of the above embodiments of the present disclosure is only for description, and does not represent the advantages and disadvantages of the embodiments.

[0117] It should be noted that in this document, the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0118] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and actual implementation can have another division manner, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0119] The units described as separate components above can or can not be physically separate, and the components shown as units can or can not be physical units; they can be located in one place or distributed on multiple network units; and part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0120] In addition, each functional unit in each embodiment of the present disclosure can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in the form of hardware or in the form of hardware plus software functional units. Those skilled in the art can understand that all or part of the steps of the above method embodiments can be completed by program instruction-related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program is executed to perform the steps of the above method embodiments; and the foregoing storage medium includes mobile storage devices, read-only memories (ROM), magnetic discs or optical discs, and various media that can store program codes.

[0121] Alternatively, the integrated units of the present disclosure, if implemented in the form of software functional modules and sold or used as independent products, can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the methods described in the embodiments of the present disclosure. The foregoing storage medium includes mobile storage devices, ROM, magnetic discs or optical discs, and various media that can store program codes.

[0122] The above is only an embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any changes or replacements within the technical range disclosed by the present disclosure can be easily thought of by those skilled in the art, and should be covered by the protection scope of the present disclosure.

Claims

1. A geometric processing method, characterized in that: Applied to a graphics processing chip, the graphics processing chip includes a plurality of processor cores, each of the processor cores is used to process a corresponding data segment in the geometric data to be processed, the method comprising: The first processor core obtains a first write end address of first output data in the storage space; the first output data is obtained by the first processor core performing geometric processing on the corresponding first data segment, and is written into the storage space by the first processor core; The first processor core broadcasts the first write end address and a first target processor identifier corresponding to a second data segment, where the second data segment is a next data segment to the first data segment; The second processor core receives the first write end address and the first target processor identifier, and writes the second output data into the storage space based on the first write end address when the first target processor identifier matches the processor identifier of the second processor core; the second output data is obtained by the second processor core performing geometric processing on the corresponding second data segment.

2. The method according to claim 1, wherein The method further comprises: After the first processor core determines the first write end address, the first processor core broadcasts the first write end address and the first target processor identifier; the first write end address is determined after the first processor core receives the first task based on the first data amount corresponding to the first output sub-data to be output by the first task and the first write start address; the first task is the last task of the first data segment.

3. The method according to claim 2, characterized in that The first processor core includes at least two sub-processing units, each of the sub-processing units is used to process a corresponding sub-data segment in the first data segment to be processed; the method further includes: In response to receiving the first task, the first sub-processing unit obtains the first data volume and the first write start address corresponding to the first output sub-data to be output by the first task; the first task is the last task of the first sub-data segment corresponding to the first sub-processing unit; the first sub-data segment is the last data segment in the first data segment.

4. The method according to claim 1, wherein The first processor core obtains a first write end address of the first output data in the storage space, including: After the first processor core writes the first output data into the storage space, it obtains the first write end address.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: After the second processor core obtains the second write end address of the second output data in the storage space, the second processor core broadcasts the second write end address of the second output data in the storage space and the second target processor identifier corresponding to the third data segment, where the third data segment is the next data segment of the second data segment.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: In a case where the first target processor identifier does not match the processor identifier of the second processor core, the second processor core continues to wait for the second output data to be written into the storage space.

7. The method according to claim 6, characterized in that The method further comprises: In a case where the first target processor identifier does not match the processor identifier of the second processor core, the second processor core performs geometric processing on the second data segment.

8. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The second processor core sends response information to the first processor core when the first target processor identifier matches the processor identifier of the second processor core; In response to receiving the response information, the first processor core stops broadcasting the first write end address and the first target processor identifier.

9. A graphics processing chip, characterized in that: The graphics processing chip includes a plurality of processor cores, each of which is used to process a corresponding data segment in the geometric data to be processed, wherein: The first processor core is configured to: obtain a first write end address of first output data in the storage space; the first output data is obtained by the first processor core performing geometric processing on the corresponding first data segment, and is written into the storage space by the first processor core; The first processor core is further configured to: broadcast the first write end address and a first target processor identifier corresponding to a second data segment, where the second data segment is a next data segment of the first data segment; The second processor core is used to: receive the first write end address and the first target processor identifier, and when the first target processor identifier matches the processor identifier of the second processor core, write the second output data into the storage space based on the first write end address; the second output data is obtained by the second processor core performing geometric processing on the corresponding second data segment.

10. The graphics processing chip according to claim 9, characterized in that: The first processor core is further used to: broadcast the first write end address and the first target processor identifier after the first processor core determines the first write end address; the first write end address is determined based on the first data volume and the first write start address corresponding to the first output sub-data to be output by the first task after the first processor core receives the first task; the first task is the last task of the first data segment.

11. The graphics processing chip according to claim 10, characterized in that: The first processor core includes at least two sub-processing units, each of the sub-processing units is used to process a corresponding sub-data segment in the first data segment to be processed, wherein: The first sub-processing unit is used to: in response to receiving the first task, obtain the first data volume and the first write start address corresponding to the first output sub-data to be output by the first task; the first task is the last task of the first sub-data segment corresponding to the first sub-processing unit; the first sub-data segment is the last data segment in the first data segment.

12. The graphics processing chip according to claim 9, characterized in that: The first processor core is further configured to obtain the first write end address after writing the first output data into the storage space.

13. The graphics processing chip according to any one of claims 9 to 12, characterized in that: The second processor core is also used to: after the second processor core writes the second output data into the storage space, broadcast a second write end address of the second output data in the storage space and a second target processor identifier corresponding to a third data segment, where the third data segment is the next data segment of the second data segment.

14. The graphics processing chip according to any one of claims 9 to 12, characterized in that: The second processor core is further configured to: when the first target processor identifier does not match the processor identifier of the second processor core, continue to wait for the second output data to be written into the storage space.

15. The graphics processing chip according to claim 14, characterized in that: The second processor core is further configured to perform geometric processing on the second data segment when the first target processor identifier does not match the processor identifier of the second processor core.

16. The graphics processing chip according to any one of claims 9 to 12, characterized in that: The second processor core is further configured to: send response information to the first processor core when the first target processor identifier matches the processor identifier of the second processor core; The first processor core is further configured to: in response to receiving the response information, stop broadcasting the first write end address and the first target processor identifier.

17. A computer device, characterized in that: The method comprises the graphics processing chip according to any one of claims 9 to 16.

Citation Information

Patent Citations

  • Data processing method and system, electronic device and storage medium

    CN114816773A

  • Task processing method and device based on multiple instances

    CN115827174A

  • Geometry processing method and device, equipment and storage medium

    CN117853309A

  • First-in first-out memory capable of achieving random address access and data processing method

    CN119376629A

  • Multi-core consistency processing method, system and device and storage medium

    CN119440881A