Data Processing Method, Device, Medium and Program Product in Pipeline

By using memory barriers and execution state instances with specific check values, the method ensures orderly hardware resource utilization and stable data processing across pipelines, addressing the inefficiencies of existing code modification-based double buffer mechanisms.

CN119883382BActive Publication Date: 2025-07-15SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510378631.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

When implementing the double buffering mechanism in the prior art, the business code needs to be greatly modified, resulting in out-of-order execution of instructions and low hardware resource utilization.

Method used

By configuring memory barrier instances and executing state instances, using the comparison mechanism of barrier verification characters and state verification characters, the data processing order of the pipeline is controlled to ensure that hardware resources are rotated in an orderly manner between pipelines and avoiding out-of-order execution.

Benefits of technology

It improves the utilization rate of hardware resources, ensures that business codes operate according to established logic, and improves the stability and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883382B_ABST
    Figure CN119883382B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of processors, and discloses a data processing method, device, medium and program product in a pipeline. The method includes: when a to-be-executed identifier indicates that there is currently data to be processed, comparing the current verification values of a barrier verification symbol and a status verification symbol, and executing the data to be processed when the current verification values of the barrier verification symbol and the status verification symbol are inconsistent; after the data to be processed is executed, sending a transfer instruction to the adjacent next pipeline, and updating the stage identifier in the execution status instance to enter the next data processing stage, where the transfer instruction is used to guide the next pipeline to update the to-be-executed identifier in its own memory barrier instance. The technical solution provided by one or more embodiments of the present application can improve the utilization rate of hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of processors, and in particular, to a data processing method, device, medium, and program product in a pipeline. Background Art

[0002] In the field of parallel computing of processors, in order to improve the utilization rate of hardware resources, a double buffer mechanism (double buffer) can be adopted to rearrange the instructions in each stage of different pipelines (pipelines), so that the data calculation instructions and data storage instructions in different pipelines overlap in time sequence, so as to make the data calculation instructions on different pipelines take turns to occupy the hardware resources as much as possible, so as to improve the utilization rate of hardware resources.

[0003] In related technologies, in order to implement the above double buffer mechanism, it is usually necessary to transform the service code to adapt to the instruction alternating processing process under the double buffer mechanism. However, when the compiler compiles the transformed service code, it may cause the instructions to be executed out of order, so that the expected effect cannot be achieved, and the utilization rate of hardware resources is still not high enough.

[0004] In view of this, there is a need for a more efficient pipeline data processing method to improve the utilization rate of hardware resources. Summary of the Invention

[0005] This application provides a data processing method, device, medium, and program product in a pipeline, which can improve the utilization rate of hardware resources.

[0006] A first aspect of this application provides a data processing method in a pipeline. The pipeline is configured with a memory barrier instance and an execution status instance. The memory barrier instance includes a to-be-executed flag and a barrier verification symbol, and the execution status instance includes a stage flag and a status verification symbol. The method includes: when the to-be-executed flag indicates that there is currently data to be processed, comparing the current verification values of the barrier verification symbol and the status verification symbol, and when the current verification values of the barrier verification symbol and the status verification symbol are inconsistent, executing the data to be processed; after the data to be processed is executed, sending a transfer instruction to the adjacent next pipeline, and updating the stage flag in the execution status instance to enter the next data processing stage, where the transfer instruction is used to guide the next pipeline to update the to-be-executed flag in its own memory barrier instance.

[0007] The technical solution provided by this embodiment can configure respective memory barrier instances and execution status instances for each pipeline participating in the work, and both of these instances contain their respective member parameters. Among them, for any pipeline, the pending data will be executed only when the verification values of the barrier verification symbol and the status verification symbol are inconsistent. After the pending data is executed, the current pipeline will send a transfer instruction to the adjacent next pipeline and enter the next data processing stage. According to the transfer instruction, the next pipeline can update the pending execution flag in its own memory barrier instance, and then can determine whether there is pending data currently based on the updated pending execution flag. It can be seen that before the pipeline executes the pending data, a comparison process of the barrier verification symbol and the status verification symbol will be carried out, and only the pipelines that meet the conditions can execute the pending data. Such a processing method can avoid the out-of-order execution of instructions and ensure that the business code can run according to the established logic. At the same time, after the pipeline executes the pending data, it will transfer the data processing process to the adjacent next pipeline to ensure that each pipeline can alternately utilize the hardware resources for data processing. Such a method can accurately and stably rotate the hardware resources between different pipelines, thereby greatly improving the utilization rate of the hardware resources.

[0008] In one embodiment, when each pipeline is initialized, the barrier verification symbols in each of the memory barrier instances are configured with a first specified value, and among each of the execution status instances, only the status verification symbol of one execution status instance is configured with a second specified value, and the status verification symbols of the remaining execution status instances are all configured with the first specified value, where the first specified value and the second specified value are different values.

[0009] In this embodiment, the member parameters in the memory barrier instances and execution status instances of each pipeline can be initialized and configured. Among them, only the status verification symbol of one execution status instance will be configured with a second specified value, and the status verification symbols of the remaining execution status instances will all be configured with the same first specified value as the barrier verification symbol. In this way, at the same time period, only one pipeline will meet the condition that the verification values of the barrier verification symbol and the status verification symbol are inconsistent, and the pipeline that meets this condition will enter the data processing process. In this way, for the same piece of hardware resource, only one pipeline will occupy the hardware resource at the same time period, and after the pipeline executes the pending data, it will rotate the hardware resource to the adjacent next pipeline through a transfer instruction. Such a processing method can orderly rotate the hardware resources among different pipelines, and at the same time can avoid the situation that different pipelines preempt the hardware resources at the same time period, improving the stability of data processing.

[0010] In one embodiment, the method further includes: if the to-be-executed flag indicates that there is no data to be processed currently, reset the assignment of the to-be-executed flag and flip the current assignment of the barrier verification symbol to another different value, where the to-be-executed flag after resetting the assignment indicates that there is data to be processed currently.

[0011] In this embodiment, if the to-be-executed flag indicates that there is no data to be processed currently, it is necessary to reset the to-be-executed flag and flip the value of the barrier verification symbol. The purpose of such processing is that the to-be-executed flag after resetting can indicate that there is data to be processed currently, so that the subsequent verification value comparison process between the barrier verification symbol and the status verification symbol can be carried out; and by flipping the value of the barrier verification symbol, it can be ensured that only one pipeline will meet the data processing conditions at the same time period, and then the hardware resources can be rotated stably.

[0012] In one embodiment, sending a transfer instruction to the next adjacent pipeline includes: identifying the pipeline identifier of the current pipeline and incrementally updating the pipeline identifier; obtaining the total number of pipelines currently participating in the work, and taking the remainder of the incrementally updated pipeline identifier with respect to the total number to obtain the target identifier of the next adjacent pipeline; sending a transfer instruction to the pipeline with the target identifier.

[0013] In this embodiment, the target identifier of the next pipeline can be accurately determined through the remainder algorithm. Based on this target identifier, the transfer instruction can be accurately sent, thereby ensuring the accuracy and stability of data processing.

[0014] In one embodiment, sending a transfer instruction to the next adjacent pipeline includes: identifying the data processing stage currently represented by the stage identifier and sending a transfer instruction for the same data processing stage in the next pipeline.

[0015] In this embodiment, since a pipeline usually includes multiple different data processing stages, after the current pipeline completes data processing in the current data processing stage, a transfer instruction should be sent to the same data processing stage of the next pipeline. This can ensure that transfer instructions can be sent back and forth between the same data processing stages, preventing disorder in the data processing stages and improving the correctness and stability of data processing.

[0016] In one embodiment, after updating the phase identifier in the execution status instance, the method further includes: determining whether the updated phase identifier indicates that the next data processing phase returns to the initial phase; if not, after entering the next data processing phase, retaining the updated phase identifier and keeping the value of the status check symbol unchanged; if so, after returning to the initial phase, resetting the phase identifier and flipping the current assignment of the status check symbol to another different value.

[0017] In this embodiment, when the last data processing phase finishes data processing, the process returns to the initial phase. When returning to the initial phase, it is necessary to reset the phase identifier to ensure that the phase identifier corresponds to the actual data processing phase. At the same time, it is also necessary to flip the value of the status check symbol to ensure that only one pipeline can meet the data processing conditions during the same period, and then the hardware resources can be rotated stably.

[0018] In one embodiment, the method further includes: when the current check values of the barrier check symbol and the status check symbol are the same, remaining in a waiting state until the check values of the barrier check symbol and the status check symbol change to be different, and then executing the data to be processed.

[0019] In this embodiment, if the check values of the barrier check symbol and the status check symbol are the same, then the current pipeline cannot perform data processing but needs to remain in a waiting state. This can ensure that only one pipeline occupies the hardware resources during the same period, and the remaining pipelines can be in a waiting state, so as to rotate the hardware resources among different pipelines in an orderly manner. At the same time, it can avoid the situation where different pipelines compete for hardware resources during the same period, improving the stability of data processing.

[0020] On the other hand, this application further provides a data processing device in a pipeline. The pipeline is configured with a memory barrier instance and an execution status instance. The memory barrier instance includes a to-be-executed identifier and a barrier check symbol, and the execution status instance includes a phase identifier and a status check symbol. The device includes:

[0021] A comparison execution unit, configured to compare the current check values of the barrier check symbol and the status check symbol when the to-be-executed identifier indicates that there is data to be processed currently, and execute the data to be processed when the current check values of the barrier check symbol and the status check symbol are different.

[0022] An instruction transfer unit, configured to send a transfer instruction to the next adjacent pipeline after the to-be-processed data is executed, and update the stage identifier in the execution status instance to enter the next data processing stage, where the transfer instruction is used to guide the next pipeline to update the to-be-executed identifier in its own memory barrier instance.

[0023] On the other hand, this application also provides an electronic device, including: a memory storing computer instructions; at least one processor configured to execute the computer instructions in the memory to execute the data processing method in the pipeline in the above-mentioned embodiments.

[0024] On the other hand, this application also provides a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the data processing method in the pipeline in the above-mentioned embodiments.

[0025] On the other hand, this application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the processor is caused to execute the data processing method in the pipeline in the above-mentioned embodiments. Description of the Drawings

[0026] In order to more clearly illustrate the specific embodiments of this application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a schematic diagram of the double-buffer mechanism in the related art;

[0028] Figure 2 It is a flowchart of the data processing method in the pipeline provided by an embodiment of this application;

[0029] Figure 3 It is a schematic diagram of the data processing process of two pipelines provided by an embodiment of this application;

[0030] Figure 4 It is a schematic diagram of the data processing timing provided by an embodiment of this application;

[0031] Figure 5 It is a schematic diagram of the functional modules of the data processing device in the pipeline provided by an embodiment of this application;

[0032] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of this application. Detailed Embodiments

[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.

[0034] In addition, in this application, the descriptions involving "first", "second", etc. are only for descriptive purposes and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the embodiments of this application, unless otherwise stated, the meaning of "a plurality" is two or more. Additionally, the use of "based on" or "according to" implies openness and inclusiveness because a process, step, calculation, or other action "based on" or "according to" one or more of the stated conditions or values may, in practice, be based on additional conditions or values beyond those stated.

[0035] Please refer to Figure 1 , in the related art, the instructions in each stage of the pipeline can be rearranged in the Figure 1 shown manner, so that the data calculation instructions overlap in time sequence with the data loading instructions and data storage instructions. In this way, after processing the data calculation instructions of one pipeline, the hardware resources can be used to process the data calculation instructions of another pipeline without waiting for a long time, thereby improving the utilization rate of the hardware resources.

[0036] To implement Figure 1 this kind of instruction processing manner shown, a great degree of modification to the business code is required in the related art. For example, in a loop structure, it is necessary to list in sequence the statements of various types of instructions for different pipelines. However, when the compiler performs compilation optimization, it is very likely to rearrange these statements, which will thus disrupt the original execution logic of the business code, and thus the instruction overlap scenario shown in Figure 1 cannot be accurately achieved. In addition, due to the out-of-order execution of the instructions, the hardware resources cannot be accurately scheduled to the expected pipeline either.

[0037] An embodiment of the present application provides a data processing method in a pipeline. This method can avoid making a large degree of transformation to business code, but only requires inserting a small number of formatted statements into the business code. At the same time, based on the memory barrier mechanism, it can ensure that the business code is executed in the expected order. Further, it can also achieve precise scheduling of hardware resources by controlling the member parameters in the instance.

[0038] The data processing method in the pipeline provided by the present application can be applied to an artificial intelligence processor, specifically any one of GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural network Processing Unit), DPU (Deep learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose Graphics Processing Unit).

[0039] In the present application, two classes can be pre-implemented:

[0040] PipelineState and PingPongPipelineBar. The functions that these two classes can achieve are specifically described below.

[0041] The PipelineState can contain member parameters stages, stage_id, and parity. Among them, stages can represent the number of data processing stages expected to control the pipeline in the current instance. In practical applications, stages can be assigned corresponding values according to the number of data processing stages. For example, if there are currently 5 data processing stages, then the assignment of stages can be 5. Stage_id can be used to represent the data processing stage where the pipeline is currently located. Generally speaking, the value range of stage_id can be [0, stages - 1]. For example, if the pipeline contains two data processing stages, then stage_id can take the value of 0 or 1. 0 represents the first data processing stage, and 1 represents the second data processing stage. The parity in PipelineState can be used to represent the status verification symbol, and this status verification symbol usually takes the value of 0 or 1. The functions implemented by PipelineState can be summarized as follows: After the current data processing stage ends, stage_id will be incremented to point to the next data processing stage. At the same time, there is also a judgment process for stage_id in PipelineState. If the incremented stage_id is the last data processing stage in the current instance, then stage_id will be reset to 0, and at the same time, the parity in PipelineState will be flipped (from 0 to 1, or from 1 to 0).

[0042] PingPongPipelineBar mainly implements two functions: the wait function and the arrive function. Among them, when the wait function is executed, it can trigger the comparison process of the verification symbol. By comparing whether two verification symbols from different sources are inconsistent, it is determined whether data processing can be performed in the current pipeline. When the arrive function is executed, it can send a transfer instruction to a specified pipeline (usually the adjacent next pipeline), thereby attempting to rotate the hardware resources from the current pipeline to the next pipeline.

[0043] Based on the above two implemented classes, in this application, two different instances can be configured for each pipeline participating in the work: the memory barrier instance (Mbar) and the execution status instance (State). Suppose there are two pipelines currently participating in the data processing process, then both of these two pipelines will have their own memory barrier instances and execution status instances. Different pipelines can manage their respective instances and thus execute the data processing process based on these two instances.

[0044] In this application, the member parameters in the memory barrier instance may include a pending execution identifier (pnd) and a barrier verification symbol (prt), and the member parameters in the execution status instance may include a stage identifier (stage) and a status verification symbol (parity). The following combines Figure 2 and Figure 3 , to illustrate the data processing method in a pipeline in an embodiment of this application.

[0045] S1: When the pending execution identifier indicates that there is currently data to be processed, compare the current verification values of the barrier verification symbol and the status verification symbol, and execute the data to be processed when the current verification values of the barrier verification symbol and the status verification symbol are inconsistent.

[0046] Please refer to Figure 3 , Figure 3 , which provides a schematic diagram of the data processing process with two pipelines as an example. In Figure 3 , six timing sub-diagrams of the data processing process are shown. The numbers (1 to 6) in the upper left corner of each timing sub-diagram can represent the timing identifiers corresponding to the timing sub-diagrams. In each timing sub-diagram, it is assumed that the currently participating pipelines are c0 and c1, and both of these pipelines have two data processing stages, stage0 and stage1. Pipelines c0 and c1 have their respective memory barrier instances (Mbar) and execution status instances (State). Among them, as described above, the memory barrier instance Mbar may include a pending execution identifier (pnd) and a barrier verification symbol (prt), and the member parameters in the execution status instance State may include a stage identifier (stage) and a status verification symbol (parity).

[0047] In an embodiment, when each pipeline is initialized, the member parameters in the memory barrier instance and the execution status instance can be initialized and assigned values. Among them, the barrier verification symbols in each memory barrier instance are all configured with a first specified value, and only one of the status verification symbols in each execution status instance is configured with a second specified value, and the status verification symbols of the remaining execution status instances are all configured with the first specified value, where the first specified value and the second specified value are different values.

[0048] For example, the first specified value may be 0, and the second specified value may be 1. In the timing sub-diagram with a timing identifier of 1 shown in Figure 3 , only the status verification symbol of pipeline c0 among the two pipelines is configured with 1, and the status verification symbol of pipeline c1 is configured with 0. At the same time, the barrier verification symbols of both pipelines are configured with 0. In practical applications, other member parameters can also be initialized. For example, the pending execution identifier in the memory barrier instance can be configured with the data volume of the data to be processed. InFigure 3 In this case, the data volume is represented by 1 (pnd = 1). For another example, the stage identifier in the execution status instance can be configured as 0 (stage0), indicating the start from the first data processing stage. Of course, the above-exemplified initialization assignments are only examples for the convenience of explanation, and do not mean that the technical solution of this application can only be assigned in the above manner. In fact, those skilled in the art can flexibly perform initialization assignments according to the actual situation when understanding the essence of the technical solution of this application.

[0049] In this embodiment, the pipeline can first determine whether the to-be-executed identifier indicates that there is data to be processed currently. Generally speaking, when the assignment of the to-be-executed identifier is not 0, it can indicate that there is data to be processed currently; while when the assignment of the to-be-executed identifier is equal to 0, it indicates that there is no data to be processed currently. Taking Figure 3 the timing sub-graph with the timing identifier of 1 as an example, the pnds in pipelines c0 and c1 are both 1, both indicating that there is data to be processed currently.

[0050] In this embodiment, the member parameters in the memory barrier instance and the execution status instance of each pipeline can be initialized and configured. Among them, only the status checker of one execution status instance will be configured as the second specified value, and the status checkers of the remaining execution status instances will be configured as the same first specified value as the barrier checker. In this way, at the same time period, only one pipeline will meet the condition that the verification values of the barrier checker and the status checker are inconsistent. The pipeline that meets this condition will enter the data processing process. In this way, for the same piece of hardware resource, only one pipeline will occupy this hardware resource at the same time period, and after the pipeline finishes executing the data to be processed, it will rotate this hardware resource to the adjacent next pipeline through a transfer instruction. This processing method can orderly rotate the hardware resource among different pipelines, and at the same time can avoid the situation that different pipelines preempt the hardware resource at the same time period, improving the stability of data processing.

[0051] When it is confirmed that there is data to be processed currently, each pipeline can compare the current verification values of the barrier checker and the status checker, and execute the data to be processed when the current verification values of the barrier checker and the status checker are inconsistent. Taking Figure 3For example, in the timing sub-graph with a timing identifier of 1, the prt of pipeline c0 is 0 and the parity is 1, and the two are inconsistent. At this time, pipeline c0 can execute the data to be processed in stage0. The prt and parity of pipeline c1 are both 0, and the two are consistent. Therefore, pipeline c1 does not meet the preconditions for data processing and cannot execute the data to be processed. That is to say, in the data processing stage of stage0, only pipeline c0 can execute the data to be processed at the same moment; while pipeline c1 can only be in a waiting state.

[0052] It can be seen that in an embodiment of the present application, when the current verification values of the barrier verification symbol and the status verification symbol are consistent, it is necessary to maintain a waiting state until the verification values of the barrier verification symbol and the status verification symbol change to be inconsistent before the data to be processed can be executed.

[0053] In this embodiment, if the verification values of the barrier verification symbol and the status verification symbol are consistent, then the current pipeline cannot perform data processing, but needs to maintain a waiting state. This can ensure that within the same time period, only one pipeline occupies the hardware resources, and the remaining pipelines can be in a waiting state, so as to rotate the hardware resources among different pipelines in an orderly manner. At the same time, it can avoid the situation where different pipelines compete for hardware resources in the same time period, improving the stability of data processing.

[0054] S3: After the data to be processed is executed, send a transfer instruction to the next adjacent pipeline, and update the stage identifier in the execution status instance to enter the next data processing stage, where the transfer instruction is used to guide the next pipeline to update the to-be-executed identifier in its own memory barrier instance.

[0055] In this embodiment, after the pipeline executes the data to be processed, it can send a transfer instruction to the next adjacent pipeline, thereby rotating the hardware resources to the next pipeline. Then, the pipeline can update the stage identifier in the execution status instance to enter the next data processing stage. For Figure 3 example, in the timing sub-graph with a timing identifier of 1, after pipeline c0 executes the data to be processed in stage0, it can send a transfer instruction to pipeline c1, and at the same time update the stage identifier to stage1, so as to enter the data processing stage of stage1 in the timing sub-graph with a timing identifier of 2. It should be noted that after entering the next data processing stage, the assignment of the status verification symbol of this pipeline will not change. Please refer to Figure 3 the timing sub-graph with a timing identifier of 2. After pipeline c0 enters stage1, the parity in the execution status instance State is still 1.

[0056] In this embodiment, after receiving a transfer instruction, the next adjacent pipeline will update the to-be-executed flag in its own memory barrier instance. Specifically, the value assignment of the to-be-executed flag can be decremented by one, indicating that the current data to be processed has been processed by the previous pipeline. Take Figure 3 the sub-graph of the timing identifier 2 as an example. After pipeline c1 receives the transfer instruction, it will decrement pnd from 1 to 0.

[0057] In one embodiment, when the to-be-executed flag indicates that there is no data to be processed currently, it is necessary to reset the value assignment of the to-be-executed flag and flip the current value assignment of the barrier verification symbol to another different value, where the to-be-executed flag after resetting the value assignment indicates that there is data to be processed currently. Please refer to Figure 3 the sub-graph of the timing identifier 2. After pipeline c1 decrements pnd from 1 to 0, pnd indicates that pipeline c1 currently has no data to be processed. At this time, if the overall operation process has not ended, the pnd of pipeline c1 will be reset to 1, and at the same time, the prt of pipeline c1 will also flip from 0 to 1, and then enter the scenario shown in the sub-graph of the timing identifier 3. In the sub-graph of the timing identifier 3, the pnd of pipeline c1 is reset to 1, and at the same time, the prt of pipeline c1 is also reset to 1. In this way, the reset pnd of pipeline c1 can indicate that there is data to be processed currently. It should be noted that pnd only represents the amount of data that needs to be processed by the pipeline in the current data processing stage. When pnd changes from 1 to 0, it indicates that there is no data that needs to be processed by the pipeline in the current data processing stage. However, in the overall data processing process, usually new data to be processed will be continuously loaded into the pipeline. Therefore, when pnd changes from 1 to 0, in order to enable the pipeline to process the newly loaded data normally, pnd needs to be reset to 1. In this way, the pipeline can continue to enter the data processing process.

[0058] In this embodiment, if the to-be-executed flag indicates that there is no data to be processed currently, it is necessary to reset the to-be-executed flag and flip the value of the barrier verification symbol. The purpose of this processing is that the reset to-be-executed flag can indicate that there is data to be processed currently, so that the verification value comparison process between the subsequent barrier verification symbol and the status verification symbol can be carried out; and by flipping the value of the barrier verification symbol, it can be ensured that only one pipeline will meet the data processing conditions at the same time period, and thus the hardware resources can be rotated stably.

[0059] In this embodiment, as Figure 3As shown in the timing sub - graph with a timing identifier of 3, after pipeline c1 resets the identifier to be executed and flips the barrier verification symbol, it can compare the current verification values of the barrier verification symbol and the status verification symbol. At this time, the barrier verification symbol prt is 1, and the status verification symbol parity is 0. Since the two are inconsistent, pipeline c1 can then start the data processing process of stage0. Similarly, after pipeline c1 completes the data processing process of stage0, it can send a transfer instruction to the adjacent pipeline c0 and can update the stage identifier in its own execution status instance (from stage0 to stage1). During this process, the status verification symbol parity of pipeline c1 remains unchanged and is still 0. Specifically, reference can be made to Figure 3 the timing sub - graph with a timing identifier of 4 in Figure 3 . In this timing sub - graph, the stage identifier of pipeline c1 is updated from Stage0 to Stage1, and Parity still remains 0.

[0060] From Figure 3 the timing sub - graph with a timing identifier of 2, it can be seen that after pipeline c0 enters stage1, since the assignments of prt and parity are still inconsistent, pipeline c0 can still execute the data processing process of stage1. From the timing perspective, the process of pipeline c0 executing stage1 is synchronous with the process of pipeline c1 executing stage0. However, in practical applications, these two processes do not preempt the same hardware resources. The reason is that in practical applications, stage0 and stage1 can correspond to different types of instructions. For example, stage0 can correspond to data calculation instructions, and stage1 can correspond to data storage instructions. When these two instructions are actually processed, they are processed by different hardware resources respectively, so that stage1 of pipeline c0 and stage0 of pipeline c1 can proceed simultaneously. The final achieved timing processing effect can be as shown in Figure 4 shown. In Figure 4 , mma can represent the processing process of data calculation instructions, and store can represent the processing process of data storage instructions. At the beginning, both pipeline c0 and c1 will perform the comparison process of verification values. The result is that only pipeline c0 can perform the processing process of stage0, thus corresponding to Figure 4 where at the beginning only pipeline c0 executes mma. When pipeline c0 enters stage1 and starts to process data storage instructions, pipeline c1 can simultaneously enter the processing process of stage0 (i.e., the processing process of mma). In this way, the store of pipeline c0 and the mma of pipeline c1 can overlap in timing, and the subsequent process follows the same pattern. From Figure 4As can be seen, for the hardware resources for processing mma, they will rotate back and forth between pipelines c0 and c1, greatly improving the utilization rate of the hardware resources.

[0061] The technical solution provided by this embodiment can configure respective memory barrier instances and execution status instances for each pipeline participating in the work, and both of these instances contain their respective member parameters. Among them, for any pipeline, the data to be processed will only be executed when the verification values of the barrier verification symbol and the status verification symbol are inconsistent. After the data to be processed is executed, the current pipeline will send a transfer instruction to the adjacent next pipeline and enter the next data processing stage. According to the transfer instruction, the next pipeline can update the to-be-executed flag in its own memory barrier instance, and then can determine whether there is data to be processed currently based on the updated to-be-executed flag. It can be seen that before the pipeline executes the data to be processed, a comparison process of the barrier verification symbol and the status verification symbol will be performed, and only the pipelines that meet the conditions can execute the data to be processed. Such a processing method can avoid the out-of-order execution of instructions and ensure that the business code can run according to the established logic. At the same time, after the pipeline executes the data to be processed, it will transfer the data processing process to the adjacent next pipeline to ensure that each pipeline can alternately utilize the hardware resources for data processing. Such a method can accurately and stably rotate the hardware resources between different pipelines, thereby greatly improving the utilization rate of the hardware resources.

[0062] In one embodiment, after updating the stage flag in the execution status instance, it can be determined whether the updated stage flag represents that the next data processing stage returns to the initial stage. If not, after entering the next data processing stage, the updated stage flag is retained and the value of the status verification symbol remains unchanged; if so, after returning to the initial stage, the stage flag is reset and the current assignment of the status verification symbol is flipped to another different value. Figure 3Taking the timing sub - graph with a timing identifier of 4 as an example, after pipeline c0 finishes the data processing in stage1, it will continue to accumulate the stage identifier. In binary, the result of accumulating the stage identifier is to change from stage1 back to stage0. At this time, the updated stage identifier indicates that the next data processing stage has returned to the initial stage (the stage of stage0). In this case, pipeline c0 needs to reset the stage identifier to stage0 (in the execution state instance State of pipeline c0 in the timing sub - graph with a timing identifier of 4, it is Stage0), and needs to flip the status checker parity from 1 to 0 (in the execution state instance State of pipeline c0 in the timing sub - graph with a timing identifier of 4, it is Parity0). If the updated stage identifier does not indicate that the next data processing stage has returned to the initial stage, then after entering the next data processing stage, the updated stage identifier can be retained, and the value of the status checker can be maintained unchanged. For example, in the timing sub - graph with a timing identifier of 2, when pipeline c0 enters stage1 from stage0, the stage identifier can be retained as the updated stage1, and at the same time, parity is not flipped and remains 1. Subsequently, the parameter change processes presented in the timing sub - graphs with timing identifiers of 5 and 6 are similar to the parameter change processes in the aforementioned timing sub - graphs, and will not be elaborated here.

[0063] In this embodiment, when the last data processing stage finishes data processing, the process will return to the initial stage. When returning to the initial stage, the stage identifier needs to be reset to ensure that the stage identifier corresponds to the actual data processing stage. At the same time, the value of the status checker needs to be flipped to ensure that only one pipeline will meet the data processing conditions at the same time, and then the hardware resources can be rotated stably.

[0064] From the above description, it can be found that by flipping the values of the barrier checker and the status checker in different scenarios, it can exactly ensure that the hardware resources of the data calculation instructions are continuously rotated between the two pipelines, thus implementing the double - buffer mechanism.

[0065] In one embodiment, in order to accurately send the transfer instruction, when sending the transfer instruction to the next adjacent pipeline, the pipeline identifier of the current pipeline can be identified and incrementally updated. Then, the total number of pipelines currently participating in the work can be obtained, and the incrementally updated pipeline identifier can be used to perform a modulo operation on the total number to obtain the target identifier of the next adjacent pipeline. In this way, the transfer instruction can be sent to the pipeline with the target identifier. Figure 3Taking the timing sub - graph with the timing identifier of 1 as an example, the pipeline identifier of pipeline c0 can be 0. After incrementing and updating, it can change from 0 to 1. The total number of pipelines currently participating in the work is 2. Taking the remainder of 1 divided by 2 gives 1. In this way, a transfer instruction can be sent to pipeline c1. Similarly, after pipeline c1 completes the data processing flow, it can increment and update 1 to get 2, and then take the remainder of 2 divided by 2 to get 0. Then a transfer instruction can be sent to pipeline c0.

[0066] In this embodiment, the target identifier of the next pipeline can be accurately determined through the modulo algorithm. Based on this target identifier, transfer instructions can be accurately sent, thus ensuring the accuracy and stability of data processing.

[0067] In one embodiment, sending a transfer instruction to the adjacent next pipeline includes: identifying the data processing stage currently represented by the stage identifier, and sending a transfer instruction for the same data processing stage in the next pipeline. Taking Figure 3 the timing sub - graph with the timing identifier of 1 as an example, when pipeline c0 sends a transfer instruction to pipeline c1, since it is currently in the data processing stage of stage0, pipeline c0 will still send a transfer instruction for the data processing stage of stage0 in pipeline c1, rather than sending a transfer instruction for the data processing stage of stage1 in pipeline c1.

[0068] In this embodiment, since a pipeline usually contains multiple different data processing stages, after the current pipeline completes data processing in the current data processing stage, a transfer instruction should be sent to the same data processing stage in the next pipeline. This can ensure that transfer instructions can be sent back and forth between the same data processing stages, preventing disorder in the data processing stages and improving the correctness and stability of data processing.

[0069] From the perspective of the implementation logic of the business code, in this application, only the execution statement of the wait function needs to be added before the original data calculation instruction, and the execution statement of the transfer function needs to be added after the data calculation instruction; similarly, the execution statement of the wait function can also be added before the original data storage instruction, and the execution statement of the transfer function can be added after the data storage instruction. In this way, the execution logic of the above - mentioned various embodiments can be implemented at the business code level. This method has a small degree of modification to the original business code, and the execution statements of the wait function and the transfer function can form a set of collaborative code, and the code modification is relatively simple.

[0070] Please refer to Figure 5, on the other hand, this application also provides a data processing device in a pipeline. The pipeline is configured with a memory barrier instance and an execution status instance. The memory barrier instance includes a to-be-executed flag and a barrier verification symbol, and the execution status instance includes a stage flag and a status verification symbol; the device includes:

[0071] A comparison execution unit 100, configured to, when the to-be-executed flag indicates that there is to-be-processed data currently, compare the current verification values of the barrier verification symbol and the status verification symbol, and execute the to-be-processed data when the current verification values of the barrier verification symbol and the status verification symbol are inconsistent.

[0072] An instruction transfer unit 200, configured to, after the to-be-processed data is executed, send a transfer instruction to the adjacent next pipeline, and update the stage flag in the execution status instance to enter the next data processing stage, where the transfer instruction is used to guide the next pipeline to update the to-be-executed flag in its own memory barrier instance.

[0073] In one embodiment, when each pipeline is initialized, the barrier verification symbols in each of the memory barrier instances are configured with a first specified value, and only one of the status verification symbols in each of the execution status instances is configured with a second specified value, and the status verification symbols of the remaining execution status instances are all configured with the first specified value, where the first specified value and the second specified value are different values.

[0074] In one embodiment, if the to-be-executed flag indicates that there is no to-be-processed data currently, reset the assignment of the to-be-executed flag, and flip the current assignment of the barrier verification symbol to another different value, where the to-be-executed flag after resetting the assignment indicates that there is to-be-processed data currently.

[0075] In one embodiment, the instruction transfer unit is specifically configured to identify the pipeline identifier of the current pipeline, and incrementally update the pipeline identifier; obtain the total number of pipelines currently participating in work, and perform a remainder operation on the total number using the incrementally updated pipeline identifier to obtain the target identifier of the adjacent next pipeline; send a transfer instruction to the pipeline with the target identifier.

[0076] In one embodiment, the instruction transfer unit is specifically configured to identify the data processing stage currently represented by the stage flag, and send a transfer instruction for the same data processing stage in the next pipeline.

[0077] In one embodiment, the instruction transfer unit is further configured to determine whether the updated stage identifier indicates that the next data processing stage returns to the initial stage. If not, after entering the next data processing stage, the updated stage identifier is retained, and the value of the status verification symbol is maintained unchanged. If so, after returning to the initial stage, the stage identifier is reset, and the current assignment of the status verification symbol is flipped to another different value.

[0078] In one embodiment, the comparison execution unit is further configured to, when the current verification values of the barrier verification symbol and the status verification symbol are consistent, remain in a waiting state until the verification values of the barrier verification symbol and the status verification symbol change to be inconsistent, and then execute the data to be processed.

[0079] On the other hand, the present application further provides an electronic device, including: a memory storing computer instructions; at least one processor configured to execute the computer instructions in the memory to perform the data processing method in the pipeline in the above embodiment.

[0080] On the other hand, the present application further provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed by a processor, the processor is caused to perform the data processing method in the pipeline in the above embodiment.

[0081] On the other hand, the present application further provides a computer program product including computer instructions. When the computer instructions are executed by a processor, the processor is caused to perform the data processing method in the pipeline in the above embodiment.

[0082] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of an electronic device provided by an alternative embodiment of the present invention. As shown in Figure 6 , the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common main board or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In

[0083] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0084] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0085] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may further include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the electronic device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0086] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.

[0087] The electronic device further includes a communication interface 30 for the electronic device to communicate with other devices or a communication network.

[0088] The embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention may be implemented in hardware, firmware, or be implemented as computer code recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and to be stored in a local storage medium, so that the method described herein may be stored in such software processed on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may further include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0089] A part of the present invention can be applied as a computer program product, for example, computer program instructions, which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0090] This application is described with reference to the flowcharts and / or block diagrams of methods and systems according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0091] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0092] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0093] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0094] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0095] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

[0096] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A data processing method in a pipeline, characterized in that, The pipeline is configured with a memory barrier instance and an execution status instance. The memory barrier instance includes a to-be-executed flag and a barrier verification symbol. The execution status instance includes a stage flag and a status verification symbol. The method includes: When the to-be-executed flag indicates that there is currently data to be processed, compare the current verification values of the barrier verification symbol and the status verification symbol, and execute the data to be processed when the current verification values of the barrier verification symbol and the status verification symbol are inconsistent; After the data to be processed is executed, send a transfer instruction to the adjacent next pipeline, and update the stage flag in the execution status instance to enter the next data processing stage, where the transfer instruction is used to guide the next pipeline to update the to-be-executed flag in its own memory barrier instance; Among them, sending a transfer instruction to the adjacent next pipeline includes: Identify the data processing stage currently represented by the stage flag, and send a transfer instruction for the same data processing stage in the next pipeline.

2. The method according to claim 1, wherein When each pipeline is initialized, the barrier verification symbols in each memory barrier instance are configured with a first specified value, and only one status verification symbol in each execution status instance is configured with a second specified value, and the status verification symbols of the remaining execution status instances are all configured with the first specified value, where the first specified value and the second specified value are different values.

3. The method according to claim 1, wherein The method further includes: If the to-be-executed flag indicates that there is currently no data to be processed, reset the assignment of the to-be-executed flag, and flip the current assignment of the barrier verification symbol to another different value, where the to-be-executed flag after resetting the assignment indicates that there is currently data to be processed.

4. The method according to claim 1, wherein Sending a transfer instruction to the adjacent next pipeline includes: Identify the pipeline identifier of the current pipeline, and incrementally update the pipeline identifier; Obtain the total number of pipelines currently participating in the work, and perform a remainder operation on the total number using the incrementally updated pipeline identifier to obtain the target identifier of the adjacent next pipeline; Send a transfer instruction to the pipeline with the target identifier.

5. The method according to claim 1 or 3, characterized in that, After updating the stage flag in the execution status instance, the method further includes: Judge whether the updated stage flag indicates that the next data processing stage returns to the initial stage. If not, after entering the next data processing stage, retain the updated stage flag and keep the value of the status verification symbol unchanged; if so, after returning to the initial stage, reset the stage flag and flip the current assignment of the status verification symbol to another different value.

6. The method according to claim 1 or 3, characterized in that, The method further includes: When the current verification values of the barrier verification symbol and the status verification symbol are consistent, remain in a waiting state until the verification values of the barrier verification symbol and the status verification symbol change to be inconsistent, and then execute the data to be processed.

7. A data processing device in a pipeline, characterized in that, The pipeline is configured with a memory barrier instance and an execution status instance. The memory barrier instance includes a to-be-executed flag and a barrier verification symbol. The execution status instance includes a stage flag and a status verification symbol. The device includes: A comparison execution unit, configured to compare the current verification values of the barrier verification symbol and the status verification symbol when the to-be-executed identifier indicates that there is currently data to be processed, and execute the to-be-processed data when the current verification values of the barrier verification symbol and the status verification symbol are inconsistent; An instruction transfer unit, configured to send a transfer instruction to the next adjacent pipeline after the to-be-processed data is executed, and update the stage identifier in the execution status instance to enter the next data processing stage, where the transfer instruction is used to guide the next pipeline to update the to-be-executed identifier in its own memory barrier instance; Wherein, sending a transfer instruction to the next adjacent pipeline includes: Identifying the data processing stage currently represented by the stage identifier, and sending a transfer instruction for the same data processing stage in the next pipeline.

8. An electronic device, characterized in that, Comprising: A memory storing computer instructions; At least one processor, configured to execute the computer instructions in the memory to execute the method according to any one of claims 1-6.

9. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instructions are executed by the processor, the processor is caused to execute the method according to any one of claims 1-6.

10. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the processor is caused to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and device for ensuring persistent memory data crash consistency

    CN119025029A

  • Trusted switch file execution method and device, electronic equipment and storage medium

    CN119089429A