Processor performance analysis method, device and electronic device

By executing benchmark test programs in the processor pipeline, generating graph data and analyzing processor performance, the problem of low silicon performance analysis efficiency in the prior art is solved, and fast and efficient performance analysis is achieved.

CN114840371BActive Publication Date: 2025-05-09LOONGSON TECH CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210492699.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2025-05-09
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

In the prior art, the processor pre-silicon performance analysis efficiency is low, resulting in long analysis time and low efficiency.

Method used

By using pipelines in the processor to execute multiple instructions in the benchmark program, generating nodes and edges for each pipeline stage, and using delay time as the weight of edges, graph data is constructed to analyze processor performance.

Benefits of technology

This method can quickly determine the factors that lead to degraded processor performance, improve analysis efficiency, and systematically present the processor operation process, supporting systematic analysis of processor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114840371B_ABST
    Figure CN114840371B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a processor performance analysis method, device and electronic device, which relates to the field of computers. The method includes: during the execution of each instruction, generating the first node of the current pipeline stage, when the execution of the current pipeline stage depends on the execution result of the historical pipeline stage that has ended, obtaining the delay time, generating an edge between the second node and the first node, and using the delay time as the weight of the edge, storing the generated nodes, edges and edge weights of each pipeline stage, obtaining graph data, and analyzing the performance of the processor according to the critical path in the graph data. During the running process of the benchmark program, graph data is constructed for each pipeline stage, and the performance of the processor is analyzed using the critical path in the graph data. The factors that cause the processor to reduce performance can be quickly determined, and the microstructure of the processor can be adjusted to improve the analysis efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to a processor performance analysis method, device and electronic equipment. Background Art

[0002] Pre-silicon performance analysis is required during the development of the processor (Central Processing Unit, CPU). The main purpose of pre-silicon performance analysis is to discover the performance bottlenecks of the existing design based on the program behavior of the benchmark program, and to improve the performance of the existing design through design balance optimization and key component optimization.

[0003] In the prior art, the pre-silicon performance analysis method mainly includes pre-silicon performance analysis based on simulator and register transfer level (RTL) code. The above method takes a long time, which makes the efficiency of pre-silicon performance analysis low. Summary of the invention

[0004] In view of the above problems, an embodiment of the present invention is proposed to provide a processor performance analysis method that overcomes the above problems or at least partially solves the above problems, so as to solve the problem of low efficiency of pre-silicon performance analysis of processors.

[0005] Correspondingly, an embodiment of the present invention also provides a processor performance analysis device and an electronic device to ensure the implementation and application of the above method.

[0006] A first aspect of an embodiment of the present invention discloses a processor performance analysis method, which is applied to a processor loaded with a benchmark test program, wherein the benchmark test program includes a plurality of instructions; the method includes:

[0007] Executing the plurality of instructions in a pipeline in the execution order of the plurality of instructions; the execution process of each of the instructions includes a plurality of pipeline stages;

[0008] During the execution of each of the instructions, generating a first node of the current pipeline stage;

[0009] In a case where the execution of the current pipeline stage depends on the execution result of a completed historical pipeline stage, obtaining a delay time between the current pipeline stage and the historical pipeline stage;

[0010] Generate an edge between a second node of the historical pipeline stage and the first node, and use the delay time as a weight of the edge; the source node of the edge is the second node, and the destination node is the first node;

[0011] The generated nodes of each pipeline stage, the edges and the weights of the edges are stored to obtain graph data, so as to analyze the performance of the processor based on a critical path between a start node and an end node in the graph data.

[0012] Optionally, the storing of the generated nodes of each pipeline stage, the edges and the weights of the edges to obtain graph data includes:

[0013] Dividing nodes of all pipeline stages in each of the instructions into a node set;

[0014] The node set and the first type of edges connecting the nodes in the node set are stored as a subgraph, and the second type of edges connecting the nodes in different sets are stored to obtain graph data consisting of the subgraph and the second type of edges.

[0015] Optionally, storing the node set and the first type of edges connecting the nodes in the node set as a subgraph, and storing the second type of edges connecting the nodes in different sets, to obtain graph data consisting of the subgraph and the second type of edges, includes:

[0016] Establishing and storing a graph data header, wherein the graph data header includes structural information of the subgraph;

[0017] Adding connection information of the second type of edge in the graph data header, where the connection information is used to describe a source node and a destination node of the second type of edge;

[0018] According to the execution order, a unique subgraph number is sequentially set for the subgraph corresponding to each of the instructions and stored, and when the subgraph includes a key node, a first information set of the first type of edge in the subgraph is correspondingly stored, and a second information set of the key node is stored;

[0019] Among them, the first information set includes the weight of the first type of edge; the key node is the destination node of the target second type edge; the second information set includes the subgraph number of the subgraph where the source node of the target second type edge is located and the weight of the target second type edge, as well as the index of the connection information of the target second type edge in the graph data header.

[0020] Optionally, after the nodes of each pipeline stage, the edges and the weights of the edges generated by the storage are obtained to obtain graph data, the method further includes:

[0021] Selecting nodes in at least one subgraph from the multiple subgraphs in sequence as nodes to be calculated according to the coding order of the nodes;

[0022] In the case where the node to be calculated has at least one incoming edge, the path distance to be selected is determined according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge; the target path distance of the target source node is the longest path distance and / or the shortest path distance among all path distances between the starting node and the target source node;

[0023] Determine the target path distance of the node to be calculated from all the path distances to be selected, and use the source node corresponding to the target path distance of the node to be calculated as the previous node of the node to be calculated; the target path distance of the node to be calculated is the maximum path distance to be selected corresponding to the longest path distance and / or the minimum path distance to be selected corresponding to the shortest path distance;

[0024] The steps of selecting the node to be calculated, determining the distance of the selected path, and determining the target path distance are repeatedly performed until the target path distances and previous nodes of all the nodes are determined.

[0025] Optionally, determining the distance of the to-be-selected path according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge includes:

[0026] If the target source node is the starting node, the weight of the incoming edge is used as the distance of the path to be selected;

[0027] If the target path distance of the target source node is predetermined, the weight of the incoming edge and the target path distance of the connected source node are summed to obtain the path distance to be selected;

[0028] If the target path distance of the target source node is not determined in advance, the summing step is performed after the target path distance of the target source node is determined to obtain the candidate path distance.

[0029] Optionally, it also includes:

[0030] When the node to be calculated is a node other than the starting node and has no incoming edge, the target path distance between the node to be calculated and the starting node is set to a first preset value, and the previous node of the node to be calculated is set to a second preset value; the first preset value is an infinitesimal value corresponding to the longest path distance value and / or an infinite value corresponding to the shortest path distance value; the second preset value indicates that the previous node of the node to be calculated is empty.

[0031] Optionally, determining the distance of the to-be-selected path according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge includes:

[0032] The step of determining the distance of the path to be selected is performed on the nodes to be calculated in the at least one selected subgraph in sequence according to the coding order.

[0033] A second aspect of an embodiment of the present invention discloses a processor performance analysis device, which is arranged on a processor loaded with a benchmark test program, wherein the benchmark test program includes a plurality of instructions; the device includes:

[0034] An execution module, configured to execute the plurality of instructions in a pipeline in the execution order of the plurality of instructions; the execution process of each of the instructions includes a plurality of pipeline stages;

[0035] A first generating module, used for generating a first node of a current pipeline stage during the execution of each of the instructions;

[0036] An acquisition module, configured to acquire a delay time between the current pipeline stage and the historical pipeline stage when the execution of the current pipeline stage depends on the execution result of the completed historical pipeline stage;

[0037] a second generating module, configured to generate an edge between a second node of the historical pipeline stage and the first node, and use the delay time as a weight of the edge; the source node of the edge is the second node, and the destination node is the first node;

[0038] A storage module is used to store the generated nodes of each pipeline stage, as well as the edges and the weights of the edges, to obtain graph data so as to analyze the performance of the processor based on the critical path between the starting node and the ending node in the graph data.

[0039] Optionally, the storage module includes:

[0040] A partitioning unit, used for partitioning the nodes of all pipeline stages in each of the instructions into a node set;

[0041] A storage unit is used to store the node set and the first type of edges connecting the nodes in the node set as a subgraph, and store the second type of edges connecting the nodes in different sets, so as to obtain graph data composed of the subgraph and the second type of edges.

[0042] Optionally, the storage unit includes:

[0043] Establishing a sub-unit, used to establish and store a graph data header, wherein the graph data header includes structural information of the sub-graph;

[0044] An adding subunit is used to add connection information of the second type of edge in the graph data header, wherein the connection information is used to describe a source node and a destination node of the second type of edge;

[0045] A storage subunit, configured to set and store a unique subgraph number for each subgraph corresponding to the instruction in the execution order, and, when the subgraph includes a key node, to store a first information set of a first type of edge in the subgraph and a second information set of the key node;

[0046] Among them, the first information set includes the weight of the first type of edge; the key node is the destination node of the target second type edge; the second information set includes the subgraph number of the subgraph where the source node of the target second type edge is located and the weight of the target second type edge, as well as the index of the connection information of the target second type edge in the graph data header.

[0047] Optionally, it also includes:

[0048] A selection unit, configured to select nodes in at least one subgraph from the plurality of subgraphs in sequence as nodes to be calculated according to the coding order of the nodes;

[0049] A first determining unit is used to determine the path distance to be selected according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge when the node to be calculated has at least one incoming edge; the target path distance of the target source node is the longest path distance and / or the shortest path distance among all path distances between the starting node and the target source node;

[0050] A second determining unit is used to determine the target path distance of the node to be calculated from all the path distances to be selected, and use the source node corresponding to the target path distance of the node to be calculated as the previous node of the node to be calculated; the target path distance of the node to be calculated is the maximum path distance to be selected corresponding to the longest path distance and / or the minimum path distance to be selected corresponding to the shortest path distance;

[0051] The repeating unit is used to repeatedly execute the steps of selecting the node to be calculated, determining the distance of the selected path, and determining the target path distance until the target path distances and previous nodes of all the nodes are determined.

[0052] Optionally, the first determination unit is specifically used to, if the target source node is the starting node, use the weight of the incoming edge as the path distance to be selected; if the target path distance of the target source node is predetermined, sum the weight of the incoming edge and the target path distance of the connected source node to obtain the path distance to be selected; if the target path distance of the target source node is not predetermined, wait until the target path distance of the target source node is determined, and then execute the summation step to obtain the path distance to be selected.

[0053] Optionally, the first determination unit is also used to set the target path distance between the node to be calculated and the starting node to a first preset value, and set the previous node of the node to be calculated to a second preset value when the node to be calculated is a node other than the starting node and has no incoming edge; the first preset value is an infinitesimal value corresponding to the longest path distance value and / or an infinite value corresponding to the shortest path distance value; the second preset value indicates that the previous node of the node to be calculated is empty.

[0054] Optionally, the first determining unit is specifically configured to sequentially execute the step of determining the distance of the to-be-selected path on the to-be-calculated nodes in the at least one selected subgraph according to the coding order.

[0055] The embodiment of the present invention further discloses an electronic device, including the processor performance analysis device as described above. The embodiment of the present invention has the following advantages:

[0056] In an embodiment of the present invention, a plurality of instructions are executed in a pipeline in the order of execution of the plurality of instructions. During the execution of each instruction, the first node of the current pipeline stage is generated. When the execution of the current pipeline stage depends on the execution result of the historical pipeline stage that has ended, the delay time between the current pipeline stage and the historical pipeline stage is obtained, and an edge between the second node and the first node of the historical pipeline stage is generated. The delay time is used as the weight of the edge. The generated nodes of each pipeline stage, as well as the edges and the weights of the edges are stored to obtain graph data, so as to analyze the performance of the processor based on the critical path between the start node and the end node in the graph data. During the running process of the benchmark program, the nodes of each pipeline stage and the edges between two pipeline stages with a dependency relationship are generated, and the delay time between the two pipeline stages is used as the weight of the edge to construct the graph data. By using the critical path in the graph data to analyze the performance of the processor, the factors that cause the processor to reduce performance can be quickly determined, and the microstructure of the processor can be adjusted to improve the analysis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A flowchart showing a method for analyzing processor performance according to an embodiment of the present invention is shown;

[0058] Figure 2 A schematic diagram of the structure of a prototype system in an embodiment of the present invention is shown;

[0059] Figure 3 A schematic diagram of instruction execution in an embodiment of the present invention is shown;

[0060] Figure 4 A schematic diagram of a graph data header in an embodiment of the present invention is shown;

[0061] Figure 5 A schematic diagram of storing graph data in an embodiment of the present invention is shown;

[0062] Figure 6 A schematic diagram of the structure of a pre-processing device in an embodiment of the present invention is shown;

[0063] Figure 7 A schematic diagram of the structure of a processor performance analysis device in an embodiment of the present invention is shown;

[0064] Figure 8 A structural block diagram of an electronic device in an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0065] In order to facilitate understanding of the embodiments of the present invention, the graph data involved in the embodiments of the present invention is first briefly introduced.

[0066] A graph is an abstract data structure used to represent the relationship between objects. It is described using nodes (Vertex) and edges (Edge). Nodes can also be called vertices. Nodes represent objects, and edges represent the relationship between objects. Data that can be abstracted into graphs is called graph data. Graph data usually includes multiple nodes and edges connecting multiple nodes. In some scenarios, edges have weights, which can represent the degree of association between the two objects corresponding to the two nodes connected by the edge.

[0067] Reference Figure 1 , shows a flowchart of a processor performance analysis method embodiment of the present invention, which is applied to a processor loaded with a benchmark test program, wherein the benchmark test program includes multiple instructions; the method can be applied in pre-silicon performance analysis of the processor; the method may include:

[0068] Step 101: Execute multiple instructions in a pipeline according to the execution order of the multiple instructions.

[0069] Among them, the execution process of each instruction includes multiple pipeline stages.

[0070] In this embodiment, the processor performance analysis method can be executed by a processor, and a benchmark program can be loaded into the processor to run the benchmark program. The processor can be an actual processor chip, or a processor simulated by a prototype system. The prototype system is, for example, a prototype system implemented by a field programmable gate array (FPGA), and the register transfer level (RTL) code of the processor is loaded into the prototype system. The FPGA can simulate the behavior of the processor obtained by RTL code tape-out based on the RTL code. The benchmark program can be loaded into the FPGA, and the benchmark program can be run after the FPGA is started. Figure 2 As shown, Figure 2 A structural schematic diagram of a prototype system in an embodiment of the present invention is shown. The prototype system 200 includes a processor 201 obtained by FPGA simulation. The processor 201 includes a reorder buffer 2011 (Re-Order Buffer, ROB) and an instruction information cache table (instruction information table) 2012. The reorder buffer 2011 can obtain relevant information of the processor 201 during the execution of the benchmark program and store it in the instruction information cache table 2012. Figure 2 Symbols i0, i1, i2, i3 and i4 represent multiple instructions in the benchmark program respectively, and the arrows represent the direction of information transmission. The processor can execute instruction i0, instruction i1, instruction i2, instruction i3 and instruction i4 in the execution order of the multiple instructions.

[0071] For example, Figure 3 As shown, Figure 3A schematic diagram of instruction execution in an embodiment of the present invention is shown. After the processor starts to run the benchmark test program, it can use pipeline technology to execute multiple instructions in the benchmark test program in the execution order of multiple instructions. When the pipeline technology is used to execute multiple instructions in the benchmark test program, the execution process of each instruction can be divided into 5 pipeline stages, namely, the dispatch stage, the ready stage, the execute stage, the complete stage and the commit stage. The dispatch stage is represented by symbol D, the ready stage is represented by symbol R, the execute stage is represented by symbol E, the complete stage is represented by symbol P, and the commit stage is represented by symbol C. When the pipeline technology is used to execute multiple instructions, the CPU processes multiple instructions in parallel. For example, in the preparation stage of instruction i0, the dispatch stage of instruction i1 can be performed. In the execution stage of instruction i0, the preparation stage of instruction i1 can be performed, and the dispatch stage of instruction i2 can be performed. By analogy, multiple instructions can be processed in parallel. The division method of the pipeline stages may include but is not limited to the above examples, and this embodiment does not limit this.

[0072] Figure 3 Each circle in the figure can represent a pipeline stage in an instruction, and the subscript of the symbol corresponds to the instruction. In the same instruction, the execution of the previous pipeline stage depends on the execution result of the next pipeline stage, and there is an association relationship between the two stages. For example, in instruction i0, the execution of the preparation stage R0 depends on the execution result of the dispatch stage D0. Between two pipeline stages of different instructions, the execution of a pipeline stage in the next instruction depends on the execution result of a pipeline stage in the previous instruction, and the two pipeline stages have an association relationship. For example, between instruction i3 and instruction i1, the preparation stage of instruction i3 needs to obtain the execution result of the completion stage of instruction i1, and there is an association relationship between the preparation stage of instruction i3 and the completion stage of instruction i1. The above is only an illustrative example, and the association relationship between each pipeline stage can be set according to actual needs, and this embodiment does not limit this.

[0073] Step 102: During the execution of each instruction, generate the first node of the current pipeline stage.

[0074] Step 103: When the execution of the current pipeline stage depends on the execution result of the completed historical pipeline stage, the delay time between the current pipeline stage and the historical pipeline stage is obtained.

[0075] The current pipeline stage is the pipeline stage being executed in the instruction. The historical pipeline stage is the pipeline stage that has ended. For example, during the execution of instruction i1, if the dispatch stage D1 has been executed and the preparation stage R1 is currently in progress, then the preparation stage R1 is the current pipeline stage and the dispatch stage D1 is the historical pipeline stage. At this time, if the completion stage P0 in instruction i0 has also been executed, then the completion stage P0 is also the historical pipeline stage.

[0076] In this embodiment, the CPU can generate a node for each pipeline stage during the execution of each instruction, and can determine the historical pipeline stage associated with the current pipeline stage. Combined with the above example, during the execution of instruction i1, if the execution of the preparation stage R1 obtains the execution result of the dispatch stage D1, it is determined that the preparation stage R1 has an associated relationship with the dispatch stage D1. Similarly, if the execution of the preparation stage R1 obtains the execution result of the completion stage P0 in instruction i0, it is determined that the preparation stage R1 has an associated relationship with the completion stage P0.

[0077] in, Figure 2 Each circle in can also represent a node corresponding to a pipeline stage, the symbol in the circle can be the node number of the node, and the line between two nodes represents the edge between the two nodes. For each pipeline stage in the instruction, the CPU can generate a node for the current pipeline stage in progress, that is, the first node. During the execution of instruction i0, for the five pipeline stages in instruction i0, five nodes D0, R0, E0, P0 and C0 can be generated respectively according to the pipeline order of multiple nodes, and the five nodes have a coding order from D0 to C0. Similarly, the nodes D1, R1, E1, P1 and C1 of the five pipeline stages of instruction i1 can be generated, as well as the five nodes corresponding to instruction i2, instruction i3 and instruction i4 respectively.

[0078] While generating the node corresponding to the pipeline stage, the historical pipeline stage associated with the current pipeline stage can be determined, and the delay time between the current pipeline stage and the historical pipeline stage can be obtained. For example, during the execution of instruction i0, the CPU can start timing at the dispatch stage D0, and when executing to the preparation stage R0, a first timing duration is obtained, and the first timing duration is the delay time between the dispatch stage D0 and the preparation stage R0. When executing to the execution stage E0, a second timing duration is obtained, and the time difference between the first timing duration and the second timing duration is the delay time between the execution stage E0 and the preparation stage R0. For the delay time between two pipeline stages in different instructions, the CPU can start timing from the dispatch stage D0, and record the initial delay time between each subsequent pipeline stage and the dispatch stage D0 in sequence. The delay time between the two pipeline stages can be determined based on the time difference between the initial delay time between the two pipeline stages and the dispatch stage D0. For example, the initial delay time between the dispatch stage D0 and the completion stage P0 is X, the initial delay time between the dispatch stage D0 and the preparation stage R1 is Y, and the time difference between the initial delay time Y and the initial delay time X is the delay time between the completion stage P0 and the preparation stage R1. The method for obtaining the delay time between two pipeline stages may include but is not limited to the above examples, and this embodiment does not limit this.

[0079] Step 104: Generate an edge between the second node and the first node of the historical pipeline stage, and use the delay time as the weight of the edge.

[0080] The source node of the edge is the second node, and the destination node is the first node. The second node is generated during the execution of the historical pipeline stage.

[0081] In this embodiment, after the CPU generates the first node of the current pipeline stage and obtains the delay time between the historical pipeline stage associated with the current pipeline stage and the current pipeline stage, it can generate an edge between the first node of the current pipeline stage and the second node of the historical pipeline stage, and the delay time can be used as the weight of the edge. In combination with the above example, after generating node R1, edge 401 between node D1 and node R1 can be generated, as well as edge 3041 between node P0 and node R1. The destination nodes of edge 3041 and edge 401 are the first node R1, the source node of edge 3041 is the second node P0, and the source node of edge 401 is the second node D1.

[0082] At the same time, the delay time between the historical pipeline stage and the current pipeline stage can be used as the weight of the edge. In combination with the above example, the delay time between the historical pipeline stage D1 and the current pipeline stage R1 can be used as the weight of edge 401, and the delay time between the historical pipeline stage P0 and the current pipeline stage R1 can be used as the weight of edge 3041. Among them, the weight of the edge can be determined by the delay time between the historical pipeline stage and the current pipeline stage, and can also be determined by other factors between the historical pipeline stage and the current pipeline stage, and this embodiment does not limit this.

[0083] In one embodiment, since the execution of the next pipeline stage in the same instruction must depend on the execution result of the previous pipeline stage, the CPU can generate a node for each pipeline stage in turn during the execution of an instruction, and for two adjacent pipeline stages in the instruction, generate an edge between the two pipeline stages, that is, an edge between two nodes corresponding to the two pipeline stages. For example, while generating 5 nodes corresponding to instruction i0, 4 edges can be generated, that is, Figure 3 As shown in the edge 301, edge 302, edge 303 and edge 304, there is an edge between two adjacent nodes, and the source node connected by each edge is the node of the previous pipeline stage, and the destination node is the node of the next pipeline stage. For example, the source node connected by edge 301 is D0, and the destination node is R0.

[0084] Step 105 : Store the generated nodes of each pipeline stage, as well as the edges and the weights of the edges, to obtain graph data, so as to analyze the performance of the processor based on the critical path between the start node and the end node in the graph data.

[0085] The start node is the node of the first pipeline stage in the first instruction, and the end node is the node of the last pipeline stage in the last instruction. Figure 3 The starting node in is the node of the dispatch phase D0, and the ending node is the node of the submission phase C4.

[0086] In this embodiment, after generating the nodes, edges and edge weights of each pipeline stage, the nodes, edges and edge weights of each pipeline stage can be stored to obtain graph data. Figure 2 The reorder cache 2011 shown in the figure executes steps 101 to 105, and stores the obtained nodes of each pipeline stage, as well as the edges and the weights of the edges into the instruction information cache table 2012. Optionally, in the generated graph data, the node numbers of all nodes can be Figure 3The encoding is performed in the manner shown in the figure. In each instruction, the encoding starts from the node of the first pipeline stage and ends at the node of the last pipeline stage. The nodes of the same pipeline stage of different instructions use different node numbers. Alternatively, combining the execution order of the instructions and the pipeline order of multiple pipeline stages, all nodes are encoded starting from the node of the first pipeline stage in the first instruction in a descending or descending order.

[0087] After the graph data is stored, the performance of the processor can be analyzed based on the critical path between the starting node and the ending node in the graph data. The critical path is the path with the longest path distance among all paths from the starting node to the ending node in the graph data. Among them, the path starts from the starting node in the graph data and gradually reaches the ending node along the outgoing edge of the node. The path distance can be calculated by the weight of the edge. The shortest path is the path with the smallest sum of the weights of all the edges in the path among all the paths; the longest path is the path with the largest sum of the weights of all the edges in the path among all the paths.

[0088] Exemplarily, a single source shortest path (SSSP) algorithm, such as the Dijkstra algorithm, can be used to determine the path with the largest weight from the start node to the end node as the critical path. At this time, the path distance of the critical path, that is, the path distance from the start node to the end node corresponds to the longest time required to execute all instructions in the benchmark program. The longest time is related to the microstructure of the CPU and the program behavior of the benchmark program. The microstructure of the CPU can be adjusted based on the critical path. For example, if the longest path is Figure 3 The thick solid line shown is from the dispatch stage D0 to the completion stage P0, from the completion stage P0 to the preparation stage R1, from the preparation stage R1 to the completion stage P1, from the completion stage P1 to the preparation stage R3, from the preparation stage R3 to the completion stage P3, from the completion stage P3 to the dispatch stage D4, and the dispatch stage D4 ends at the submission stage C4. Figure 2 The host shown Figure 3 The graph structure shown in the figure is displayed, and the longest path is highlighted. Users can determine the reason for the longest path, that is, the reason for the CPU performance degradation, based on the displayed graph data. For example, if it is determined that the path distance is large due to the large delay time between the completion stage P1 and the preparation stage R3, the CPU microstructure can be adjusted to delete the association between the completion stage P1 and the preparation stage R3, or to shorten the delay time between the completion stage P1 and the preparation stage R3. After the adjustment, the longest path can be Figure 3The dotted line in the figure indicates the path. After the adjustment, the prototype system can be restarted, and steps 101 to 105 can be repeated to obtain new graph data, and then the analysis can be continued to adjust the microstructure of the CPU. After multiple iterations, the CPU performance can be optimized. The specific method for analyzing the performance of the CPU based on the critical path in the graph data may include but is not limited to the above examples, and this embodiment does not limit this.

[0089] In summary, in the embodiment of the present invention, according to the execution order of multiple instructions, a pipeline is used to execute multiple instructions. During the execution of each instruction, the first node of the current pipeline stage is generated. When the execution of the current pipeline stage depends on the execution result of the historical pipeline stage that has ended, the delay time between the current pipeline stage and the historical pipeline stage is obtained, and the edge between the second node and the first node of the historical pipeline stage is generated, and the delay time is used as the weight of the edge. The generated nodes of each pipeline stage, as well as the edges and the weights of the edges are stored to obtain graph data, so as to analyze the performance of the processor based on the critical path between the starting node and the ending node in the graph data. During the running process of the benchmark test program, the nodes of each pipeline stage and the edges between two pipeline stages with a dependency relationship are generated, and the delay time between the two pipeline stages is used as the weight of the edge, so as to construct the graph data. By using the critical path in the graph data to analyze the performance of the processor, the factors that cause the processor to reduce performance can be quickly determined, and the microstructure of the processor can be adjusted to improve the analysis efficiency.

[0090] At the same time, since the graph data is built based on all pipeline stages of all instructions in the entire benchmark program, it can systematically and completely present the process of the processor running the benchmark program, which is convenient for systematic analysis of the processor's performance. In addition, during the running of the benchmark program, there is no need to adjust the benchmark program, so the processor can run the complete benchmark program and achieve accurate analysis of the processor's performance.

[0091] Optionally, step 105 may include:

[0092] Divide the nodes of all pipeline stages in each instruction into a node set;

[0093] The node set and the first type of edges connecting the nodes in the node set are stored as a subgraph, and the second type of edges connecting the nodes in different sets are stored, so as to obtain graph data consisting of the subgraph and the second type of edges.

[0094] In this embodiment, all nodes corresponding to each instruction can be divided into a node set according to multiple instructions in the benchmark test program, and the generated edges can be divided into first-class edges and second-class edges. The first-class edges connect nodes in the node set, and the second-class edges connect nodes in different node sets. The nodes of each node set and the edges connecting the nodes in the node set form a subgraph. Figure 3 As shown, the nodes D0, R0, E0, P0 and C0 corresponding to instruction i0 form a node set, and the first-class edges connecting the node set include edges 301, 302, 303 and 304. The nodes D0, R0, E0, P0 and C0, as well as the first-class edges 301, 302, 303 and 304 can be divided into a subgraph 300. The nodes D1, R1, E1, P1 and C1 corresponding to instruction i1 form another node set, and the node set corresponding to instruction i1 and the first-class edges connecting the nodes in the node set can be divided into another subgraph 400. The second-class edges connecting the nodes in subgraph 300 and the nodes in subgraph 400 include edges 3041, 3042 and 3043. Similarly, a subgraph corresponding to each instruction and all the second-class edges can be obtained.

[0095] In the embodiment of the present invention, the nodes in the graph data are in subgraph units, and each time a subgraph node and a first-class edge are added, Figure 3 The graph data shown can be called equal-increment graph data. In the process of storing graph data, nodes and first-class edges in each subgraph are stored in units of subgraphs, and all second-class edges in the entire graph data can be stored at the same time. Dividing the graph data into multiple subgraphs and second-class edges with the same graph structure can facilitate compressed storage of the entire graph data according to the graph structures of the multiple subgraphs.

[0096] Optionally, the step of storing the node set and the first type of edges connecting the nodes in the node set as a subgraph, and storing the second type of edges connecting the nodes in different sets, and obtaining the graph data consisting of the subgraph and the second type of edges may include:

[0097] Create and store a graph data header, which includes subgraph structure information;

[0098] Add the connection information of the second type of edge in the graph data header, where the connection information is used to describe the source node and the destination node of the second type of edge;

[0099] In the order of execution, a unique subgraph number is sequentially set for the subgraph corresponding to each instruction and stored, and when the subgraph includes a key node, a first information set of the first type of edge in the subgraph is correspondingly stored, and a second information set of the key node is stored;

[0100] Among them, the first information set includes the weight of the first type of edge; the key node is the destination node of the target second type edge; the second information set includes the subgraph number of the subgraph where the source node of the target second type edge is located and the weight of the target second type edge, as well as the index of the connection information of the target second type edge in the graph data header.

[0101] In this embodiment, a reference node number corresponding to the node number of the starting node can be generated according to the node number of each starting node included in the starting subgraph, and a graph data header can be established according to the connection information and the reference node number of each starting first-class edge included in the starting subgraph. The starting subgraph is the first subgraph of multiple subgraphs, for example Figure 3 The subgraph 300 corresponding to the first instruction i0 shown in the figure, the starting nodes are nodes D0, R0, E0, P0 and C0 in the starting subgraph, and the starting first-class edges are first-class edges 301, 302, 303 and 304 in the starting subgraph. Multiple subgraphs have the same graph structure, and the same reference node number can be used in the graph data header to represent nodes at the same position in different subgraphs. For example, Figure 3 Nodes D0, D1, D2, D3 and D4 in the graph can be represented by reference node number D, and nodes R0, R1, R2, R3 and R4 can be represented by reference node number R. Multiple reference node numbers can be used to represent nodes at the same position in multiple subgraphs to describe the node sets in the multiple subgraphs. At the same time, the connection relationship between multiple reference node numbers can be used to describe the set of first-class edges in the multiple subgraphs.

[0102] In one embodiment, the graph data header may be composed of a number array and a connection information list, wherein the number array is used to sequentially store multiple reference node numbers, and the connection information list is used to store first connection information of a first type of edge and second connection information of a second type of edge. The first connection information is used to describe the reference node numbers of the source node and the destination node of the first type of edge connection, and the second connection information is used to describe the reference node numbers of the source node and the destination node of the second type of edge connection. The number array is, for example, a one-dimensional array N[5]={D, R, E, P, C}, wherein the first element in the number array stores the reference node number D of the first node in the subgraph, and the second element stores the reference node number R of the second node in the subgraph, and so on and so forth, and the reference node number corresponding to each node in the subgraph may be stored to describe the node set of multiple subgraphs. The node number of each node in the starting subgraph may be directly used as the corresponding reference node number, or a unique number may be set for each node in the starting subgraph as the reference node number. Multiple reference node numbers are sequentially set in the number array, so the node corresponding to the reference node number may be determined according to the position of the reference node number in the number array. For example, if the reference node number E is in the third position in the number array, then the reference node number E can be determined as the reference node number of the third node in the subgraph; if the reference node number P is in the fourth position in the number array, then the reference node number P can be determined as the reference node number of the fourth node in the subgraph.

[0103] As shown in Table 1, Table 1 is an exemplary connection information list.

[0104] index 1 2 3 4 5 Connection Information DR RE EP PC DD

[0105] Table 1

[0106] In Table 1, the first connection information and the second connection information are connection information of the same format. Each connection information is used to describe the source node and the destination node of an edge. The first bit of the connection information is the reference node number of the source node of the edge connection, and the second bit is the reference node number of the destination node of the edge connection. For example, the connection information (DR) is the first connection information, the first bit D indicates that the reference node number of the source node of the edge connection is D, and the second bit R indicates that the reference node number of the destination node of the edge connection is R. The index of the connection information is used to obtain the connection information from the connection information list. For example, when the index of a first-class edge is 1, the connection information (DR) can be determined from Table 1 according to the index 1, so that it can be determined that the source node of the first-class edge connection is the first node in the subgraph where it is located, and the destination node is the second node in the subgraph where it is located. Further, according to the subgraph number of the subgraph where the first-class edge is located, the source node and the destination node of the first-class edge connection can be determined. For example, if the subgraph number of the subgraph where the first-class edge is located is 1, it can be determined that the source node of the first-class edge connection is D1 and the destination node is R1.

[0107] Among them, when the connection information is the second connection information, it is necessary to determine the subgraph number of the source subgraph where the source node of the second type edge connection is located, and the subgraph number of the destination subgraph where the destination node is located. For example, when the index of a second type edge is 5, the second connection information (DD) can be determined from Table 1. Further, if it is determined that the subgraph number of the source subgraph where the source node of the second type edge connection is located is 1, and the subgraph number of the destination subgraph where the destination node is located is 2, then it can be determined that the source node of the second type edge connection is D1 and the destination node is D2.

[0108] Optionally, according to the connection information and reference node number of each first-type edge included in the starting subgraph, the step of establishing the graph data header may include:

[0109] Get multiple reference node numbers; each reference node number corresponds to a node at the same position in multiple subgraphs;

[0110] Establish a coding matrix; the row numbers and column numbers with the same values ​​in the coding matrix correspond to the same reference node number;

[0111] Determine the matrix position corresponding to the first-class edge from the encoding matrix, and add the index corresponding to the first-class edge in the matrix position; the row number of the matrix position is the reference node number corresponding to the source node of the first-class edge, and the column number of the matrix position is the reference node number corresponding to the destination node of the first-class edge.

[0112] In one embodiment, the graph data header may be established in the form of a coding matrix. Figure 4 As shown, Figure 4 A schematic diagram of a graph data header in an embodiment of the present invention is shown, where the encoding matrix is ​​a 5-row 5-column encoding matrix, the row numbers from the 1st row to the 5th row are reference node numbers D, R, E, P and C, respectively, and the column numbers from the 1st column to the 5th column are reference node numbers D, R, E, P and C, respectively. The row numbers and column numbers with the same values ​​are encoded by the same reference node number, for example, the row number of the 3rd row and the column number of the 3rd column are both reference node number E. The row numbers of the encoding matrix correspond to the source nodes connected by the first type of edge and the second type of edge, and the column numbers correspond to the destination nodes connected by the first type of edge and the second type of edge. For example, the Pth row indicates that the reference node number of the source node connected by the first type of edge or the second type of edge is P, and the Rth column indicates that the reference node number of the destination node connected by the first type of edge or the second type of edge is R.

[0113] After the coding matrix is ​​established, the matrix position of the starting first-class edge can be determined from the coding matrix according to the reference node numbers of the source node and the destination node connected by the starting first-class edge, and a unique index is added to the corresponding position. For example, the source node connected by the starting first-class edge 302 is R0, the destination node is E0, the reference node number of the source node is R, and the reference node number of the destination node is E. Then the matrix position of the starting first-class edge 302 is in the Rth row and the Eth column. A unique index "3" can be set in the Rth row and the Eth column in the coding matrix. The index "3" corresponds to the first-class edge between the node R0 and the node E0, and also corresponds to the first-class edge between the node R1 and the node E1 at the same position, as well as the first-class edge between the node R2 and the node E2, the first-class edge between the node R3 and the node E3, and the first-class edge between the node R4 and the node E4. Among them, the above-mentioned reference node number can also be replaced by Arabic numerals 1, 2, 3, 4 and 5. The specific type of the reference node number can be set according to the requirements, and this embodiment does not limit this.

[0114] In an embodiment of the present invention, when a graph data header is established in the form of a coding matrix, the source nodes and destination nodes corresponding to the first connection information and the second connection information can be stored through the row and column relationship of the coding matrix, so that a graph data header with a smaller data volume can be obtained, thereby reducing the storage space required to store the graph data.

[0115] In another embodiment, the starting subgraph can be directly used as the graph data header. Since the graph structure of the starting subgraph is the same as that of other subgraphs, the graph structure of all other subgraphs can be determined based on the structure of the starting subgraph. The node set of the starting subgraph and the set of first-class edges can be directly used as the graph data header. When adding the connection information of the second-class edges to the graph data header, the connection information of the second-class edges can be stored in adjacent positions of the starting subgraph, and the connection information of the second-class edges can only include the number of the source node and the number of the destination node of the second-class edges. Among them, the format of the graph data header may include but is not limited to the above examples, and the graph data header may be flexibly designed according to the first-class edges and the second-class edges in the subgraph. The specific form of the graph data header may include but is not limited to the above examples.

[0116] In this embodiment, for the second type of edge, it can be first determined whether the second connection information of the second type of edge has been stored in the graph data header. If the second connection information of the second type of edge is not stored, the second connection information is added to the graph data header. Combined with the above example, if for the subgraph 200, in the process of executing instruction i1, a node R1 corresponding to the preparation stage R1 can be generated, and a second type of edge 3041 can be generated. The source node connected to the second type of edge 3041 is P0, and the destination node is R1. At this time, it is possible to find out whether the second connection information (PR) of the second type of edge 3041 is included in the graph data header. For example, if the graph data header consists of a number array and a connection information list, it can be determined that the second connection information (PR) corresponding to the first type of edge 3021 is not included in Table 1. An item of connection information (PR) can be added to the connection information list shown in Table 1, and a unique index "6" can be set to obtain a connection information list as shown in Table 2.

[0117] index 1 2 3 4 5 6 Connection Information DR RE EP PC DD PR

[0118] Table 2

[0119] Similarly, if the image data header is a coding matrix, when it is detected Figure 3 When the Pth row and the Rth column in the coding matrix shown are empty, a unique index 6 may be set at the position of the Pth row and the Rth column. On the contrary, if the second connection information (PR) already exists in the connection information list or the coding matrix, the second connection information (PR) may not be added to the connection information list or the coding matrix.

[0120] In this embodiment, a sub-graph number that uniquely identifies the starting sub-graph may be set, and the sub-graph number of the starting sub-graph may be stored. For other sub-graphs after the starting sub-graph, a unique sub-graph number may be set for the sub-graph. The sub-graph numbers of the multiple sub-graphs may be encoded in sequence to set a unique sub-graph code for each sub-graph. Figure 5 As shown, Figure 5 A schematic diagram of graph data storage in an embodiment of the present invention is shown, where subgraph number 501 is subgraph number 0 of subgraph 300, subgraph number 502 is subgraph number 1 of subgraph 400, and the subgraph numbers of the plurality of subgraphs are encoded in sequence according to the encoding order. The information set includes a first information set of first-type edges and a second information set of second-type edges. Figure 4As shown, the subgraph 300 includes the first type edge 301, the first type edge 302, the first type edge 303 and the first type edge 304, and the first storage area 503 corresponding to the subgraph number 501 sequentially stores the first information set 5031 of the first type edge 301, the first information set 5032 of the first type edge 302, the first information set 5033 of the first type edge 303, and the first information set 5034 of the first type edge 304. The information "1" included in the first information set 5031 indicates that the weight of the first type edge 301 is 1, and the information "0" included in the first information set 5032 indicates that the weight of the first type edge 302 is 0. When reading the subgraph 300 from the graph data, the reference node number can be first obtained from the graph data header, and then each node in the subgraph 300 can be determined based on the reference node number and the subgraph number 0, and the four first information sets in the first storage area corresponding to the subgraph number 0 can be obtained in sequence based on the subgraph number 0, and the first type of edge corresponding to the first information set can be determined based on the position of each first information set. For example, the weight in the first information set 5032 stored in the second position of the first storage area is the weight of the second first type edge 302 in the subgraph 300, and the weight in the first information set 5033 stored in the third position of the first storage area is the weight of the third first type edge 303 in the subgraph 300. Among them, the first information set can also include other information of the first type of edge, which is not limited in this embodiment.

[0121] In one embodiment, the first information set may also include the index of the first type of edge in the graph data header. In combination with the above example, an information may be added to the first information set 5031, which is the index 2 of the first type of edge 301. After setting the index in the first information set, the first connection information of the first type of edge can be obtained from the graph data header according to the index in the first information set, and the reference node numbers of the source node and the destination node of the first type of edge connection can be determined. Further, the destination node and the source node of the first type of edge connection can be determined according to the subgraph number corresponding to the first information set. For example, if the first information set of a first type of edge includes index 2, the first connection information (DR) can be determined from the graph data header according to index 2, and further, according to the subgraph number 0 corresponding to the first information set, it can be determined that the source node of the first type of edge connection is the first node D0 in the subgraph 300, and the destination node is the second node R0 in the subgraph 300.

[0122] Similarly, in the first storage area corresponding to subgraph number 502, the weight of each first-class edge in subgraph 400 is stored in sequence. In the second storage area 504 corresponding to subgraph number 502, a second information set 5041 of second-class edge 3041, a second information set 5042 of second-class edge 3042, and a second information set 5043 of second-class edge 3043 in subgraph 400 are stored in sequence. The first information in the second information set of each second-class edge is the subgraph number of the source subgraph connected by the second-class edge, the second information indicates the index of the second-class edge in the graph data header, and the third information is the weight of the second-class edge. For example, the first information 0 in the second information set 5041 of second-class edge 3041 indicates that the source node connected by the second-class edge 3041 is in subgraph 300 with subgraph number 0, and index 6 corresponds to Figure 3 In the Pth row and Rth column of the coding matrix shown in FIG. 1 , in the process of reading the graph data, firstly, according to the second information (i.e., index 6) in the second information set 5041, Figure 4 The second connection information PR is determined in the encoding matrix shown in FIG. 1 or the connection information list shown in Table 1. Further, according to the first information (i.e., subgraph number 0), the source node of the second type of edge can be determined to be P0, and according to subgraph number 1, the destination node of the second type of edge can be determined to be P1, and the weight of the second type of edge 3041 can be determined to be 10. Other information of the second type of edge can also be stored in the second information set, and this embodiment does not limit this.

[0123] When storing equal-incremental graph data in the manner described above, the structural information of multiple subgraphs can be stored through the graph data header, and the same structural information can be stored for multiple subgraphs of the same structure. At the same time, an information set can be stored based on the subgraph number, and the information set includes the weight information of the first type of edges in each subgraph, and the weight information of the second type of edges and other difference information between multiple subgraphs. Storing a copy of structural information for each subgraph can be avoided, and the amount of graph data can be reduced, thereby reducing storage space. It should be noted that other methods can also be used to store the equal-incremental graph data as described above, and this embodiment does not limit this.

[0124] Optionally, after the graph data is stored in the above manner, the method may further include:

[0125] Selecting nodes in at least one subgraph from the plurality of subgraphs in sequence according to the coding order of the nodes as nodes to be calculated;

[0126] In the case where the node to be calculated has at least one incoming edge, the path distance to be selected is determined according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge; the target path distance of the target source node is the longest path distance and / or the shortest path distance among all path distances between the starting node and the source node;

[0127] Determine the target path distance of the node to be calculated from all the path distances to be selected, and use the source node corresponding to the target path distance of the node to be calculated as the previous node of the node to be calculated; the target path distance of the node to be calculated is the maximum path distance to be selected corresponding to the longest path distance and / or the minimum path distance to be selected corresponding to the shortest path distance;

[0128] The steps of selecting a node to be calculated, determining the distance of a path to be selected, and determining the target path distance are repeated until the target path distances and previous nodes of all nodes are determined.

[0129] In one embodiment, the Figure 2 The preprocessing device shown performs the steps of selecting a node to be calculated, determining a distance of a selected path, and determining a target path distance until the target path distances and previous nodes of all nodes are determined. The preprocessing device can obtain a node in at least one subgraph as a node to be calculated from the instruction information cache table, and obtain the first-class edge and the second-class edge, the weight of the first-class edge and the weight of the second-class edge from the instruction information cache table. Figure 3 In each subgraph shown, there is a first-class edge between two adjacent nodes, and the source node of the first-class edge is located before the target node in the coding order. For example, the source node of the first-class edge 302 is node R0, and the target node is E0. At the same time, the source node of the second-class edge is located before the target node in the coding order, for example, the source node of the second-class edge 3041 is node P0, and the destination node is node R1. In the coding order, node P0 is located before node R1. The preprocessing device can start from the starting node, and select one or more nodes in the subgraphs from all subgraphs as nodes to be calculated each time according to the coding order of the nodes, and the number of subgraphs selected is a positive integer. When selecting according to the coding order of the nodes, the selection can be made according to the coding order of multiple subgraphs, that is, starting from the first subgraph, at least one node in the subgraph is selected as the node to be calculated each time. Alternatively, directly according to the coding order of all nodes, starting from the starting node, at least one node in the subgraph is selected as the node to be calculated each time.

[0130] Exemplarily, each time a node in a subgraph can be selected as a node to be calculated, the first time the node in the subgraph corresponding to instruction i0 is selected as the node to be calculated, the second time the node in the subgraph corresponding to instruction i1 is selected as the node to be calculated, and so on, the nodes in the subgraph corresponding to all instructions can be selected as nodes to be calculated. Alternatively, each time two nodes in subgraphs can be selected as nodes to be calculated, the first time the nodes in the two subgraphs corresponding to instruction i0 and instruction i1 are selected as nodes to be calculated, the second time the nodes in the two subgraphs corresponding to instruction i0 and instruction i1 are selected as nodes to be calculated, and so on, the nodes in the subgraphs corresponding to all instructions can be selected as nodes to be calculated. The number of subgraphs selected each time can be the same or different, and this embodiment does not limit this.

[0131] In this embodiment, after selecting the node to be calculated, the incoming edge of the node to be calculated can be determined from the graph data, and the incoming edge is the edge connected when the node is used as the destination vertex. Figure 3 As shown, if the node to be calculated is node R1, the incoming edges of node R1 to be calculated include first-class edges 401 and second-class edges 3041. If the node to be calculated is node E0, the incoming edges of node E0 to be calculated only include first-class edges 302. When equal-increment graph data is stored in a unit compression structure format, the nodes to be calculated and the incoming edges can be determined from the graph data based on the subgraph number. For example, when selecting a node in subgraph 400 as a node to be calculated, first, multiple reference node numbers D, R, E, P, and C can be obtained from the graph data header, and then, combined with the subgraph number 1 of subgraph 400, a subscript is added to each reference node number to obtain each node number D1, R1, E1, P1, and C1 in subgraph 400, and the node to be calculated in subgraph 400 is determined. At the same time, since multiple nodes in subgraph 400 are connected sequentially, multiple first-class edges in subgraph 400 can be determined, that is, Figure 1 The first first-class edge D1-R1, the second first-class edge R1-E1, the third first-class edge E1-P1 and the fourth first-class edge P1-C1 are shown. In addition, the first information set 505, the first information set 506, the first information set 507 and the first information set 508 can be obtained from the first storage area corresponding to the subgraph number 1. The multiple first information sets are sequentially stored in the first storage area, and the weight of the first first-class edge D1-R1 can be determined to be the weight 1 in the first information set 505. Similarly, the weight of each first-class edge in the subgraph 400 can be determined.

[0132] Furthermore, multiple second information sets can be obtained from the second storage area corresponding to the subgraph number 1. Taking the second information set 5041 as an example, the second information 6 (the information is an index) can be extracted from the second information set 5041, and the connection information (PR) can be obtained from the graph data header according to the index. Then, according to the first information 0 in the second information set 5041 (the information is the subgraph number of the source subgraph), the source node of the second type of edge can be determined to be node P0, and according to the subgraph number 1, the destination node of the second type of edge can be determined to be node R1, so that the second type of edge can be determined to be P0-R1, that is, Figure 3 . In addition, the third information 10 can be obtained from the second information set 5041, and the third information 10 is the weight of the second edge 3041. Similarly, the weights of the second edge 3042 and the second edge 3043 corresponding to the subgraph 400 can be obtained. At this point, all nodes and first edges in the subgraph 400 have been determined, and all second edges corresponding to the subgraph 400 have been determined, and the incoming edges of each node in the subgraph 400 can be further determined. For example, the incoming edges of the node D1 include the second edge 3042, and the incoming edges of the node R1 include the first edge 401 and the second edge 3041, and the weight of each incoming edge can also be determined.

[0133] It should be noted that when equal-increment graph data is stored in other formats, one or more nodes in the subgraph can be selected from the graph data by a method that matches the storage format, and the incoming edges of the nodes and the weights of the incoming edges can be determined.

[0134] In this embodiment, after determining the node to be calculated, it can be determined whether the node to be calculated has an incoming edge, so as to determine the candidate path distance of the node to be calculated according to the weight of the incoming edge and the target path distance of the target source node connected by the incoming edge, and then select and determine the target path distance from all the candidate path distances. When it is necessary to determine the shortest path between the end node and the start node, the target path distance is the minimum candidate path distance; when it is necessary to determine the longest path between the end node and the start node, the target path distance is the maximum candidate path distance. Among them, after determining the target path distance of a node, the obtained target path distance can be stored in a preset position, so that when determining the candidate path distances of other nodes, if the node is determined to be the source node of the incoming edge, the target path distance of the node can be obtained from the preset position.

[0135] Optionally, the step of determining the distance of the to-be-selected path according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge may include:

[0136] If the target source node is the starting node, the weight of the incoming edge is used as the distance of the path to be selected;

[0137] If the target path distance with a target source node is predetermined, the weight of the incoming edge and the target path distance of the target source node are summed to obtain the distance of the selected path;

[0138] If the target path distance of the target source node is not determined in advance, after the target path distance of the target source node is determined, the summing step is performed to obtain the candidate path distance.

[0139] Exemplarily, according to the coding order of multiple nodes, the node in the first subgraph of the multiple subgraphs can be selected as the node to be calculated. Since the starting node of the multiple nodes has no incoming edge and no previous node, the target path distance of the starting node can be set to 0, and the previous node is empty. When the target source node connected by the incoming edge is determined to be the starting node, the weight of the incoming edge can be used as the path distance to be selected. In combination with the above example, all nodes in the subgraph 300 can be determined as nodes to be calculated. At this time, the incoming edge of each node in the subgraph 300 and the weight of the incoming edge can be determined. When calculating the target path distance of node R0, since the incoming edge of node R0 only has the first type of edge 301, the path distance to be selected of node R0 is the weight of the first type of edge 301, and since node R0 has only one incoming edge, it has only one path distance to be selected. The target path distance of node R0 can be determined as the weight of the first type of edge 301, and the target path distance of node R0 can be stored in a preset position.

[0140] Further, when determining the candidate path distance of node E0, the incoming edge can be determined to be the first type edge 302, and the target source node of the incoming edge 302 is node R0. At this time, the target path distance of node R0 can be read from the preset position, and the target path distance of node R0 and the weight of the incoming edge 302 are summed to obtain a candidate path distance of node E0. Similarly, since node E0 has only one incoming edge, the calculated candidate path distance can be directly used as the target path distance of node E0, and the target path distance of node E0 is stored in the preset position. By analogy, the target path distance of each node in the subgraph 300 can be determined. Among them, the target path distance of each node can be stored in different positions to store the target path distances of all nodes that have been determined.

[0141] Among them, since nodes in multiple subgraphs can be selected as nodes to be calculated each time, when determining the path distance to be calculated of a certain node, there may be a situation where the target path distance of the source node connected by the incoming edge has not been determined. In combination with the above example, when the nodes in subgraph 300 and subgraph 400 are determined as nodes to be calculated at the same time, multiple nodes can be carried out in parallel, and the path distance to be calculated can be determined at the same time. When determining the path distance to be calculated of node R1, the incoming edge of node R1 includes the second type edge 3041, and the source node of the second type edge 3041 is node P0, and the target path distance of node P0 needs to be obtained. At this time, the target path distance of node P0 may not be determined. In this case, after waiting for the target path distance of node P0 to be determined, the weight of the second type edge 3041 and the target path distance of node P0 can be summed to obtain a path distance to be calculated for node R1. On the contrary, when the target path distance of node P0 has been determined, the target path distance of node P0 and the weight of the second type edge 3041 can be directly summed to obtain a path distance to be calculated for node R1. When determining the candidate path distance of node R1, the target path distance of node P0 can be continuously read from the preset position. If the target path distance of node P0 is not read, wait for a preset time until the target path distance of node P0 is read, and then sum the weight of the second-class incoming edge 3041 and the target path distance of node P0 to obtain a candidate path distance of node R1. The specific value of the preset time can be set according to demand, and this embodiment does not limit this.

[0142] In this embodiment, for a node to be calculated having multiple incoming edges, each incoming edge of the node to be calculated can be determined, and the source node of each incoming edge and the target path distance of the source node can be determined. Taking node R1 as an example, when node R1 is used as a node to be calculated, it can be determined that node R1 has two incoming edges, namely, a first-type edge 401 and a second-type edge 3041. At this time, it can be determined that the target source node of the first-type edge 401 is node D1, and the target source node of the second-type edge 3041 is node P0. The predetermined target path distance of node D1 and the target path distance of node P0 can be read from a preset position, and the target path distance of node D1 and the weight of the first-type edge 401 can be summed to obtain a path distance to be selected for node R1, and the target path distance of node P0 and the weight of the second-type edge 3041 can be summed to obtain another path distance to be selected for node R1.

[0143] When it is determined that the node to be calculated has a candidate path distance, the candidate path distance can be directly used as the target path distance of the node to be calculated. When it is determined that the node to be calculated has multiple candidate path distances, the maximum and / or minimum candidate path distances can be selected as the target path distance. In combination with the above example, if it is necessary to determine the longest path between the end node and the start node, the maximum candidate path distance can be selected from the two candidate path distances of node R1 as the target path distance of node R1, and the source node corresponding to the maximum candidate path distance can be used as the previous node of node R1. For example, if the candidate path distance calculated based on the target path distance of node D1 and the weight of the first type edge 401 is the largest, node D1 can be used as the previous node of node R1. On the contrary, if it is necessary to determine the shortest path between the end node and the start node, the minimum candidate path distance can be selected from the two candidate path distances as the target path distance of node R1, and the source node corresponding to the minimum candidate path distance can be used as the previous node of node R1.

[0144] In one embodiment, it may be necessary to simultaneously determine the longest path and the shortest path between the end node and the start node. In this case, the largest candidate path distance can be used as a target path distance of node R1, and the smallest candidate path distance can be used as another target path distance of node R1. In addition, the node corresponding to the largest candidate path distance can be used as the previous node corresponding to the longest path, and the node corresponding to the smallest candidate path distance can be used as the previous node corresponding to the shortest path. In this case, when determining the candidate path distance of node E1, the incoming edge of node E1 only has the first-class edge between node E1 and node R1, and the source node of the incoming edge is node R1. Two target path distances can be selected, and the two candidate path distances corresponding to node E1 are calculated respectively. The candidate path distance with the largest of the two candidate path distances is used as the target path distance corresponding to the longest path, and the smallest candidate path distance is used as the target path distance corresponding to the shortest path.

[0145] If a node has multiple candidate path distances, and the multiple candidate path distances are equal, the multiple candidate path distances can be used as the target path distances at the same time, and the source nodes corresponding to the multiple candidate path distances can be saved as the previous nodes at the same time. For example, if the two candidate path distances of node R1 are the same, node D1 and node P0 can be saved as the previous nodes of node R1 at the same time.

[0146] Optionally, the method may further include:

[0147] When the node to be calculated is a node other than the starting node and has no incoming edge, the target path distance between the node to be calculated and the starting node is set to a first preset value, and the previous node of the node to be calculated is set to a second preset value; the first preset value is an infinitesimal value corresponding to the longest path distance value and / or an infinite value corresponding to the shortest path distance value; the second preset value indicates that the previous node of the node to be calculated is empty.

[0148] In one embodiment, among all nodes except the starting node, there may be nodes without incoming edges, for example, there may be no second-type edge 3042 between node D0 and node D1. In this case, when node D1 is determined as the node to be calculated, since node D1 has no incoming edges, the target path distance between node D1 and the starting node D0 cannot be determined. At this time, the target path distance between node D1 and the starting node D0 can be set to a first preset value, and the previous node of node D1 can be set to a second preset value. When it is necessary to determine the longest path between the starting node and the ending node, the first preset value can be an infinitesimal value; when it is necessary to determine the shortest path between the starting node and the ending node, the first preset value can be an infinite value; when it is necessary to simultaneously determine the shortest path and the longest path between the starting node and the ending node, the first preset value can include both an infinite value and an infinitesimal value. The infinitesimal value and the infinite value are used to assist in determining the target path distance of other nodes. For example, when determining the distance of the candidate path of node R1, when it is necessary to determine the longest path between the starting node and the ending node, the infinitesimal value of node D1 and the weight of the first type edge 401 can be summed to obtain the distance of the candidate path of node D1, which is also an infinitesimal value; when it is necessary to determine the shortest path between the starting node and the ending node, the infinite value of node D1 and the weight of the first type edge 401 can be summed to obtain the distance of the candidate path of node D1, which is also an infinitesimal value. When the target path distance of a certain node is an infinite value and an infinitesimal value, it means that there is no path between the node and the starting node. When the target path distance of a certain node is an infinite value or an infinitesimal value, the previous node of the node can be set to a second preset value, indicating that the node has no previous node, that is, it means that there is no path between the node and the starting node.

[0149] In this embodiment, the steps of selecting the node to be calculated, determining the distance of the path to be selected, and determining the target path distance can be repeatedly performed until the target path distance and the previous node of all nodes are determined. At this time, the longest path and / or the shortest path between the starting node and the ending node can be determined based on all the previous nodes. In combination with the above example, if the longest path needs to be determined, the target path distance of each node is the maximum path distance to be selected, and the previous node of each node is the source node corresponding to the maximum path distance to be selected. At this time, the path composed of all the previous nodes is the longest path from the starting node to the ending node, and the target path distance of the ending node is the longest path distance between the starting node and the ending node. If the shortest path needs to be determined, the target path distance of each node is the minimum path distance to be selected, and the previous node of each node is the source node corresponding to the minimum path distance to be selected. At this time, the path composed of all the previous nodes is the shortest path from the starting node to the ending node, and the target path distance of the ending node is the shortest path distance between the starting node and the ending node. If it is necessary to determine the longest path and the shortest path at the same time, then for the longest path, the path composed of the previous nodes corresponding to the maximum candidate path distance of each node is the longest path, and the maximum target path distance of the end node is the longest path distance between the start node and the end node; for the shortest path, the path composed of the previous nodes corresponding to the minimum candidate path distance of each node is the shortest path, and the minimum target path distance of the end node is the shortest path distance between the start node and the end node.

[0150] In the embodiment of the present invention, for equal increment graph data, the shortest or longest path distance between a node and a starting node is determined by the weight of the incoming edge, and the path distance of the node does not need to be updated multiple times, which is highly efficient. In addition, since the path distance between the node and the starting node is calculated using the incoming edge of the node, the nodes in multiple subgraphs can be processed in parallel, which can further improve the processing efficiency.

[0151] Optionally, the step of determining the distance of the to-be-selected path according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge may include:

[0152] The step of determining the distance of the path to be selected is performed for each node to be calculated in each determined subgraph in sequence according to the coding order.

[0153] In one embodiment, when nodes in multiple subgraphs are determined to be nodes to be calculated, the distance of the path to be calculated of each subgraph can be determined in sequence according to the coding order. In combination with the above example, when determining the distance of the path to be calculated of a certain node, since the target path distance of the source node connected to the incoming edge of the node is not determined, it is necessary to wait for a certain time, and further when the target path distance of the node is needed to determine the distance of the path to be calculated of other nodes, further waiting is required. Therefore, the distance of the path to be calculated of each subgraph can be determined in sequence according to the coding order. For example, when the nodes in subgraph 300 and subgraph 400 are determined to be nodes to be calculated at the same time, the distance of the path to be calculated of the node to be calculated in subgraph 300 can be determined in the coding order first, and then the distance of the path to be calculated of the node to be calculated in subgraph 400 can be determined. At this time, when determining the distance of the path to be calculated of node R1, it is not necessary to wait for the determination of the target path distance of node P0.

[0154] Optionally, the step of determining the distance of the to-be-selected path for each to-be-calculated node in each determined subgraph in turn may include:

[0155] For each subgraph, the step of determining the distance of the to-be-selected path is performed on the nodes to be calculated in the subgraph in sequence according to the coding order of the nodes to be calculated in the subgraph.

[0156] In one embodiment, when determining the distance of the candidate path of the node to be calculated in each subgraph, the distance of the candidate path of each node in the subgraph can be determined in sequence according to the coding order. In combination with the above example, for the nodes to be calculated in subgraph 300, the distance of the candidate path of node D0 can be determined first, then the distance of the candidate path of node R0 can be determined, and then the distance of the candidate path of node E0 can be determined. By sequentially determining the distance of the candidate path of each node according to the coding order of multiple nodes to be calculated in subgraph 300, it is possible to avoid reducing waiting time and improve efficiency.

[0157] In the embodiment of the present invention, when nodes in multiple subgraphs are determined to be nodes to be calculated, the distance of the candidate path of the node to be calculated in each subgraph can be determined in sequence according to the coding order, which can reduce waiting time and improve efficiency.

[0158] In one embodiment, the Figure 2 The preprocessing device shown performs the steps of selecting a node to be calculated, determining the distance of a selected path, determining the distance of a target path, and determining the target path distance of all nodes and the step of the previous node. The preprocessing device includes a processing module, the processing module includes at least one processing unit, and the processing unit includes a plurality of processing sub-units. The processing module is connected to the instruction information cache table 2012, and the processing module can obtain graph data from the instruction information cache table 2012.

[0159] In this embodiment, the processing module is used to obtain nodes in a subgraph as nodes to be calculated for the idle processing unit in the coding order of multiple nodes from the instruction information cache table in sequence in the case that there is an idle processing unit in at least one processing unit, and obtain the incoming edge information of the nodes to be calculated for the idle processing unit. Further, the processing module is also used to assign the nodes to be calculated and the incoming edge information of the nodes to be calculated to a processing sub-unit in the idle processing unit, so that the processing sub-unit determines the path distance to be selected according to the weight of the incoming edge included in the incoming edge information and the target path distance of the connected target source node when the node to be calculated has at least one incoming edge; the target path distance of the target source node is the longest path distance and / or the shortest path distance among all path distances between the starting node and the target source node among the multiple nodes.

[0160] like Figure 6 As shown, Figure 6 FIG. 1 shows a schematic diagram of the structure of a preprocessing device in an embodiment of the present invention. The preprocessing device is composed of a processing module 600 and a cache module 700. The processing module 600 is connected to the cache module 700. The cache module 700 is also connected to the cache module 700. Figure 2 The memory module 2013 is shown connected. The pre-processing device can be a part of the prototype system or can exist independently of the prototype system.

[0161] The processing module 600 may include: Figure 6 The processing unit 601, the processing unit 602 and the processing unit 603 shown, Figure 6 The symbol PU stands for a processing unit, and the symbol PE stands for a processing sub-unit. After starting to process the graph data, the processing module can monitor the three processing units. When a processing unit is determined to be idle, the processing unit can be used as an idle processing unit. At this time, the processing module can obtain a node in a subgraph of the graph data from the instruction information cache table according to the encoding order of multiple nodes, and obtain the incoming edge information of the node to be calculated, and assign the obtained node as the node to be calculated to the idle processing unit. Combined with Figure 6As shown, when the processing of graph data is started, the three processing units in the processing module are all in an idle state. The processing module can extract the nodes in the first three subgraphs corresponding to instruction i0, instruction i1 and instruction i2 in sequence as nodes to be calculated in the coding order, and allocate the nodes to be calculated in the subgraph corresponding to instruction i0 to processing unit 601, allocate the nodes to be calculated in the subgraph corresponding to instruction i1 to processing unit 602, and allocate the nodes to be calculated in the subgraph corresponding to instruction i2 to processing unit 603. At the same time, the processing module can also obtain the incoming edge information of the node to be calculated from the instruction information cache table. The incoming edge information may include the source node connected to the incoming edge of the node to be calculated, and the weight of the incoming edge. The process of the processing module obtaining the node and the incoming edge information of the node can refer to the above example, and this embodiment will not be repeated here.

[0162] In this embodiment, after receiving the assigned node to be calculated and the incoming edge information, the processing subunit can first determine whether the node to be calculated has an incoming edge. If there is an incoming edge, determine the target path distance of the target source node connected by the incoming edge, and sum the target path distance and the weight of the incoming edge to obtain the selected path distance of the node to be calculated. The process of the processing subunit determining the target path distance of the node to be calculated and the previous node can refer to the above example, and this embodiment will not be repeated here.

[0163] Optionally, the processing subunit is also used to set the target path distance between the node to be calculated and the starting node to a first preset value and set the previous node of the node to be calculated to a second preset value when the node to be calculated is a node other than the starting node and has no incoming edge; the first preset value is an infinitesimal value corresponding to the longest path distance value and / or an infinite value corresponding to the shortest path distance value; the second preset value indicates that the previous node of the node to be calculated is empty.

[0164] Optionally, after selecting a target value as the target path distance of the node to be calculated from all the path distances to be selected, the target path distance of the node to be calculated can be cached in a cache module, so that when calculating the target path distances of other nodes to be calculated, if it is determined that the node to be calculated is a source node, the target path distance of the node to be calculated is obtained from the cache module.

[0165] Among them, the processing module is also used to apply for a storage area for each node to be calculated in the cache module, and output the target path distance and previous node of the node to be calculated to the cache module; the cache module is used to receive and store the target path distance and previous node of the node to be calculated in the storage area of ​​the node to be calculated; the processing module is also used to obtain the target path distance of the target source node from the cache module.

[0166] The processing subunit is specifically used to, if the target source node is the starting node, use the weight of the incoming edge as the distance of the path to be selected; if the processing module obtains the target path distance of the target source node from the cache module, then sum the weight of the incoming edge and the target path distance of the target source node to obtain the distance of the path to be selected; if the processing module does not obtain the target path distance of the target source node from the cache module, then wait for the processing module to obtain the target path distance of the target source node from the cache module, and then execute the summation step to obtain the distance of the path to be selected.

[0167] In one embodiment, after obtaining a node in a subgraph as a node to be calculated, the processing module may send an allocation request to the cache module to apply for multiple continuous storage areas from the cache module to store the target path distance and the previous node of each node to be calculated in the subgraph. After receiving the allocation request, the cache module may allocate a storage area in the cache module for each node to be calculated in the subgraph, and multiple storage areas of multiple nodes to be calculated in the same subgraph are sequentially located in the same area of ​​the cache module according to the coding order of the nodes, and each storage area stores the target path distance and the previous node of a node to be calculated.

[0168] like Figure 6 As shown, the cache module 700 includes multiple cache lines, the history length is the number of cache lines, and the number of storage areas in each cache line corresponds to the number of all nodes to be calculated included in a subgraph. An exclusive bit can be set at the head of each cache line, and each storage area included in the cache line can store valid bits, distances, and previous node data in sequence. The exclusive bit is used to mark whether the cache line is free. For example, when the exclusive bit is 1, it means that the cache line has been assigned to one of the subgraphs, and when it is 0, it means that the cache line is free. The valid bit is used to mark whether the data in the storage area is valid. For example, when the valid bit is 1, it means that the data in the storage area is valid, and when it is 0, it means that the data in the storage area is invalid. The distance is the target path distance of the node to be calculated, and the previous node array is one or more previous nodes of the node to be calculated. The head pointer in the cache module points to the free cache line in the cache module, and the head pointer number of the head pointer is initialized to 0. After receiving the allocation request sent by the processing module, the cache module can set the exclusive bit in the free cache line pointed to by the head pointer to 1, allocate the cache line to the subgraph, and send the current head pointer number to the processing module, then the head pointer number is increased by 1, and the head pointer points to the next free cache line. Each cache line is used to store the target path distance and previous node of a node in a subgraph, so each head pointer number corresponds to the target path distance and previous node of a node in a subgraph.

[0169] In one embodiment, the spatial size of each storage area is consistent and arranged in sequence. The processing module can determine the storage address of the storage area of ​​each node to be calculated in the cache line based on the head pointer number and the offset. The offset corresponds to the position of the node to be calculated in the subgraph. For example, if the node to be calculated is the first node in the subgraph, the offset corresponds to the first storage area in the cache line. If the node to be calculated is the third node in the subgraph, the offset corresponds to the third storage area in the cache line. The processing module assigns the node to be calculated in a subgraph to a processing unit, and after receiving the head pointer number sent by the cache module after assigning the cache line to the subgraph, the corresponding relationship between the processing unit and the head pointer number can be recorded. When a processing subunit in the processing unit determines the target path distance and the previous node of a node to be calculated, the processing module can first determine the position of the node to be calculated processed by the processing subunit in the subgraph, determine the offset according to the position, and then send a data packet including the corresponding head pointer number, offset, target path distance and previous node to the cache module. After receiving the data packet, the cache module first determines the corresponding cache line according to the head pointer number, and determines the target storage area in the cache line according to the offset, and stores the received target path distance and previous node in the target storage area. After storing the target path distance and previous node in a certain storage area, the cache module can set the valid position of the storage area to 1, indicating that the storage area has stored the target path distance and previous node of the node to be calculated, and the data in the storage area is valid data.

[0170] In this embodiment, when determining the selected path distance of a node to be calculated, the processing subunit needs to obtain the target path distance of the target source node connected by the incoming edge. At this time, the processing module can generate a first data request, which includes the head pointer number corresponding to the subgraph where the node to be calculated is located, and the offset of the node to be calculated. Correspondingly, after receiving the first data request, the cache module first determines the target cache line from all cache lines according to the head pointer number, and then determines the storage area of ​​the target source node from the target cache line according to the offset. When the valid bit of the storage area is 1, the target path distance of the target source node is read from the storage area, and the target path distance of the target source node is sent to the processing module, and the processing module forwards the target path distance to the corresponding processing subunit. On the contrary, when the cache module determines that the valid bit of the storage area is 0, it determines that the target path distance of the target source node has not been stored, and can send a response message to the processing module, so that the processing module controls the processing subunit to wait, and resends the first data request to the cache module after a preset time length until the target path distance of the target source node is obtained, and the target path distance is provided to the processing subunit.

[0171] In an embodiment of the present invention, a cache module is provided in the preprocessing device, and the cache module can cache the target path distance and the previous node of the node that have been calculated during the graph data processing process, and when the target path distance of a certain node is needed to calculate the target path distance of other nodes, the target path distance of the node can be quickly provided to the processing module, and direct acquisition from the memory module can be avoided, thereby improving efficiency. The data interaction process between the cache module and the processing module may include but is not limited to the above examples.

[0172] In this embodiment, the cache module is connected to the memory module. The cache module is also used to send the target path distances and previous nodes of all nodes to be calculated in the subgraph to the memory module when the target path distances and previous nodes of all nodes to be calculated in the subgraph are stored in the cache module, so that the memory module stores the target path distances and previous nodes of all nodes to be calculated in the subgraph; the cache module is also used to obtain the target path distance of the target source node from the memory module to provide the target path distance of the target source node to the processing module.

[0173] In this embodiment, the cache module is also used to determine that the target distance paths and previous nodes of all nodes in the corresponding subgraph have been determined when all valid bits in a cache line are 1, and the target path distances and previous nodes of all nodes in the cache line can be output to the memory module for storage. Figure 6 As shown, the memory module consists of a memory controller and a memory. The tail pointer in the cache module can point to a cache line to be written into the memory. When all valid bits in the cache line are 1, it means that the target path distance and the previous node of all nodes in the subgraph corresponding to the cache line have been determined. The cache module can send the target path distance and the previous node of all nodes stored in the cache line to the memory module, and the memory controller stores the received data in the memory. After writing the data in the cache line into the memory, the exclusive bit of the cache line can be cleared to make the cache line idle, and each valid bit in the cache line can be cleared to allocate the cache line to other subgraphs.

[0174] Optionally, when the cache module sends the target path distance and previous node in a cache line to the memory module, it can also send the head pointer number of the cache line to the memory module, so that the memory module stores the target path distance and previous node of the node in each subgraph received in sequence according to the head pointer number.

[0175] Among them, after receiving the first data request, the cache module can send a second data request to the memory module, and the second data request can also include a head pointer number and an offset. The memory control can obtain the target path distance and the previous node of the target source node from the memory according to the head pointer number and the offset, and return the target path distance and the previous node of the target source node to the cache module, and the cache module sends the target path distance and the previous node of the target source node to the processing module. Alternatively, the memory module can store the target path distance and the previous node of all nodes in each subgraph in the form of an array or a matrix, and the cache module can determine the position of the target path distance of the target source node in the array or the matrix according to the encoding position of the target source node in the graph data, so as to obtain the target path distance of the target source node from the memory module according to the determined position. The specific way in which the memory module stores the target path distance, and the data interaction process between the memory module and the cache module can be set according to the needs, and this embodiment does not limit this.

[0176] In the embodiment of the present invention, the processing module includes at least one processing unit, and the processing unit in the processing module includes multiple processing sub-units. For equal-increment graph data, combined with the specificity of the equal-increment graph data, the path distance between the node and the starting node is determined by the incoming edge of the node, so that the multiple processing sub-units in each processing unit can process all nodes in a sub-graph in parallel, thereby improving the processing efficiency of the graph data.

[0177] like Figure 2 As shown, the memory module includes a memory controller and a memory. The memory can be located in the prototype system or independent of the prototype system. The memory module and the cache module can be connected through a system bus. The cache module can send the target path distance and previous node of each node to the memory controller, and the memory controller stores the target path distance and previous node of the node into the memory. In addition, the preprocessing device can also obtain graph data from the instruction information cache table and store the graph data into the memory. The prototype system also includes a transmission unit, and the host is connected to the prototype system through the transmission unit. The host can obtain the target path distance and previous node of the node from the memory through the transmission unit, and obtain the graph data. After obtaining the graph data, the host can display it as follows Figure 3 The graph data shown in the figure can highlight the critical path in the graph data.

[0178] Among them, the host can be an electronic device such as a computer or server, and the host can also directly obtain graph data from the instruction information cache table 2012, and then execute the steps of selecting the nodes to be calculated, determining the distance of the selected path, and determining the target path distance until the target path distance and the previous node of all nodes are determined.

[0179] In one embodiment, a cache module may not be provided in the preprocessing device, and the processing module is directly connected to the memory module. After the processing subunit determines the target path distance and the previous node of a node to be calculated, the processing module may output the target path distance and the previous node of the node to be calculated to the memory module, so that the memory module stores the target path distance and the previous node of the node to be calculated to the target position in the memory module. When other processing subunits are determining the selected path distances of other nodes to be calculated, if the node to be calculated is determined to be a source node, the processing module may read the target path distance of the node to be calculated from the target position, and provide the target path distance of the node to be calculated to other processing subunits.

[0180] Optionally, after the processing module determines the target path distance and previous node of all nodes to be calculated in a subgraph, the processing module may send the target path distance and previous node of all nodes in the subgraph to the memory module. The memory module may sequentially store the target path distance and previous node of all nodes in a subgraph through a continuous storage area, taking the subgraph as a unit. Accordingly, when a certain processing subunit needs to obtain the target path distance of the source node, the processing module may send a data request to the memory module, and the memory module may send the target path distance of the source node to the processing module in response to the data request, and the processing module provides the target path distance of the source node to the processing unit. The specific method of storing the target path distance and the previous node in the memory module, and the specific process of the processing module obtaining the target path distance of the source node from the memory module can be set according to demand, and this embodiment does not limit this.

[0181] like Figure 7 As shown, Figure 7 A schematic diagram of the structure of a processor performance analysis device in an embodiment of the present invention is shown. The device 700 is arranged on a processor loaded with a benchmark test program. The benchmark test program includes multiple instructions, including:

[0182] An execution module 701 is used to execute multiple instructions in a pipeline according to the execution order of the multiple instructions; the execution process of each instruction includes multiple pipeline stages;

[0183] A first generating module 702, used to generate a first node of a current pipeline stage during the execution of each instruction;

[0184] The acquisition module 703 is used to acquire the delay time between the current pipeline stage and the historical pipeline stage when the execution of the current pipeline stage depends on the execution result of the completed historical pipeline stage;

[0185] A second generating module 704 is used to generate an edge between the second node and the first node of the historical pipeline stage, and use the delay time as the weight of the edge; the source node of the edge is the second node, and the destination node is the first node;

[0186] The storage module 705 is used to store the generated nodes of each pipeline stage, as well as the edges and the weights of the edges, to obtain graph data so as to analyze the performance of the processor based on the critical path between the start node and the end node in the graph data.

[0187] Optionally, the storage module 705 includes:

[0188] A partitioning unit, used for partitioning the nodes of all pipeline stages in each instruction into a node set;

[0189] The storage unit is used to store the node set and the first type of edges connecting the nodes in the node set as a subgraph, and store the second type of edges connecting the nodes in different sets, so as to obtain graph data composed of the subgraph and the second type of edges.

[0190] Optionally, the storage unit comprises:

[0191] Establishing a sub-unit for establishing and storing a graph data header, wherein the graph data header includes structural information of the sub-graph;

[0192] Add a subunit, used to add connection information of the second type of edge in the graph data header, where the connection information is used to describe the source node and the destination node of the second type of edge;

[0193] A storage subunit, used to set and store a unique subgraph number for each subgraph corresponding to each instruction in the order of execution, and when the subgraph includes a key node, store a first information set of the first type of edge in the subgraph and a second information set of the key node;

[0194] Among them, the first information set includes the weight of the first type of edge; the key node is the destination node of the target second type edge; the second information set includes the subgraph number of the subgraph where the source node of the target second type edge is located and the weight of the target second type edge, as well as the index of the connection information of the target second type edge in the graph data header.

[0195] Optionally, the apparatus 700 further includes: a selection unit, configured to select nodes in at least one subgraph as nodes to be calculated from the plurality of subgraphs in sequence according to the coding order of the nodes;

[0196] A first determining unit is used to determine the path distance to be selected according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge when the node to be calculated has at least one incoming edge; the target path distance of the target source node is the longest path distance and / or the shortest path distance among all path distances between the starting node and the target source node;

[0197] A second determining unit is used to determine a target path distance of the node to be calculated from all the path distances to be selected, and to use a source node corresponding to the target path distance of the node to be calculated as a previous node of the node to be calculated; the target path distance of the node to be calculated is a maximum path distance to be selected corresponding to the longest path distance and / or a minimum path distance to be selected corresponding to the shortest path distance;

[0198] The repeating unit is used to repeatedly execute the steps of selecting a node to be calculated, determining the distance of a selected path, and determining the target path distance until the target path distances and previous nodes of all nodes are determined.

[0199] Optionally, the first determination unit is specifically used to, if the target source node is the starting node, use the weight of the incoming edge as the path distance to be selected; if the target path distance of the target source node is predetermined, sum the weight of the incoming edge and the target path distance of the connected source node to obtain the path distance to be selected; if the target path distance of the target source node is not predetermined, wait until the target path distance of the target source node is determined, and then execute the summation step to obtain the path distance to be selected.

[0200] Optionally, the first determination unit is also used to set the target path distance between the node to be calculated and the starting node to a first preset value and set the previous node of the node to be calculated to a second preset value when the node to be calculated is a node other than the starting node and has no incoming edge; the first preset value is an infinitesimal value corresponding to the longest path distance value and / or an infinite value corresponding to the shortest path distance value; the second preset value indicates that the previous node of the node to be calculated is empty.

[0201] Optionally, the first determining unit is specifically configured to sequentially perform the step of determining the distance of the to-be-selected path on the to-be-calculated nodes in the selected at least one subgraph in a coding order.

[0202] Figure 8 The electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a message transceiver, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0203] Reference Figure 8, the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0204] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0205] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0206] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0207] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0208] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0209] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0210] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0211] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0212] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0213] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of an electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0214] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform a graph data processing method.

[0215] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0216] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0217] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0218] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal device to operate in a predictable manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0219] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0220] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0221] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0222] The above is a detailed introduction to a graph data processing method and device, an electronic device and a storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A processor performance analysis method, characterized in that: The method is applied to a processor loaded with a benchmark test program, wherein the benchmark test program includes a plurality of instructions; the method includes: Executing the plurality of instructions in a pipeline in the execution order of the plurality of instructions; the execution process of each of the instructions includes a plurality of pipeline stages; During the execution of each of the instructions, generating a first node of the current pipeline stage; In a case where the execution of the current pipeline stage depends on the execution result of a completed historical pipeline stage, obtaining a delay time between the current pipeline stage and the historical pipeline stage; Generate an edge between a second node of the historical pipeline stage and the first node, and use the delay time as a weight of the edge; the source node of the edge is the second node, and the destination node is the first node; storing the generated nodes of each pipeline stage, the edges and the weights of the edges to obtain graph data, so as to analyze the performance of the processor based on a critical path between a start node and an end node in the graph data; The starting node is the node of the first pipeline stage in the first instruction, and the ending node is the node of the last pipeline stage in the last instruction; The critical path is a path with the longest path distance among all paths from the starting node to the ending node in the graph data.

2. The method according to claim 1, characterized in that The nodes of each pipeline stage generated by the storage, as well as the edges and the weights of the edges, are obtained to obtain graph data, including: Dividing nodes of all pipeline stages in each of the instructions into a node set; The node set and the first type of edges connecting the nodes in the node set are stored as a subgraph, and the second type of edges connecting the nodes in different sets are stored to obtain graph data consisting of the subgraph and the second type of edges.

3. The method according to claim 2, characterized in that The step of storing the node set and the first type of edges connecting the nodes in the node set as a subgraph, and storing the second type of edges connecting the nodes in different sets, to obtain graph data consisting of the subgraph and the second type of edges, comprises: Establishing and storing a graph data header, wherein the graph data header includes structural information of the subgraph; Adding connection information of the second type of edge in the graph data header, where the connection information is used to describe a source node and a destination node of the second type of edge; According to the execution order, a unique subgraph number is sequentially set for the subgraph corresponding to each of the instructions and stored, and when the subgraph includes a key node, a first information set of the first type of edge in the subgraph is correspondingly stored, and a second information set of the key node is stored; Among them, the first information set includes the weight of the first type of edge; the key node is the destination node of the target second type edge; the second information set includes the subgraph number of the subgraph where the source node of the target second type edge is located and the weight of the target second type edge, as well as the index of the connection information of the target second type edge in the graph data header.

4. The method according to claim 2, characterized in that: After the nodes of each pipeline stage generated by the storage, as well as the edges and the weights of the edges are obtained, the method further includes: Selecting nodes in at least one subgraph from the multiple subgraphs in sequence as nodes to be calculated according to the coding order of the nodes; In the case where the node to be calculated has at least one incoming edge, the path distance to be selected is determined according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge; the target path distance of the target source node is the longest path distance and / or the shortest path distance among all path distances between the starting node and the target source node; Determine the target path distance of the node to be calculated from all the path distances to be selected, and use the source node corresponding to the target path distance of the node to be calculated as the previous node of the node to be calculated; the target path distance of the node to be calculated is the maximum path distance to be selected corresponding to the longest path distance and / or the minimum path distance to be selected corresponding to the shortest path distance; The steps of selecting the node to be calculated, determining the distance of the selected path, and determining the target path distance are repeatedly performed until the target path distances and previous nodes of all the nodes are determined.

5. The method according to claim 4, characterized in that The determining the distance of the to-be-selected path according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge includes: If the target source node is the starting node, the weight of the incoming edge is used as the distance of the path to be selected; If the target path distance of the target source node is predetermined, the weight of the incoming edge and the target path distance of the connected source node are summed to obtain the path distance to be selected; If the target path distance of the target source node is not determined in advance, the summing step is performed after the target path distance of the target source node is determined to obtain the candidate path distance.

6. The method according to claim 4, characterized in that Also includes: When the node to be calculated is a node other than the starting node and has no incoming edge, the target path distance between the node to be calculated and the starting node is set to a first preset value, and the previous node of the node to be calculated is set to a second preset value; the first preset value is an infinitesimal value corresponding to the longest path distance value and / or an infinite value corresponding to the shortest path distance value; the second preset value indicates that the previous node of the node to be calculated is empty.

7. The method according to any one of claims 4 to 6, characterized in that: The determining the distance of the to-be-selected path according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge includes: The step of determining the distance of the path to be selected is performed on the nodes to be calculated in the at least one selected subgraph in sequence according to the coding order.

8. A processor performance analysis device, characterized in that: The device is arranged on a processor loaded with a benchmark test program, wherein the benchmark test program includes a plurality of instructions; the device includes: An execution module, configured to execute the plurality of instructions in a pipeline in the execution order of the plurality of instructions; the execution process of each of the instructions includes a plurality of pipeline stages; A first generating module, used for generating a first node of a current pipeline stage during the execution of each of the instructions; An acquisition module, configured to acquire a delay time between the current pipeline stage and the historical pipeline stage when the execution of the current pipeline stage depends on the execution result of the completed historical pipeline stage; a second generating module, configured to generate an edge between a second node of the historical pipeline stage and the first node, and use the delay time as a weight of the edge; the source node of the edge is the second node, and the destination node is the first node; A storage module, used for storing the generated nodes of each pipeline stage, the edges and the weights of the edges, to obtain graph data, so as to analyze the performance of the processor based on the critical path between the starting node and the ending node in the graph data; The starting node is the node of the first pipeline stage in the first instruction, and the ending node is the node of the last pipeline stage in the last instruction; The critical path is a path with the longest path distance among all paths from the starting node to the ending node in the graph data.

9. The device according to claim 8, characterized in that The storage module comprises: A partitioning unit, used for partitioning the nodes of all pipeline stages in each of the instructions into a node set; A storage unit is used to store the node set and the first type of edges connecting the nodes in the node set as a subgraph, and store the second type of edges connecting the nodes in different sets, so as to obtain graph data composed of the subgraph and the second type of edges.

10. The device according to claim 9, characterized in that The storage unit comprises: Establishing a sub-unit, used to establish and store a graph data header, wherein the graph data header includes structural information of the sub-graph; An adding subunit is used to add connection information of the second type of edge in the graph data header, wherein the connection information is used to describe a source node and a destination node of the second type of edge; A storage subunit, configured to set and store a unique subgraph number for each subgraph corresponding to the instruction in the execution order, and, when the subgraph includes a key node, to store a first information set of a first type of edge in the subgraph and a second information set of the key node; Among them, the first information set includes the weight of the first type of edge; the key node is the destination node of the target second type edge; the second information set includes the subgraph number of the subgraph where the source node of the target second type edge is located and the weight of the target second type edge, as well as the index of the connection information of the target second type edge in the graph data header.

11. The device according to claim 9, characterized in that Also includes: A selection unit, configured to select nodes in at least one subgraph from the plurality of subgraphs in sequence as nodes to be calculated according to the coding order of the nodes; A first determining unit is used to determine the path distance to be selected according to the weight of the incoming edge and the target path distance of the target source node connected to the incoming edge when the node to be calculated has at least one incoming edge; the target path distance of the target source node is the longest path distance and / or the shortest path distance among all path distances between the starting node and the target source node; A second determining unit is used to determine the target path distance of the node to be calculated from all the path distances to be selected, and use the source node corresponding to the target path distance of the node to be calculated as the previous node of the node to be calculated; the target path distance of the node to be calculated is the maximum path distance to be selected corresponding to the longest path distance and / or the minimum path distance to be selected corresponding to the shortest path distance; The repeating unit is used to repeatedly execute the steps of selecting the node to be calculated, determining the distance of the selected path, and determining the target path distance until the target path distances and previous nodes of all the nodes are determined.

12. The device according to claim 11, characterized in that The first determination unit is specifically used to, if the target source node is the starting node, use the weight of the incoming edge as the path distance to be selected; if the target path distance of the target source node is predetermined, sum the weight of the incoming edge and the target path distance of the connected source node to obtain the path distance to be selected; if the target path distance of the target source node is not predetermined, wait until the target path distance of the target source node is determined, and then execute the summing step to obtain the path distance to be selected.

13. The device according to claim 11, characterized in that The first determination unit is also used to set the target path distance between the node to be calculated and the starting node to a first preset value and set the previous node of the node to be calculated to a second preset value when the node to be calculated is a node other than the starting node and has no incoming edge; the first preset value is an infinitesimal value corresponding to the longest path distance value and / or an infinite value corresponding to the shortest path distance value; the second preset value indicates that the previous node of the node to be calculated is empty.

14. The device according to any one of claims 11 to 13, characterized in that: The first determining unit is specifically configured to sequentially perform the step of determining the distance of the to-be-selected path on the to-be-calculated nodes in the at least one selected subgraph according to the coding order.

15. An electronic device, characterized in that: It comprises a processor performance analysis device as described in any one of claims 8 to 14.

Citation Information

Patent Citations

  • Network path determination and switching method and device, equipment, medium and program product

    CN113347083A

  • Kernel performance test method, computing equipment and storage medium

    CN113868068A