Calculation acceleration module and debugging method and device of calculation acceleration module
By deploying the controller between the debugging interface of the computing acceleration module and multiple processors, the problem of low debugging efficiency of multi-processors in the computing acceleration module is solved, and efficient processor debugging is achieved.
Patent Information
- Application Number
- CN202412000505.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The increase in the number of processors configured in the computing acceleration module results in a significant increase in the debug interface usage, which in turn reduces the debugging efficiency of the computing acceleration module to the processor.
The controller is deployed between the debugging interface of the calculation acceleration module and multiple processors. The controller connects the debugging device of the calculation acceleration module through the debugging interface, receives debugging requests and converts them into target control instructions for each processor, sends instructions to the processor and receives execution results, and converts them into debugging results and feeds back to the debugging device.
Through a debugging interface, multiple processors deployed in the computing acceleration module are debugged simultaneously, which avoids the problem of increased interface usage and improves the debugging efficiency of the computing acceleration module to the processor.
Smart Images

Figure CN120045407A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and more specifically, to a computing acceleration module, and a debugging method and device for the computing acceleration module. Background Art
[0002] The computing acceleration module is a key component in modern data centers and high-performance computing systems, and is designed to improve data processing speed and efficiency through highly integrated high-performance computing units. At present, a processor with data computing functions is usually deployed in the computing acceleration module architecture. In order to realize the special computing functions of the computing acceleration module, it is usually necessary to debug the computing performance of the processor in the computing acceleration module. In order to realize the debugging function of the processor, it is necessary to configure a debugging interface on the computing acceleration module, connect to the processor through the debugging interface, and thus realize the debugging of the computing performance of the processor through the debugging interface. With the increase in the demand for the computing performance of the computing acceleration module, the number of configurations in the computing acceleration module has gradually increased, that is, multiple processors are configured in the same computing acceleration module at the same time. In this case, in order to realize the debugging function of the computing performance of the computing acceleration module, it is necessary to configure a debugging interface corresponding to each processor on the computing acceleration module. However, the current configuration protocol of the computing acceleration module stipulates computing acceleration, which leads to a significant increase in the port occupancy of the computing acceleration module, which in turn leads to low debugging efficiency of the computing acceleration module for the processor. Summary of the invention
[0003] The embodiments of the present application provide a computing acceleration module, a debugging method and a device for the computing acceleration module, so as to at least solve the problem of low efficiency in regulating the computing performance of the computing acceleration module in the related art.
[0004] According to one embodiment of the present application, a computer acceleration module is provided, including:
[0005] A debugging interface, a plurality of processors and a controller, wherein the controller is connected to the plurality of processors respectively, and the controller is also used to connect to the debugging device of the computing acceleration module through the debugging interface;
[0006] The debugging interface is used to receive a debugging request from a debugging device for the operating performance of the plurality of processors;
[0007] The controller is used to convert a target control instruction for each of the processors according to the debugging request; and send the corresponding target control instruction to each of the processors;
[0008] The processor is configured to execute the received target control instruction to obtain a target execution result; and send the target execution result to the controller;
[0009] The controller is further used to receive the target execution results fed back by the multiple processors, convert the multiple target execution results into target debugging results corresponding to the debugging request; and call the debugging interface to feed back the target debugging results to the debugging device.
[0010] Optionally, the controller includes an instruction configuration component, and the instruction configuration component is connected to the debugging interface and the plurality of processors respectively;
[0011] The instruction configuration component is used to extract target operation information carried in the debugging request, wherein the target operation information is used to indicate the debugging operation to be performed by each processor when executing the debugging project requested by the debugging request; and convert the target control instruction for each processor according to the target operation information.
[0012] Optionally, the instruction configuration component is also used to locate identification information carried by the target operation information, wherein the identification information is used to identify the processor; search the target operation information for a target operation field that has a binding relationship with the identification information, wherein the target operation field is used to indicate the corresponding debugging operation to be performed by the processor; and determine the target control instruction corresponding to the target operation field from the operation fields and control instructions that have a corresponding relationship.
[0013] Optionally, the instruction configuration component is further used to extract a target operation field carried in the target operation information, wherein the target operation field is used to indicate a debugging operation to be performed by the processor; determine a reference control instruction corresponding to the target operation field from operation fields and control instructions having a corresponding relationship; and configure the reference control instruction as the target control instruction of each of the processors.
[0014] Optionally, the controller further includes a data conversion component and a plurality of data transmission channels, the data conversion component is respectively connected to the debug interface and the plurality of data transmission channels, the data transmission channels are configured in a one-to-one correspondence with the processors, and the data transmission channels are connected to the corresponding processors;
[0015] The data transmission channel is used to receive the target execution result reported by the corresponding processor;
[0016] The data conversion component is used to extract candidate execution results from the target execution results received on each data transmission channel according to data requirement information of the debugging device, wherein the data requirement information is used to indicate the debugging device's requirement for the processor execution result when executing the debugging item corresponding to the debugging request; and mark the candidate execution result using the channel identifier of the data transmission channel to which the candidate execution result belongs to obtain the target debugging result.
[0017] Optionally, the data transmission channel includes a buffer and a data transmission port, the buffer is connected to the data transmission port and the data conversion component respectively, and the data transmission port is also connected to the processor corresponding to the data transmission channel;
[0018] The data transmission port is used to send the target execution result to the buffer after receiving the target execution result;
[0019] The buffer is used to cache the received target execution results in a chronological order;
[0020] The data conversion component is used to extract the candidate execution results of the target data requirement from the target execution results cached in the buffer in a chronological order; use the channel identifier to mark the candidate execution results of the target data requirement to obtain the target debugging result, wherein the data requirement information includes the target data requirement.
[0021] According to another embodiment of the present application, a debugging method for a computing acceleration module is provided, and the method is applied to a controller deployed in the computing acceleration module, comprising:
[0022] Converting target control instructions for each processor according to the debugging request, wherein the computing acceleration module includes a debugging interface, a plurality of the processors and the controller, the controller is respectively connected to the plurality of the processors, the controller is also connected to the debugging device of the computing acceleration module through the debugging interface, the debugging request is a request sent by the debugging device received by the debugging interface, and the debugging request is used to request debugging of the operating performance of the plurality of the processors;
[0023] Sending the corresponding target control instruction to each processor, wherein the processor is used to execute the received corresponding target control instruction to obtain a corresponding target execution result;
[0024] receiving the target execution results fed back by the multiple processors, and converting the multiple target execution results into target debugging results corresponding to the debugging request;
[0025] The debugging interface is called to feed back the target debugging result to the debugging device.
[0026] According to another embodiment of the present application, a debugging device for a computing acceleration module is provided.
[0027] The device is applied to a controller deployed in a computing acceleration module, and the device includes:
[0028] A conversion module, used for converting target control instructions for each processor according to a debugging request, wherein the computing acceleration module includes a debugging interface, a plurality of the processors and the controller, the controller is respectively connected to the plurality of the processors, the controller is also connected to a debugging device of the computing acceleration module through the debugging interface, the debugging request is a request sent by the debugging device received by the debugging interface, and the debugging request is used for requesting debugging of the operating performance of the plurality of the processors;
[0029] A sending module, used for sending the corresponding target control instruction to each processor, wherein the processor is used for executing the received corresponding target control instruction to obtain a corresponding target execution result;
[0030] A processing module, configured to receive the target execution results fed back by the multiple processors, and convert the multiple target execution results into target debugging results corresponding to the debugging request;
[0031] A feedback module is used to call the debugging interface to feed back the target debugging result to the debugging device.
[0032] According to another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when run.
[0033] According to another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0034] Through the present application, by deploying a controller between the debugging interface of the computing acceleration module and multiple processors, the controller can be connected to the debugging device of the computing acceleration module through the debugging interface. When the debugging interface receives a debugging request from the debugging device for requesting the operating performance of multiple processors, the controller can receive the debugging request and convert it into a target control instruction for each processor, and send the corresponding target control instruction to each processor. Then, after receiving the corresponding target control instruction, the processor sends the corresponding target execution result obtained to the controller, and the controller can convert the received multiple target execution results into target debugging results corresponding to the debugging request, and send the target debugging results to the debugging device. Through the above method, it is possible to debug multiple processors deployed in the computing acceleration module at the same time through a debugging interface, avoid the interface occupancy on the computing acceleration module when debugging the computing performance of the computing acceleration module, and ensure the computing performance of the computing acceleration module. Therefore, the problem of low efficiency in regulating the computing performance of the computing acceleration module in the related art can be solved, and the effect of improving the efficiency of regulating the computing performance of the computing acceleration module can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic diagram of a computing acceleration module according to an embodiment of the present application;
[0036] Figure 2 is a flowchart of a debugging method for a computing acceleration module according to an embodiment of the present application;
[0037] Figure 3 It is an improved UART debugging topology solution 1 according to an embodiment of the present application;
[0038] Figure 4 This is an improved UART debugging topology solution 2 according to an embodiment of the present application;
[0039] Figure 5 A UART TX data encoding and decoding process according to an embodiment of the present application is shown in FIG. Figure 1 ;
[0040] Figure 6 A UART TX data encoding and decoding process according to an embodiment of the present application is shown in FIG. Figure 2 ;
[0041] Figure 7 A UART RX data encoding and decoding process according to an embodiment of the present application is shown in FIG. Figure 3 ;
[0042] Figure 8 It is a hardware structure block diagram of a server device of a debugging method for a computing acceleration module according to an embodiment of the present application;
[0043] Fig. 9 It is a structural block diagram of a debugging device for a computing acceleration module according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0045] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0046] In this embodiment, a computing acceleration module is provided. Figure 1 is a schematic diagram of a computing acceleration module according to an embodiment of the present application, such as Figure 1 As shown, the module includes: a debugging interface, multiple processors and a controller, wherein the controller is connected to the multiple processors respectively, and the controller is also used to connect to the debugging device of the computing acceleration module through the debugging interface;
[0047] The debugging interface is used to receive a debugging request from a debugging device for the operating performance of the plurality of processors;
[0048] The controller is used to convert a target control instruction for each of the processors according to the debugging request; and send the corresponding target control instruction to each of the processors;
[0049] The processor is configured to execute the received target control instruction to obtain a target execution result; and send the target execution result to the controller;
[0050] The controller is further used to receive the target execution results fed back by the multiple processors, convert the multiple target execution results into target debugging results corresponding to the debugging request; and call the debugging interface to feed back the target debugging results to the debugging device.
[0051] Through the above steps, by deploying a controller between the debugging interface of the computing acceleration module and multiple processors, the controller can be connected to the debugging device of the computing acceleration module through the debugging interface. When the debugging interface receives a debugging request from the debugging device for requesting the operating performance of multiple processors, the controller can receive the debugging request and convert it into a target control instruction for each processor, and send the corresponding target control instruction to each processor. After receiving the corresponding target control instruction, the processor sends the corresponding target execution result obtained to the controller, and the controller can convert the received multiple target execution results into target debugging results corresponding to the debugging request, and send the target debugging results to the debugging device. Through the above method, it is possible to debug multiple processors deployed in the computing acceleration module at the same time through a debugging interface, avoid the interface occupancy on the computing acceleration module when debugging the computing performance of the computing acceleration module, and ensure the computing performance of the computing acceleration module. Therefore, the problem of low efficiency in regulating the computing performance of the computing acceleration module in the related art can be solved, and the effect of improving the efficiency of regulating the computing performance of the computing acceleration module can be achieved.
[0052] Optionally, in an embodiment of the present application, a computing acceleration module is a hardware component used to accelerate specific types of computing tasks. The computing acceleration module may be, but is not limited to, an OAM (OCP Accelerator Module, Open Computing Project Accelerator Module), a GPU accelerator card (Graphics Processing Unit Accelerator Card), etc. This solution does not limit this.
[0053] Optionally, in an embodiment of the present application, the target control instruction is used to control the debugging operation performed by the corresponding processor according to the debugging requirements of the debugging request.
[0054] Optionally, in an embodiment of the present application, the debug interface is a connecting component used to implement communication between the controller and the debug device. The debug interface is used to receive a debug request sent by the debug device and send it to the controller, and send the debug result returned by the controller to the debug device. In the present application, the debug interface can be but is not limited to a UART (Universal Asynchronous Receiver / Transmitter) interface or a JTAG (Joint Test Action Group) interface, etc., that is, in the present application, at least UART and JTAG are provided for debugging, and this solution does not limit this.
[0055] Optionally, in an embodiment of the present application, the processor is a unit that performs specific computing tasks inside the computing acceleration module. After debugging is completed, the processor can perform computing tasks issued by the host device to which the computing acceleration module is connected under the target operating performance. In the present application, the processor can be, but is not limited to, a GPU (Graphics Processing Unit), a CPU (Central Processing Unit), a DPU (Data Processing Unit), an FPGA (Field-Programmable Gate Array) with integrated data computing functions, etc., and this solution does not limit this.
[0056] Optionally, in an embodiment of the present application, the controller is a computing device that integrates the function of converting debugging requests into control instructions of each processor, and converting the execution results returned by each processor into debugging results. In the present application, the controller can be but is not limited to CPLD (Complex Programmable Logic Device), FPGA (Field-Programmable Gate Array), MCU (Microcontroller), etc., and this solution does not limit this.
[0057] Optionally, in an embodiment of the present application, the debugging device is a device that can communicate with the computing acceleration module and is used to monitor, control and debug the internal hardware of the module. In the present application, the debugging device can be a BMC (Baseboard Management Controller) or HOST deployed in the server.
[0058] Optionally, in the embodiments of the present application, Figure 1 As shown, the module includes: a debugging interface, multiple processors and a controller, the controller is connected to the multiple processors respectively, and the controller is connected to the debugging device of the computing acceleration module through the debugging interface. After receiving the debugging request sent by the debugging device for requesting to debug the running performance of multiple processors, the debugging interface sends the debugging request to the controller. After receiving the debugging request sent by the debugging interface, the controller converts the debugging request into a target control instruction corresponding to each processor in the computing acceleration module, and sends the corresponding target control instruction (such as Figure 1 The controller shown in FIG. 1 sends target control instruction 1 to processor 1 and sends target control instruction 2 to processor 2), and then each processor returns the target execution result obtained after executing the corresponding target control instruction to the controller (such as Figure 1The processor 1 shown sends the target execution result 1 after executing the target control instruction 1 to the controller, and the processor 2 sends the target execution result 2 after executing the target control instruction 2 to the controller).
[0059] Optionally, in an embodiment of the present application, when debugging the operating performance of the processor on the computing acceleration module, the debugging device may need different processors to perform different debugging operations respectively. Therefore, the debugging device may have different data requirements for feedback from different processors, that is, the debugging device may have different requirements for debugging results of different processors. For example, the debugging device needs to obtain the operating temperature data of processor 1 and the power supply voltage data of processor 2, and the execution results returned by the processors may contain data information that is not required by the debugging device. Therefore, after receiving the target execution results sent from each processor, the controller will convert multiple target execution results into target debugging results according to the data requirement information of the debugging device, that is, according to the data requirements of the debugging device for each processor, the execution results that match the data requirement information are extracted from the target execution results sent by multiple processors, and the extracted multiple execution results are converted into target debugging results. This can be achieved in the following ways: extracting the data requirement type of each processor when the debugging device debugs the debugging project corresponding to the debugging request from the data requirement information, wherein the data requirement type is used to indicate the execution result type that the debugging device needs each processor to return; extracting the candidate execution result that matches the execution result type indicated by the data requirement type from the target execution result sent by the processor; encapsulating the candidate execution result to obtain a reference execution result, wherein the reference execution result records the correspondence between the candidate execution result and the corresponding processor identifier and timestamp, wherein the processor identifier is used to identify the processor to which the candidate execution result belongs, and the timestamp is used to indicate the time when the processor sends the candidate execution result to the controller; merging multiple reference execution results to obtain the target debugging result. Thus, the debugging result received by the debugging device can meet the data requirement of the debugging device when debugging the debugging project corresponding to the debugging request.
[0060] Optionally, in an embodiment of the present application, the debugging device can perform multiple rounds of debugging on multiple processors until the operating performance of the multiple processors is debugged to the target operating performance, wherein the multiple rounds of debugging can be implemented in the following manner: S1: the debugging device sends a debugging request to the controller through the debugging interface; S2: the controller receives the debugging request and converts it into a target control instruction corresponding to each processor in the computing acceleration module, and sends the corresponding target control instruction to each processor; S3: after receiving the corresponding target control instruction, the processor sends the corresponding target execution result obtained to the controller; S4: after receiving the target execution result fed back by the processor, the controller sends the multiple target execution results according to the data requirement information of the debugging device. Convert the target debugging result into a target debugging result, and send the target debugging result to the debugging device; S5: the debugging device detects whether the operating performance of multiple processors reaches the target operating performance according to the target debugging result; S51: when the detection result indicates that the operating performance of multiple processors has reached the target operating performance, issue corresponding computing tasks to each processor; S52: when the detection result indicates that there is a first processor whose operating performance does not reach the target operating performance among the multiple processors, send a first debugging request to the controller through the debugging interface, the first debugging request is used to request debugging of the operating performance of the first processor, and repeat steps S2 to S5 until the operating performance of multiple processors is debugged to the target operating performance.
[0061] As an optional embodiment, the controller includes an instruction configuration component, and the instruction configuration component is connected to the debugging interface and the plurality of processors respectively;
[0062] The instruction configuration component is used to extract target operation information carried in the debugging request, wherein the target operation information is used to indicate the debugging operation to be performed by each processor when executing the debugging project requested by the debugging request; and convert the target control instruction for each processor according to the target operation information.
[0063] Optionally, in an embodiment of the present application, the controller also includes an instruction configuration component, which is respectively connected to the debugging interface and multiple processors. When the debugging device sends a debugging request, the instruction configuration component parses the target operation information in the debugging request. The target operation information includes specific operations that need to be performed by each processor in the debugging project, and then generates control instructions for each processor corresponding to the operations to be performed.
[0064] Optionally, in an embodiment of the present application, the instruction configuration component is a component for parsing and allocating received debugging device instructions. The instruction configuration component may be another controller in the computing acceleration module, that is, in the computing acceleration module, controller A (instruction configuration component) is used to parse and allocate received debugging device instructions to each processor, and another controller B is used to convert the target execution result fed back by the processor into a target debugging result and send it to the debugging device; alternatively, the instruction configuration component may also be a functional area in the controller that integrates the function of parsing and allocating received debugging device instructions, and this scheme does not limit this.
[0065] Through the above content, the instruction configuration component can accurately parse and convert the target operation information in the debugging request, generate precise control instructions for each processor, avoid unnecessary data transmission and processing, and improve the debugging efficiency of the computing acceleration module for multiple processors.
[0066] As an optional embodiment, the instruction configuration component is further used to locate identification information carried by the target operation information, wherein the identification information is used to identify the processor; search the target operation information for a target operation field that has a binding relationship with the identification information, wherein the target operation field is used to indicate the corresponding debugging operation to be performed by the processor; and determine the target control instruction corresponding to the target operation field from the operation fields and control instructions that have a corresponding relationship.
[0067] Optionally, in an embodiment of the present application, the target operation information also includes identification information for identifying a reference processor among multiple processors. The identification information may be the name, number or other unique identification of the processor, and then the instruction configuration component can search for the target operation field bound to the identification information based on the identification information, that is, determine the business operation to be executed corresponding to the reference processor, and then obtain the target control instruction corresponding to the target operation field from a pre-constructed mapping table of the correspondence between operation fields and control instructions, thereby realizing the generation of one-to-one corresponding operation instructions for each processor.
[0068] In the above manner, the controller identifies the operations to be executed by each processor from the target operation information and generates operation instructions that accurately match each processor. This one-to-one adaptation mechanism ensures the targeting of control instructions, avoids the ambiguity or mismatch of instructions in traditional debugging, and improves the response speed of the processor and the debugging efficiency of the processor.
[0069] As an optional embodiment, the instruction configuration component is also used to extract the target operation field carried in the target operation information, wherein the target operation field is used to indicate the debugging operation to be performed by the processor; determine a reference control instruction corresponding to the target operation field from the operation fields and control instructions with a corresponding relationship; and configure the reference control instruction as the target control instruction of each of the processors.
[0070] Optionally, in an embodiment of the present application, the target operation information includes debugging operations to be performed by multiple processors. The instruction configuration component first extracts all target operation fields from the target operation information, and determines the reference control instruction that matches each target operation field based on the correspondence table between the operation field and the control instruction, and configures the determined reference control instruction as the target control instruction of each processor, that is, the configuration method of the target control instruction here is to send the reference control instruction corresponding to all identified target operation fields to each processor, that is, each processor needs to execute all operations carried in the operation information. After each processor executes all received target control instructions, it will return the target execution result to the controller. After receiving the target execution result, the controller will perform summary processing, such as extracting the execution result that matches the data requirement information from the target execution results sent by multiple processors according to the data requirements of the debugging device for each processor, and converting the extracted multiple execution results into target debugging results.
[0071] Through the above method, each processor executes all business operation instructions, and the system can comprehensively monitor and manage all processors, collect complete system status data, and contribute to the overall health detection and resource optimization of the system.
[0072] As an optional embodiment, the controller further includes a data conversion component and a plurality of data transmission channels, the data conversion component is respectively connected to the debug interface and the plurality of data transmission channels, the data transmission channels are configured in a one-to-one correspondence with the processors, and the data transmission channels are connected to the corresponding processors;
[0073] The data transmission channel is used to receive the target execution result reported by the corresponding processor;
[0074] The data conversion component is used to extract candidate execution results from the target execution results received on each data transmission channel according to data requirement information of the debugging device, wherein the data requirement information is used to indicate the debugging device's requirement for the processor execution result when executing the debugging item corresponding to the debugging request; and mark the candidate execution result using the channel identifier of the data transmission channel to which the candidate execution result belongs to obtain the target debugging result.
[0075] Optionally, in the embodiment of the present application, multiple data transmission channels are established inside the controller, and each channel is directly connected to a processor in the system, realizing a one-to-one correspondence between the channel and the processor. Each channel is assigned a unique channel identifier, such as "Channel_GPU0", "Channel_GPU1". These channel identifiers are used by the data conversion component to identify the data source, add tags to the target execution results, and ensure that the debugging device can accurately parse the execution status of each processor.
[0076] Optionally, in an embodiment of the present application, after the processor executes the target control instruction sent by the controller, it will report the execution results such as temperature data, frequency information, etc. to the controller through the corresponding data transmission channel. After the data conversion component in the controller receives these target execution results, it first obtains the channel identifier of the data transmission channel, and then uses this identifier to mark the execution result, such as "Channel_GPU0: Temperature = 50°C". After marking the target execution results, the data conversion component will further convert these data into a format that the debugging device can understand, such as converting UART signals into USB signals or network signals. The converted data is the target debugging result, which is then fed back to the debugging device through the debugging interface.
[0077] Optionally, in an embodiment of the present application, the data conversion component is a component used to convert the target execution result into the target debugging result, and the data conversion component may be another controller in the computing acceleration module; or, the data conversion component may also be a functional area in the controller that integrates the function of converting the target execution result into the target debugging result, and this solution does not limit this.
[0078] Through the above method, by introducing data conversion components and multiple data transmission channels, accurate marking and efficient feedback of processor execution results are achieved, the traditional debugging process is simplified, data conversion or processing in the intermediate links is avoided, and the debugging efficiency of multiple processors is improved.
[0079] As an optional embodiment, the data transmission channel includes a buffer and a data transmission port, the buffer is connected to the data transmission port and the data conversion component respectively, and the data transmission port is also connected to the processor corresponding to the data transmission channel;
[0080] The data transmission port is used to send the target execution result to the buffer after receiving the target execution result;
[0081] The buffer is used to cache the received target execution results in a chronological order;
[0082] The data conversion component is used to extract the candidate execution results of the target data requirement from the target execution results cached in the buffer in a chronological order; use the channel identifier to mark the candidate execution results of the target data requirement to obtain the target debugging result, wherein the data requirement information includes the target data requirement.
[0083] Optionally, in an embodiment of the present application, a data transmission channel includes a cache and a data transmission port. Each processor sends the target execution result generated after executing the target control instruction to the cache through its corresponding data transmission port. The cache is responsible for receiving the target execution result sent from the data transmission port, and caching multiple target execution results in the time order of data arrival, so as to effectively store and manage the execution results from different processors and ensure the logical order of data processing; the data conversion component will first extract the target data requirement for indicating the number of target execution results that the debugging device expects to obtain from the data requirement information, and then extract the corresponding number of execution results from the cache in chronological order, and then use the channel identifier to mark each execution result.
[0084] Optionally, in an embodiment of the present application, the cache is a component used to store the target execution results sent by the processor in chronological order. The cache can be a cache device deployed in each data transmission channel, such as a hard disk, flash memory, etc., or it can be a plurality of sub-cache spaces divided in a large cache space of the controller, and each sub-cache space is allocated to a different data transmission channel. This solution does not limit this.
[0085] In the above manner, by introducing a buffer and a data transmission port in the data transmission channel, the buffer can store the target execution results received from the processor in a time sequence, and the data conversion component can batch extract these cached execution results. This batch processing method greatly reduces the number of interactions between the data conversion component and the buffer, thereby reducing the delay of data processing and improving the debugging efficiency of multiple processors.
[0086] As an optional embodiment, a debugging method for a computing acceleration module is provided in this embodiment. Figure 2 is a flowchart of a debugging method for a computing acceleration module according to an embodiment of the present application, wherein the method is applied to a controller deployed in the computing acceleration module, such as Figure 2 As shown, the process includes the following steps:
[0087] Step S202, converting a target control instruction for each processor according to a debugging request, wherein the computing acceleration module includes a debugging interface, a plurality of the processors and the controller, the controller is respectively connected to the plurality of the processors, the controller is also connected to a debugging device of the computing acceleration module through the debugging interface, the debugging request is a request sent by the debugging device received by the debugging interface, and the debugging request is used to request debugging of the operating performance of the plurality of the processors;
[0088] Step S204, sending the corresponding target control instruction to each processor, wherein the processor is used to execute the received corresponding target control instruction to obtain a corresponding target execution result;
[0089] Step S206, receiving the target execution results fed back by the multiple processors, and converting the multiple target execution results into target debugging results corresponding to the debugging request;
[0090] Step S208: calling the debugging interface to feed back the target debugging result to the debugging device.
[0091] Through the above content, by deploying a controller between the debugging interface of the computing acceleration module and multiple processors, the controller can be connected to the debugging device of the computing acceleration module through the debugging interface. When the debugging interface receives a debugging request from the debugging device for requesting the operating performance of multiple processors, the controller can receive the debugging request and convert it into a target control instruction for each processor, and send the corresponding target control instruction to each processor. Then, after receiving the corresponding target control instruction, the processor sends the corresponding target execution result obtained to the controller, and the controller can convert the received multiple target execution results into target debugging results corresponding to the debugging request, and send the target debugging results to the debugging device. Through the above method, it is possible to debug multiple processors deployed in the computing acceleration module at the same time through a debugging interface, avoid the interface occupancy on the computing acceleration module when debugging the computing performance of the computing acceleration module, and ensure the computing performance of the computing acceleration module. Therefore, the problem of low efficiency in regulating the computing performance of the computing acceleration module in the related art can be solved, and the effect of improving the efficiency of regulating the computing performance of the computing acceleration module can be achieved.
[0092] Optionally, converting a target control instruction for each of the processors according to the debugging request includes:
[0093] Extract target operation information carried in the debugging request, wherein the target operation information is used to indicate the debugging operation to be performed by each processor when executing the debugging project requested by the debugging request; and convert the target control instruction for each processor according to the target operation information.
[0094] Optionally, converting the target control instruction for each processor according to the target operation information includes:
[0095] Locate identification information carried by the target operation information, wherein the identification information is used to identify the processor; search the target operation information for a target operation field having a binding relationship with the identification information, wherein the target operation field is used to indicate a corresponding debugging operation to be performed by the processor; and determine the target control instruction corresponding to the target operation field from operation fields and control instructions having a corresponding relationship.
[0096] Optionally, converting the target control instruction for each processor according to the target operation information includes:
[0097] Extracting a target operation field carried in the target operation information, wherein the target operation field is used to indicate a debugging operation to be performed by the processor; determining a reference control instruction corresponding to the target operation field from operation fields and control instructions having a corresponding relationship; and configuring the reference control instruction as the target control instruction of each of the processors.
[0098] Optionally, converting the plurality of target execution results into a target debugging result corresponding to the debugging request includes:
[0099] Receiving the target execution result reported by the corresponding processor;
[0100] Obtain a channel identifier corresponding to the data transmission channel, wherein the channel identifier is used to indicate the processor connected to the data transmission channel; use the channel identifier to mark the target execution result received by the corresponding data transmission channel that meets the data requirement information, and obtain the target debugging result.
[0101] Optionally, the using the channel identifier to mark the target execution result that meets the data requirement information and is received by the corresponding data transmission channel to obtain the target debugging result includes:
[0102] Cache the received target execution results in chronological order;
[0103] Extracting the target execution result of the target data requirement from the execution results cached in the buffer in chronological order; marking the target execution result of the target data requirement using the channel identifier to obtain the target debugging result, wherein the data requirement information includes the target data requirement.
[0104] As an optional embodiment, this embodiment also provides a hardware debugging design method for dual-GPU architecture OAM based on the OAM specification, which can achieve complete debugging of the two GPU chips inside the entire OAM under the premise of complying with the OAM2.0 specification.
[0105] Figure 3 This is an improved UART debugging topology solution 1 according to an embodiment of the present application, such as Figure 3 As shown in the figure, the UART lines of the two GPUs inside the OAM are connected to the CPLD or FPGA inside the OAM after Levelshift. The CPLD or FPGA parses the UART TX signals from the two GPUs, and then re-encodes them and sends them out to the high-speed daughter board connector Conn0 via UART. The RX signal from the high-speed daughter board connector Conn0 is directly transparently transmitted to the two GPUs by the CPLD or FPGA inside the OAM. The UBB connects the UART signal of the OAM according to the standard signal definition and implements the UART-based debugging function of the internal chip of the OAM according to the standard topology.
[0106] Figure 4 This is an improved UART debugging topology solution 2 according to an embodiment of the present application, such as Figure 4 As shown in the figure, the UART lines of the two GPUs inside the OAM are connected to the CPLD or FPGA inside the OAM after Levelshift. The CPLD or FPGA parses the UART TX signals from the two GPUs, and then re-encodes them and sends them out through UART to the high-speed daughter board connector Conn0. At the same time, the RX signal from the high-speed daughter board connector Conn0 is also parsed by the CPLD or FPGA inside the OAM, and then re-encoded and sent to the two GPUs. The UBB connects the UART signal of the OAM according to the standard signal definition and implements the UART-based debugging function of the internal chip of the OAM according to the standard topology.
[0107] In order to clearly illustrate the implementation of this design method, Figure 3 and Figure 4 Design a solution to illustrate the implementation steps, as follows:
[0108] for Figure 3In the scheme shown, the UART level on the GPU inside the OAM of the dual-GPU architecture is generally 1.2V or 1.8V, and the UART TX signals of GPU0 and GPU1 are respectively connected to Levelshift, converted to 3.3V by Levelshift and connected to the CPLD or FPGA inside the OAM. The CPLD or FPGA decodes the UART TX signals received from the two GPUs, and adds the GPU serial number (adds "GPU0:" and "GPU1:") prefixes before the data sent by the two GPUs through UART (the data sent by the GPU can also be divided into blocks, and then the GPU serial number prefix is added to the whole block), and then the modified data is encoded and sent to the high-speed gusset connector Conn0 through UART; in addition, the CPLD or FPGA inside the OAM will transparently transmit the UARTRX signal sent from the high-speed gusset connector Conn0 to the two internal GPUs; the hardware on the UBB is designed according to the specification, and the UART signal of the OAM is connected to the UART to USB chip CP2108 and then converted into a USB signal to MICROUSB and HostBMC;
[0109] for Figure 4 The scheme shown, in Figure 3 The solution is further expanded to enrich the freedom of UART debugging. In OAM, in addition to receiving the UARTTX signals sent by the two GPUs to the CPLD or FPGA for decoding and re-encoding, the UARTRX signals sent from the high-speed daughter board connector Conn0 are also decoded and re-encoded inside the CPLD or FPGA. Figure 3 The scheme shown will not be repeated here; the UARTRX signal will also be decoded after being received by the CPLD or FPGA, and more debugging functions can be achieved by decoding and identifying specific instructions; if the host indicates through the "OAM_GPU0_ONLY" instruction that only the OAM internal GPU0 needs to return information, then its corresponding data acquisition instruction (such as "Core_Temp") will only be sent to GPU0 after being decoded and re-encoded by the OAM internal CPLD or FPGA, and will not be sent to GPU1. Subsequently, only GPU0 will return the corresponding data information and send it to the Host end through the OAM internal CPLD or FPGA, and GPU1 will not return data.
[0110] Here are some examples:
[0111] Figure 5 A UART TX data encoding and decoding process according to an embodiment of the present application is shown in FIG. Figure 1 , Figure 6A UART TX data encoding and decoding process according to an embodiment of the present application is shown in FIG. Figure 2 ,like Figure 5 and Figure 6 As shown, for the UART TX signal sent by the GPU chip inside the OAM, the CPLD or FPGA performs decoding and re-encoding operations on the signal.
[0112] Figure 7 A UART RX data encoding and decoding process according to an embodiment of the present application is shown in FIG. Figure 3 ,like Figure 7 As shown, for the UART RX signal coming from the high-speed daughter board connector Conn0, the CPLD or FPGA performs decoding and re-encoding operations on the signal.
[0113] The above diagrams and examples only illustrate the decoding and re-encoding functions of the CPLD inside the OAM. The decoding and re-encoding processing logic of the actual CPLD is not limited and all fall within the scope of the requirements of this solution.
[0114] This solution designs a dual-GPU OAM hardware debugging design method based on the OAM specification, and realizes the complete debugging of the two GPUs inside the OAM with one UART signal under the premise of complying with the OAM interface specification.
[0115] This solution can decode and re-encode UART communication data through the CPLD inside the OAM, which can make the OAM interface signal definition follow the OAM specification, ensure that the OAM adaptation work is simpler, increase the OAM shipment volume, and increase the manufacturer's profits; at the same time, it can realize the complete debugging of the two GPUs inside the OAM through one UART signal, and realize more debugging functions based on the decoding and re-encoding of the CPLD.
[0116] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 8 1 is a hardware structure block diagram of a server device of a method for debugging a computing acceleration module according to an embodiment of the present application. Figure 8 As shown, the server device may include one or more ( Figure 8 Only one is shown in the figure) a processor 802 (the processor 802 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 804 for storing data, wherein the server device may also include a transmission device 806 and an input / output device 808 for communication functions. A person skilled in the art will understand that Figure 8 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 8 More or fewer components as shown, or with Figure 8 Different configurations are shown.
[0117] The memory 804 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the debugging method of the computing acceleration module in the embodiment of the present application. The processor 802 executes various functional applications and data processing by running the computer program stored in the memory 804, that is, to implement the above method. The memory 804 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 804 may further include a memory remotely arranged relative to the processor 802, and these remote memories may be connected to the server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0118] The transmission device 806 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the server device. In one example, the transmission device 806 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 806 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0120] In this embodiment, a debugging device for a computing acceleration module is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0121] Fig. 9 is a structural block diagram of a debugging device for a computing acceleration module according to an embodiment of the present application, wherein the device is applied to a controller deployed in the computing acceleration module, such as Fig. 9 As shown, the device comprises:
[0122] A conversion module, used for converting target control instructions for each processor according to a debugging request, wherein the computing acceleration module includes a debugging interface, a plurality of the processors and the controller, the controller is respectively connected to the plurality of the processors, the controller is also connected to a debugging device of the computing acceleration module through the debugging interface, the debugging request is a request sent by the debugging device received by the debugging interface, and the debugging request is used for requesting debugging of the operating performance of the plurality of the processors;
[0123] A sending module, used for sending the corresponding target control instruction to each processor, wherein the processor is used for executing the received corresponding target control instruction to obtain a corresponding target execution result;
[0124] A processing module, configured to receive the target execution results fed back by the multiple processors, and convert the multiple target execution results into target debugging results corresponding to the debugging request;
[0125] A feedback module is used to call the debugging interface to feed back the target debugging result to the debugging device.
[0126] Through the above device, by deploying a controller between the debugging interface of the computing acceleration module and multiple processors, the controller can be connected to the debugging device of the computing acceleration module through the debugging interface. When the debugging interface receives a debugging request from the debugging device for requesting the operating performance of multiple processors, the controller can receive the debugging request and convert it into a target control instruction for each processor, and send the corresponding target control instruction to each processor. Then, after receiving the corresponding target control instruction, the processor sends the corresponding target execution result obtained to the controller, and the controller can convert the received multiple target execution results into target debugging results corresponding to the debugging request, and send the target debugging results to the debugging device. Through the above method, it is possible to debug multiple processors deployed in the computing acceleration module through a debugging interface at the same time, avoid the interface occupancy on the computing acceleration module when debugging the computing performance of the computing acceleration module, and ensure the computing performance of the computing acceleration module. Therefore, the problem of low efficiency in regulating the computing performance of the computing acceleration module in the related art can be solved, and the effect of improving the efficiency of regulating the computing performance of the computing acceleration module can be achieved.
[0127] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0128] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0129] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0130] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0131] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.
[0132] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0133] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A computing acceleration module, characterized in that: include: A debugging interface, a plurality of processors and a controller, wherein the controller is connected to the plurality of processors respectively, and the controller is also used to connect to the debugging device of the computing acceleration module through the debugging interface; The debugging interface is used to receive a debugging request from a debugging device for the operating performance of the plurality of processors; The controller is used to convert a target control instruction for each of the processors according to the debugging request; Sending the corresponding target control instruction to each of the processors; The processor is used to execute the received target control instruction to obtain a target execution result; and send the target execution result to the controller; The controller is further configured to receive the target execution results fed back by the multiple processors, and convert the multiple target execution results into target debugging results corresponding to the debugging request; The debugging interface is called to feed back the target debugging result to the debugging device.
2. The computing acceleration module according to claim 1, characterized in that: The controller comprises an instruction configuration component, and the instruction configuration component is respectively connected to the debugging interface and the plurality of processors; The instruction configuration component is used to extract target operation information carried in the debugging request, wherein the target operation information is used to indicate the debugging operation to be performed by each processor when executing the debugging project requested by the debugging request; and convert the target control instruction for each processor according to the target operation information.
3. The computing acceleration module according to claim 2, characterized in that: The instruction configuration component is also used to locate identification information carried by the target operation information, wherein the identification information is used to identify the processor; search the target operation information for a target operation field that has a binding relationship with the identification information, wherein the target operation field is used to indicate the corresponding debugging operation to be performed by the processor; and determine the target control instruction corresponding to the target operation field from the operation fields and control instructions that have a corresponding relationship.
4. The computing acceleration module according to claim 2, characterized in that: The instruction configuration component is further used to extract a target operation field carried in the target operation information, wherein the target operation field is used to indicate a debugging operation to be performed by the processor; determine a reference control instruction corresponding to the target operation field from operation fields and control instructions having a corresponding relationship; and configure the reference control instruction as the target control instruction of each of the processors.
5. The computing acceleration module according to claim 1, characterized in that: The controller further comprises a data conversion component and a plurality of data transmission channels, wherein the data conversion component is respectively connected to the debug interface and the plurality of data transmission channels, the data transmission channels are configured in a one-to-one correspondence with the processors, and the data transmission channels are connected to the corresponding processors; The data transmission channel is used to receive the target execution result reported by the corresponding processor; The data conversion component is used to extract candidate execution results from the target execution results received on each data transmission channel according to data requirement information of the debugging device, wherein the data requirement information is used to indicate the debugging device's requirement for the processor execution result when executing the debugging item corresponding to the debugging request; and mark the candidate execution result using the channel identifier of the data transmission channel to which the candidate execution result belongs to obtain the target debugging result.
6. The computing acceleration module according to claim 5, characterized in that: The data transmission channel includes a buffer and a data transmission port, the buffer is connected to the data transmission port and the data conversion component respectively, and the data transmission port is also connected to the processor corresponding to the data transmission channel; The data transmission port is used to send the target execution result to the buffer after receiving the target execution result; The buffer is used to cache the received target execution results in a chronological order; The data conversion component is used to extract the candidate execution results of the target data demand from the target execution results cached in the buffer in a chronological order; The candidate execution result of the target data requirement is marked using the channel identifier to obtain the target debugging result, wherein the data requirement information includes the target data requirement.
7. A debugging method for a computing acceleration module, characterized in that: The method is applied to a controller deployed in a computing acceleration module, and the method comprises: Converting target control instructions for each processor according to the debugging request, wherein the computing acceleration module includes a debugging interface, a plurality of the processors and the controller, the controller is respectively connected to the plurality of the processors, the controller is also connected to the debugging device of the computing acceleration module through the debugging interface, the debugging request is a request sent by the debugging device received by the debugging interface, and the debugging request is used to request debugging of the operating performance of the plurality of the processors; Sending the corresponding target control instruction to each processor, wherein the processor is used to execute the received corresponding target control instruction to obtain a corresponding target execution result; receiving the target execution results fed back by the multiple processors, and converting the multiple target execution results into target debugging results corresponding to the debugging request; The debugging interface is called to feed back the target debugging result to the debugging device.
8. A debugging device for a computing acceleration module, characterized in that: The device is applied to a controller deployed in a computing acceleration module, and the device includes: A conversion module, used for converting a target control instruction for each processor according to a debugging request, wherein the computing acceleration module includes a debugging interface, a plurality of the processors and the controller, the controller is respectively connected to the plurality of the processors, and the controller is also connected to a debugging device of the computing acceleration module through the debugging interface, the debugging request is a request sent by the debugging device received by the debugging interface, and the debugging request is used to request debugging of the operating performance of the plurality of the processors; a sending module, used for sending the corresponding target control instruction to each of the processors, wherein the processor is used to execute the received corresponding target control instruction to obtain a corresponding target execution result; A processing module, configured to receive the target execution results fed back by the multiple processors, and convert the multiple target execution results into target debugging results corresponding to the debugging request; A feedback module is used to call the debugging interface to feed back the target debugging result to the debugging device.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in claim 7 when executed by a processor.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in claim 7 are implemented.
Citation Information
Patent Citations
Method, system and dispatcher for simulating multiple processors in parallel
CN102331961A
Accelerator card debugging method and device, computer equipment and storage medium
CN117349091A
Switching circuit for test diagnosis, logic device and circuit board card
CN118349410A
Acceleration module debugging system, method and equipment and computer readable storage medium
CN119201726A