Evaluation method, device and equipment of processing unit on dpu and medium

By acquiring data from the hardware entries of the DPU's processing unit and performing classification and calculation, the problem of miscalculation or undercalculation in DPU network card latency assessment is solved, thereby improving the accuracy and efficiency of processing unit performance assessment.

CN120295877BActive Publication Date: 2026-07-21YUSUR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUSUR TECH CO LTD
Filing Date
2025-03-31
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the overall latency assessment of DPU network cards suffers from miscalculation or undercalculation, resulting in the inability to accurately assess the performance of the processing unit and failing to meet the performance assessment requirements of various application scenarios.

Method used

The data recorded by obtaining the message processing type identifier, start processing timestamp, and end processing timestamp from the hardware table corresponding to the processing unit of the DPU is classified according to the message processing type identifier, and the timestamp difference of each business processing stage is calculated. Finally, the average is performed to determine the target latency.

Benefits of technology

It improves the accuracy and efficiency of DPU processing unit performance evaluation, ensures the precision of latency calculation, and is suitable for various application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295877B_ABST
    Figure CN120295877B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method and device for evaluating a processing unit on a DPU, and a medium. The DPU includes at least one processing unit. The method includes: obtaining n pieces of recorded data from a hardware table item corresponding to a processing unit of the DPU; classifying the n pieces of data according to a message processing type identifier, to obtain m pieces of data corresponding to each message processing type identifier; calculating a start processing timestamp and an end processing timestamp corresponding to each service processing stage based on the message processing type identifier in the m pieces of data, to obtain a timestamp difference of each service processing stage corresponding to the message processing type identifier; and performing average processing based on the timestamp difference of each service processing stage, to obtain a target delay of each service processing stage of a message processing type corresponding to each message processing type identifier in the processing unit. Thus, the delay of each service processing stage of different message processing types in each processing unit on the DPU can be accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of cloud computing technology, and in particular to a method, apparatus, device and medium for evaluating processing units on a DPU (Data Processing Unit). Background Technology

[0002] With the rapid development of cloud computing technology, the Data Processing Unit (DPU), as a new type of network interface card (NIC), utilizes hardware offloading technology to handle large-scale data-intensive tasks in cloud environments. It offloads data processing tasks from the CPU (Central Processing Unit) to the DPU hardware accelerator, thereby freeing up data processing resources from the CPU, improving system performance and efficiency, and reducing latency. Therefore, DPUs will gradually replace traditional NICs and become the primary network devices on hosts in cloud computing environments.

[0003] In related technologies, the overall latency of incoming and outgoing packets to and from the DPU network card is used to evaluate the overall performance of the DPU. However, the latency within a certain processing unit on the DPU can only be estimated based on the approximate workload of each processing unit, which may result in miscalculations or omissions, making it difficult to guarantee accuracy. Consequently, it cannot meet the performance evaluation requirements of a certain processing unit on the DPU for various application scenarios. Summary of the Invention

[0004] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides an evaluation method, apparatus, device, and medium for processing units on a DPU.

[0005] This disclosure provides an evaluation method for a processing unit on a DPU, wherein the DPU includes at least one processing unit, and the method includes:

[0006] Obtain n data records from the hardware table entry corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, a start processing timestamp and an end processing timestamp of each service processing stage corresponding to the message processing type identifier in the processing unit;

[0007] The n data entries are classified according to the message processing type identifier to obtain m data entries corresponding to each message processing type identifier; where m is a positive integer less than or equal to n.

[0008] The timestamp difference for each business processing stage corresponding to the message processing type identifier in the m data is calculated based on the start and end timestamps of each business processing stage.

[0009] The target latency for each business processing stage is obtained by averaging the timestamp differences of each of the business processing stages corresponding to each message processing type identifier in the processing unit.

[0010] This disclosure also provides an evaluation apparatus for a processing unit on a DPU, the DPU including at least one processing unit, the apparatus comprising:

[0011] The acquisition module is used to acquire n records of data from the hardware table entry corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, a start processing timestamp and an end processing timestamp of each service processing stage corresponding to the message processing type identifier in the processing unit;

[0012] The classification module is used to classify the n data according to the message processing type identifier to obtain m data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n;

[0013] The calculation module is used to calculate the timestamp difference between each business processing stage corresponding to the message processing type identifier in the m data based on the start processing timestamp and end processing timestamp of each business processing stage.

[0014] The processing module is used to perform average processing based on the timestamp difference of each of the service processing stages to obtain the target latency of each service processing stage of the message processing type corresponding to each message processing type identifier in the processing unit.

[0015] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the evaluation method for a processing unit on a DPU as provided in this disclosure.

[0016] This disclosure also provides a computer-readable storage medium storing a computer program for executing an evaluation method for a processing unit on a DPU as provided in this disclosure.

[0017] This disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the evaluation method for the processing unit on the DPU described in the preceding aspect.

[0018] The technical solution provided in this disclosure has the following advantages compared with the prior art: The evaluation scheme for the processing unit on the DPU provided in this disclosure obtains n data records from the hardware table corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, the start processing timestamp and end processing timestamp of each service processing stage corresponding to the message processing type identifier in the processing unit; the n data are classified according to the message processing type identifier to obtain m data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n; the timestamp difference of each service processing stage corresponding to the message processing type identifier is calculated based on the start processing timestamp and end processing timestamp of each service processing stage corresponding to the message processing type identifier in the m data; the average processing is performed based on the timestamp difference of each service processing stage to obtain the target latency of each service processing stage of each message processing type corresponding to each message processing type identifier in the processing unit. Therefore, the latency corresponding to each business processing stage of different message processing types of the processing unit can be determined based on the data in the hardware table corresponding to the processing unit, ensuring the accuracy and efficiency of latency calculation. Furthermore, the final latency can be determined by averaging the timestamp differences of each business processing stage of the same message processing type for multiple data items, further ensuring the accuracy of latency. This improves the accuracy of performance evaluation of the processing units on the DPU in various application scenarios. Attached Figure Description

[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0020] Figure 1 A flowchart illustrating an evaluation method for a processing unit on a DPU provided in this embodiment of the present disclosure;

[0021] Figure 2 A flowchart illustrating another method for evaluating a processing unit on a DPU provided in an embodiment of this disclosure;

[0022] Figure 3 Example diagram for evaluating another processing unit on a DPU provided in this disclosure embodiment;

[0023] Figure 4 A schematic diagram of the structure of an evaluation device for a processing unit on a DPU provided in an embodiment of this disclosure;

[0024] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0031] Specifically, DPU stands for Data Processor, a major category of newly developed dedicated processors. Following CPUs and GPUs (Graphics Processing Units), it is the third crucial computing chip in data center scenarios, providing a computing engine for high-bandwidth, low-latency, and data-intensive computing environments. A DPU network interface card (NIC) is a NIC equipped with a DPU data processor, providing servers with high-throughput, low-latency data processing capabilities.

[0032] Currently, the method of roughly estimating the latency of a specific processing unit on the DPU by considering the overall latency of the DPU and the workload of that unit is prone to miscalculation or underestimation, making it difficult to guarantee accuracy. This results in a technical problem that fails to meet the performance evaluation requirements of a specific processing unit on the DPU for various application scenarios.

[0033] To address the aforementioned issues, the evaluation method for processing units on a DPU provided in this disclosure can determine the latency corresponding to different message processing types of the processing unit based on the data in the hardware entries corresponding to the processing unit, ensuring the accuracy and efficiency of latency calculation. Furthermore, it can determine the final latency by averaging the timestamp differences of multiple data entries of the same message processing type, further ensuring the accuracy of latency. This improves the accuracy of performance evaluation of processing units on a DPU in various application scenarios.

[0034] Specifically, Figure 1 This is a flowchart illustrating an evaluation method for a processing unit on a DPU provided in an embodiment of the present disclosure. The method can be executed by an evaluation device for the processing unit on the DPU, wherein the device can be implemented in software and / or hardware, and is generally integrated into an electronic device. The DPU includes at least one processing unit, such as... Figure 1 As shown, the method includes:

[0035] Step 101: Obtain n records of data from the hardware table corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes the message processing type identifier, the start processing timestamp and end processing timestamp of each business processing stage corresponding to the message processing type identifier in the processing unit.

[0036] In this embodiment of the disclosure, the DPU includes one or more processing units, each processing unit having a corresponding hardware entry, which is an entry for storing and managing the processing unit.

[0037] Understandably, when a message enters a processing unit in the DPU and meets the sampling conditions, the start and end timestamps of each business processing stage corresponding to each message processing type are recorded as a data entry and directly stored in a register. At a certain frequency or after the message processing is completed, all the data in the register is stored in the hardware table entry corresponding to the processing unit. That is, each processing unit in the DPU pipeline can be configured independently and recorded in its own hardware table entry, without interfering with each other and with minimal impact.

[0038] In this embodiment of the disclosure, one or more data entries, i.e. n data entries, can be obtained from the hardware table of the processing unit after sampling or according to a certain evaluation cycle. It is understood that the n data entries can be data of the same message processing type or data of multiple message processing types. Typically, the n data entries correspond to multiple different message processing types.

[0039] In this embodiment of the disclosure, the data includes a message processing type identifier, which can uniquely identify a message processing type. In this embodiment of the disclosure, the message processing type can be determined according to each business processing stage. For example, it includes three business processing stages: "parsing", "encapsulation" and "matching" to determine the message processing type as decapsulation. That is, during the message processing process of the processing unit, the message processing type of the message can be determined according to one or more business processing stages.

[0040] The business processing stage may include one or more processing types such as message reception, message processing business, and message sending. Each piece of data also includes the start processing timestamp and end processing timestamp of each business processing stage corresponding to the message processing type identifier in the processing unit. That is, each message processing type corresponds to the start processing time and end processing timestamp of each business processing stage, such as the start processing timestamp and end processing timestamp of business A, the start processing timestamp and end processing timestamp of business B, etc.

[0041] Step 102: Classify the n data according to the message processing type identifier to obtain m data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n.

[0042] In this embodiment of the disclosure, n data entries may include data with different message processing type identifiers. For example, if n is 100, message processing type identifier X corresponds to 50 data entries, message processing type identifier Y corresponds to 20 data entries, and message processing type Z corresponds to 30 data entries. Thus, the 100 data entries are divided into three parts according to the three different message processing type identifiers. For example, the m data entries corresponding to message processing type identifier X are 50 data entries. More specifically, service A has 50 data entries.

[0043] Step 103: Calculate the timestamp difference for each business processing stage based on the start and end timestamps of the message processing type identifier in each of the m data entries.

[0044] In this embodiment of the disclosure, for each of the m data entries, the message processing type identifier corresponds to the start processing timestamp and end processing timestamp of each business processing stage, and the timestamp difference corresponding to each business processing stage is obtained. Taking the m data entries corresponding to the message processing type identifier X as 50 data entries as an example, each data entry may have 50 timestamp differences for example, business processing stage A, and there may also be businesses with less than 50 timestamp differences.

[0045] Step 104: Averaging the timestamp differences of each business processing stage to obtain the target latency for each business processing stage of the message processing type corresponding to each message processing type identifier in the processing unit.

[0046] In this embodiment of the disclosure, taking the aforementioned example again, by averaging the differences of 50 timestamps, the target latency of service A in the service processing stage of the message processing type corresponding to message processing type identifier X can be obtained. Similarly, the target latency of each service processing stage of the message processing type corresponding to all message processing type identifiers in the processing unit can be obtained.

[0047] The evaluation scheme for the processing unit on the DPU provided in this embodiment obtains n data records from the hardware table corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, the start processing timestamp and end processing timestamp of each service processing stage corresponding to the message processing type identifier in the processing unit; the n data records are classified according to the message processing type identifier to obtain m data records corresponding to each message processing type identifier; where m is a positive integer less than or equal to n; the timestamp difference of each service processing stage corresponding to the message processing type identifier is calculated based on the start processing timestamp and end processing timestamp of each service processing stage corresponding to the message processing type identifier in the m data records; the average processing is performed based on the timestamp difference of each service processing stage to obtain the target latency of each service processing stage corresponding to each message processing type identifier in the processing unit. Therefore, the latency of each business processing stage corresponding to different message processing types of the processing unit can be determined based on the data in the hardware table corresponding to the processing unit, ensuring the accuracy and efficiency of latency calculation. Furthermore, the final latency of each business processing stage can be determined by averaging the timestamp differences of each business processing stage for the same message processing type of multiple data, further ensuring the accuracy of latency. This improves the accuracy of performance evaluation of the processing units on the DPU in various application scenarios.

[0048] Figure 2 A flowchart illustrating another method for evaluating a processing unit on a DPU provided in this disclosure embodiment, the method comprising:

[0049] Step 201: When a message enters any processing unit in the DPU, obtain the total number of messages currently received. If the total number of messages currently received is greater than or equal to a preset threshold, obtain the message processing type of the message and record the start processing timestamp and end processing timestamp corresponding to each business processing stage.

[0050] In this embodiment of the disclosure, the quantity threshold can be selected and set according to actual needs. When the total number of currently received messages is greater than or equal to the preset quantity threshold, it means that data recording can begin. Therefore, the timestamps of received messages, the timestamps of when message processing of service A begins, the timestamps of when message processing of service A ends, and the timestamps of when message sending ends can be started.

[0051] In this embodiment of the disclosure, the message processing type of the message can be determined according to the various business processing stages corresponding to each message, and the start processing timestamp and end processing timestamp of each business processing stage corresponding to the message processing type can be recorded.

[0052] Step 202: Generate a data entry based on the message processing type, start processing timestamp, and end processing timestamp of each service processing stage according to the message processing type identifier, and store it in the DPU register. After the processing unit finishes sending the message, retrieve all data entries from the register and store them in the hardware table entry corresponding to the processing unit of the DPU.

[0053] In some embodiments, the total number of samples of the sampled data in the register is obtained, the preset sampling threshold corresponding to the message is obtained, and when the total number of samples is greater than or equal to the sampling threshold, the storage of the data corresponding to the message is stopped, and the processing unit is controlled to stop recording data for the message.

[0054] Specifically, different sampling thresholds can be set for different messages to further meet personalized needs and ensure processing efficiency and effectiveness. In other words, the sampling frequency and sampling threshold can be flexibly configured according to actual needs, such as sampling every 100th message, with a maximum of 512 messages recorded.

[0055] Specifically, when a message enters a processing unit in the DPU and meets the sampling conditions (e.g., the 100th message), the timestamp information on the current DPU is recorded. During the processing of the message in that unit, the current message processing type is recorded. When the message finishes processing in that unit, the current timestamp is recorded again. Due to the sensitivity to latency, in order to minimize the time-consuming errors caused by the sampling process, the processing of each sampling point should be as fast as possible. The records of these intermediate sampling points are first written to the DPU's register. After all the data related to this message has been sampled, the data of the sampling points recorded during this sampling process are finally summarized into a single data record and recorded in the hardware table of the DPU.

[0056] Therefore, in a single processing unit on the DPU, the timestamps and message processing types of the whole and each business stage are recorded by sampling. Due to the high accuracy requirements of the DPU for the sampled data, the process of writing to the register first (fast) and summarizing and recording at the last time (slow) is adopted to reduce errors. In addition, each processing unit of the DPU pipeline is configured separately, minimizing the impact range. The sampling frequency, sampling threshold and sampling stage of each processing unit of the DPU pipeline can be flexibly customized to further meet personalized needs.

[0057] Step 203: Obtain n records of data from the hardware table corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes the message processing type identifier, the start processing timestamp and end processing timestamp of each business processing stage corresponding to the message processing type identifier in the processing unit.

[0058] Step 204: Classify the n data according to the message processing type identifier to obtain m data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n.

[0059] Step 205: Based on the start and end timestamps of each business processing stage corresponding to the message processing type identifier in the m data, calculate the timestamp difference for each business processing stage corresponding to the message processing type identifier. Then, average the timestamp differences for each business processing stage to obtain the target latency for each business processing stage of each message processing type corresponding to the message processing type identifier in the processing unit.

[0060] In some embodiments, the start and end timestamps of each business processing stage in each of the m data are obtained. The difference between the start and end timestamps of each business processing stage is calculated to obtain the timestamp difference of each business processing stage. Then, all timestamp differences of each business processing stage are averaged to obtain the target latency corresponding to each business processing stage in the processing unit.

[0061] Specifically, each processing unit in the DPU pipeline records the message processing type and the current timestamp information on the DPU system through sampling, and records it in the hardware table. After sampling, the recorded data is extracted from the hardware table and analyzed, classified according to different message processing types, and the latency corresponding to each message processing type is calculated.

[0062] In other words, after sampling is completed, sampling can be stopped, the recorded data can be extracted from the hardware table entries, and then the extracted data can be classified according to different message processing types. For each processing type, the timestamp difference between the entry and exit of each message in the processing unit can be calculated, and the average value can be calculated to obtain the average delay. This gives the complete delay statistics for different message processing types on the DPU.

[0063] Therefore, within each processing unit of the DPU pipeline, timestamp information can be flexibly recorded before and after different processes according to their own needs. This allows for flexible statistical analysis of latency differences at different stages of different data processing. It can be deployed in a production environment with minimal business impact and can assess the latency differences of different message processing units on the DPU.

[0064] As an example scenario, in processing unit D on the DPU, before a complete round of processing begins, it is determined whether the current sampling conditions are met: 1) If the sampling conditions are not met, each service is processed normally, and the next round of processing continues after the processing is completed; 2) If the sampling conditions are met, the received message timestamp 1 is recorded in the register, the message reception timestamp 2 is recorded in the register after the message is received, the message before processing service m is recorded in the register, the message after processing service m is completed is recorded in the register, and the current message processing type is recorded in the register. Before the message is sent from processing unit D, the message timestamp 5 is recorded in the register, and the message after sending is completed is recorded in the register. Finally, the 6 timestamps recorded during this sampling process and the corresponding message processing type are summarized into a data record and recorded in the hardware table of the DPU. At this point, the processing of the message in this round in processing unit D is completed, and the next round of processing continues.

[0065] After the above processing unit D completes sampling, the sampled data of processing unit D can be extracted from the hardware table of DPU, and then each sampled data is analyzed: complete time consumption: timestamp 6 minus timestamp 1; message reception time consumption: timestamp 2 minus timestamp 1; m service processing time consumption: timestamp 4 minus timestamp 3; message transmission time consumption: timestamp 6 minus timestamp 5.

[0066] Then, each data item is classified according to the message processing type, and the average value of various time consumption statistics is calculated. For example, if there are 500 sample records for message processing type t1, the time consumption of q service processing (service processing stage) in these 500 records is accumulated and the average value is calculated to obtain the time consumption of q service processing in the case of message processing type t1 in processing unit D.

[0067] Therefore, after calculating the various processing times within the DPU, we obtain high-precision, latency-sensitive performance data of the processing unit D for different message processing, and we can also perform targeted optimization analysis for time-consuming processing.

[0068] For example, such as Figure 3 As shown, when a message enters the processing unit, the timestamp when the processing unit receives the message is recorded, the timestamp before processing 1 begins is recorded, the timestamp after processing 1 ends is recorded, and so on, until the timestamp before processing p begins is recorded and the timestamp after processing p ends is recorded. Based on the aforementioned processing flow, the processing type of the message can be determined, and the timestamp information of each business processing stage can be determined. Finally, the timestamp when the message is sent out from the processing unit is recorded, and the message is sent out from the processing unit.

[0069] Therefore, within each processing unit of the DPU pipeline, data such as timestamps at different stages are first recorded in the DPU's registers through sampling (high efficiency), and then finally summarized into the DPU's hardware table entries to reduce errors. After sampling, the sampled data is manually extracted and classified, which can more accurately calculate the latency statistics of different stages of different data processing in the processing unit.

[0070] Figure 4 This is a schematic diagram of an evaluation device for a processing unit on a DPU provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is generally integrated into an electronic device.

[0071] like Figure 4 As shown, the DPU includes at least one processing unit, and the device includes:

[0072] The acquisition module 401 is used to acquire n data records from the hardware table entry corresponding to the processing unit of the DPU; wherein n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, a start processing timestamp and an end processing timestamp of each service processing stage corresponding to the message processing type identifier in the processing unit;

[0073] The classification module 402 is used to classify the n data according to the message processing type identifier to obtain m data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n;

[0074] The calculation module 403 is used to calculate the timestamp difference between each business processing stage corresponding to the message processing type identifier in the m data based on the start processing timestamp and end processing timestamp of each business processing stage.

[0075] The processing module 404 is used to perform average processing based on the timestamp difference of each of the service processing stages to obtain the target latency of each service processing stage of the message processing type corresponding to each message processing type identifier in the processing unit.

[0076] Optionally, the device further includes:

[0077] The quantity acquisition module is used to acquire the total number of currently received messages when a message enters any of the processing units in the DPU.

[0078] The recording module is used to obtain the message processing type of the message when the total number of currently received messages is greater than or equal to a preset number threshold, and to record the start processing timestamp and end processing timestamp corresponding to each service processing stage.

[0079] The generation module is used to generate a data entry based on the message processing type of the message, the start processing timestamp and the end processing timestamp of each service processing stage, and store it in the register of the DPU.

[0080] Optionally, the device further includes:

[0081] The storage module is used to retrieve all data from the register and store them in the hardware table corresponding to the processing unit of the DPU after the processing unit has finished sending the message.

[0082] Optionally, the device further includes:

[0083] The sample count acquisition module is used to acquire the total number of samples of the sampled data in the register;

[0084] The threshold acquisition module is used to acquire the preset sampling threshold corresponding to the message;

[0085] The control module is used to stop storing the data corresponding to the message and control the processing unit to stop recording data on the message when the total number of samples is greater than or equal to the sampling threshold.

[0086] The evaluation apparatus for the DPU on-processing unit provided in this disclosure can execute the evaluation method for the DPU on-processing unit provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0087] Figure 5 This is a schematic diagram of an electronic device including a DPU, provided as an embodiment of the present disclosure. See below for details. Figure 5The diagram illustrates a structural schematic suitable for implementing the electronic device 500 in the embodiments of this disclosure. The electronic device 500 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0088] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0089] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0090] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from ROM 502. When the computer program is executed by processing device 501, it performs the functions defined in the evaluation method of the processing unit on the DPU of embodiments of this disclosure.

[0091] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0092] The aforementioned computer-readable medium carries one or more programs. When the aforementioned one or more programs are executed by the electronic device, the electronic device causes the following: it retrieves n data entries from the hardware table corresponding to the processing unit of the DPU; it classifies the n data entries according to the message processing type identifier to obtain m data entries corresponding to each message processing type identifier; it calculates the timestamp difference for each business processing stage corresponding to the message processing type identifier based on the start and end timestamps of each business processing stage corresponding to the message processing type identifier in the m data entries; and it averages the timestamp differences for each business processing stage to obtain the target latency for each business processing stage corresponding to each message processing type identifier in the processing unit.

[0093] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0095] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0096] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0097] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0098] According to one or more embodiments of this disclosure, this disclosure provides an electronic device, including:

[0099] processor;

[0100] Memory used to store the processor's executable instructions;

[0101] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the evaluation method of the processing unit on the DPU as provided in any of the present disclosure.

[0102] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium storing a computer program for performing an evaluation method for a processing unit on a DPU as described in any of the present disclosure.

[0103] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0104] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0105] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for evaluating processing units on a DPU, characterized in that, The DPU includes at least one processing unit, and the method includes: n data entries are obtained from the hardware table corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, a start processing timestamp and an end processing timestamp of each service processing stage corresponding to the message processing type identifier in the processing unit; wherein, each processing unit has a corresponding hardware table entry, which is an entry for storing and managing the processing unit. When a message enters any processing unit in the DPU and meets the sampling conditions, the start processing timestamp and end processing timestamp of each service processing stage corresponding to each message processing type are recorded as a data entry and directly stored in a register. At a certain frequency or after the message processing is completed, all data in the register is stored in the hardware table entry corresponding to the processing unit. The n data entries are classified according to the message processing type identifier to obtain m data entries corresponding to each message processing type identifier; where m is a positive integer less than or equal to n. Based on the start and end timestamps of each business processing stage corresponding to the message processing type identifier in the m data, the timestamp difference corresponding to each business processing stage of the message processing type identifier is calculated. The average time is calculated based on the timestamp difference of each of the service processing stages to obtain the target latency of each of the service processing stages of the message processing type corresponding to each message processing type identifier in the processing unit.

2. The evaluation method for processing units on a DPU according to claim 1, characterized in that, The method further includes: When a message enters any of the processing units in the DPU, the total number of messages currently received is obtained. When the total number of currently received messages is greater than or equal to a preset threshold, the message processing type of the message is obtained, and the start processing timestamp and end processing timestamp corresponding to each service processing stage are recorded. The message processing type, the start timestamp and end timestamp of each service processing stage are used to generate a data entry based on the message processing type identifier and stored in the register of the DPU.

3. The evaluation method for processing units on a DPU according to claim 2, characterized in that, The method further includes: After the processing unit finishes sending the message, it retrieves all data from the register and stores them in the hardware entry corresponding to the processing unit of the DPU.

4. The evaluation method for processing units on a DPU according to claim 2, characterized in that, The method further includes: Obtain the total number of samples of the sampled data in the register; Obtain the preset sampling threshold corresponding to the message; When the total number of samples is greater than or equal to the sampling threshold, the storage of the data corresponding to the message is stopped, and the processing unit is controlled to stop recording data for the message.

5. An evaluation apparatus for a processing unit on a DPU, characterized in that, The DPU includes at least one processing unit, and the device includes: The acquisition module is used to acquire n data records from the hardware table entries corresponding to the processing units of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, a start processing timestamp and an end processing timestamp corresponding to each service processing stage of the message processing type identifier in the processing unit; wherein each processing unit has a corresponding hardware table entry, which is an entry for storing and managing the processing units. When a message enters any processing unit in the DPU and meets the sampling conditions, the start processing timestamp and end processing timestamp corresponding to each service processing stage of each message processing type are recorded as a data record and directly stored in a register. At a certain frequency or after the message processing is completed, all data in the register is stored in the hardware table entry corresponding to the processing unit. The classification module is used to classify the n data according to the message processing type identifier to obtain m data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n; The calculation module is used to calculate the timestamp difference between each business processing stage corresponding to the message processing type identifier in the m data based on the start processing timestamp and end processing timestamp of each business processing stage. The processing module is used to perform average processing based on the timestamp difference of each of the service processing stages to obtain the target latency of each service processing stage of the message processing type corresponding to each message processing type identifier in the processing unit.

6. The evaluation apparatus for the processing unit on the DPU according to claim 5, characterized in that, The device further includes: The quantity acquisition module is used to acquire the total number of currently received messages when a message enters any of the processing units in the DPU. The recording module is used to obtain the message processing type of the message when the total number of currently received messages is greater than or equal to a preset number threshold, and to record the start processing timestamp and end processing timestamp corresponding to each service processing stage. The generation module is used to generate a data entry based on the message processing type of the message, the start processing timestamp and the end processing timestamp of each service processing stage, and store it in the register of the DPU.

7. The evaluation apparatus for the processing unit on the DPU according to claim 6, characterized in that, The device further includes: The storage module is used to retrieve all data from the register and store them in the hardware table corresponding to the processing unit of the DPU after the processing unit has finished sending the message.

8. The evaluation apparatus for the processing unit on the DPU according to claim 6, characterized in that, The device further includes: The sample count acquisition module is used to acquire the total number of samples of the sampled data in the register; The threshold acquisition module is used to acquire the preset sampling threshold corresponding to the message; The control module is used to stop storing the data corresponding to the message and control the processing unit to stop recording data on the message when the total number of samples is greater than or equal to the sampling threshold.

9. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the evaluation method of the processing unit on the DPU as described in any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the evaluation method of the processing unit on the DPU as described in any one of claims 1-4.