Evaluation method and device of processing unit on DPU, equipment and medium
By obtaining data from the DPU's processing unit hardware table entries, classification and calculation, the mis-estimation and misest estimation problems of DPU network card delay evaluation are solved, and the accuracy and efficiency of processing unit performance evaluation is achieved, which is suitable for various application scenarios.
Patent Information
- Application Number
- CN202510389656.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the prior art, there is incorrect or missed delay evaluation of DPU network cards, resulting in the inability to accurately evaluate the performance of the processing unit and the performance evaluation needs of various application scenarios cannot be met.
By obtaining recorded data from the hardware table entries corresponding to the processing unit of the DPU, including message processing type identification, start processing time stamp and end processing time stamp, classification and calculation, the timestamp difference value of each service processing stage is obtained, and the average processing is performed to obtain the target delay.
It improves the accuracy and efficiency of the delay calculation of the DPU processing unit, ensures the accuracy of performance evaluation, and is suitable for various application scenarios.
Smart Images

Figure CN120295877A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of cloud computing technology, and in particular, to an evaluation method, device, equipment and medium for a processing unit on a DPU (Data Processing Unit). Background Art
[0002] With the rapid development of cloud computing technology, as a brand-new network card device, the DPU's hardware offloading technology can handle large-scale data tasks in the cloud environment, transfer data processing tasks from the CPU (Central Processing Unit) to the DPU hardware accelerator, thus liberating the data processing tasks from the CPU, improving system performance and efficiency, and reducing the latency of data processing at the same time. Therefore, the DPU will gradually replace traditional network cards and become the network device installed on the host in the cloud computing environment.
[0003] In related technologies, the overall latency of packets entering and leaving the DPU network card is counted to evaluate the overall performance of the DPU. Therefore, the latency inside a certain processing unit on the DPU can only estimate a rough data based on the approximate workload of each processing unit, resulting in the phenomenon of misestimation or underestimation, making it difficult to guarantee accuracy, and thus unable to meet the performance evaluation requirements of a certain processing unit on the DPU for various application scenarios. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an evaluation method, device, equipment and medium for a processing unit on a DPU.
[0005] An embodiment of the present disclosure provides an evaluation method for a processing unit on a DPU. The DPU includes at least one processing unit, and the method includes:
[0006] Obtaining n pieces of recorded data from the hardware entries corresponding to the processing units of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a packet processing type identifier, start processing timestamps and end processing timestamps of each service processing stage corresponding to the packet processing type identifier in the processing unit;
[0007] Classifying the n pieces of data according to the packet processing type identifier to obtain m pieces of data corresponding to each packet processing type identifier; where m is a positive integer less than or equal to n;
[0008] Calculating based on the start processing timestamps and end processing timestamps of each service processing stage corresponding to the packet processing type identifier in the m pieces of data to obtain the timestamp difference of each service processing stage corresponding to the packet processing type identifier;
[0009] Perform averaging processing based on the time - stamp differences of each of the business - processing phases to obtain the target latency of each business - processing phase of the message - processing type corresponding to each message - processing type identifier in the processing unit.
[0010] An embodiment of the present disclosure also provides an evaluation device for a processing unit on a DPU. The DPU includes at least one processing unit, and the device includes:
[0011] An acquisition module, configured to obtain n pieces of recorded data from the hardware table entries corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message - processing type identifier, the start - processing time - stamp and the end - processing time - stamp of each business - processing phase corresponding to the message - processing type identifier in the processing unit;
[0012] A classification module, configured to classify the n pieces of data according to the message - processing type identifier to obtain m pieces of data corresponding to each message - processing type identifier; where m is a positive integer less than or equal to n;
[0013] A calculation module, configured to calculate based on the start - processing time - stamp and the end - processing time - stamp of each business - processing phase corresponding to the message - processing type identifier in the m pieces of data to obtain the time - stamp difference of each business - processing phase corresponding to the message - processing type identifier;
[0014] A processing module, configured to perform averaging processing based on the time - stamp differences of each business - processing phase to obtain the target latency of each business - processing phase of the message - processing type corresponding to each message - processing type identifier in the processing unit.
[0015] An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the evaluation method of the processing unit on the DPU provided in the embodiment of the present disclosure.
[0016] An embodiment of the present disclosure also provides a computer - readable storage medium, where the storage medium stores a computer program, and the computer program is used to execute the evaluation method of the processing unit on the DPU provided in the embodiment of the present disclosure.
[0017] An embodiment of the present disclosure also provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements the evaluation method of the processing unit on the DPU described in the foregoing aspect.
[0018] The technical solutions provided in the embodiments of the present disclosure have the following advantages compared with the prior art: The evaluation solution for the processing unit on the DPU provided in the embodiments of the present disclosure obtains n pieces of recorded data from the hardware entries corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, start processing timestamps and end processing timestamps of each service processing stage corresponding to the message processing type identifier in the processing unit; classifying the n pieces of data according to the message processing type identifier to obtain m pieces of data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n; calculating based on the start processing timestamps and end processing timestamps of each service processing stage corresponding to the message processing type identifier in the m pieces of data to obtain the timestamp difference of each service processing stage corresponding to the message processing type identifier; performing an averaging process based on the timestamp differences of each service processing stage to obtain the target latency of each service processing stage of the message processing type corresponding to each message processing type identifier in the processing unit. Thus, it is possible to determine the latency corresponding to each service processing stage of different message processing types of the processing unit based on the data in the hardware entries corresponding to the processing unit, ensuring the accuracy and efficiency of latency calculation, and being able to determine the final latency by averaging the timestamp differences of each service processing stage of the same message processing type based on multiple pieces of data, further ensuring the accuracy of the latency, so as to improve the accuracy of the performance evaluation of the processing unit on the DPU in various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0020] Figure 1 It is a schematic flowchart of an evaluation method for a processing unit on a DPU provided in an embodiment of the present disclosure;
[0021] Figure 2 It is a schematic flowchart of another evaluation method for a processing unit on a DPU provided in an embodiment of the present disclosure;
[0022] Figure 3 It is an example diagram of another evaluation of a processing unit on a DPU provided in an embodiment of the present disclosure;
[0023] Figure 4 It is a schematic structural diagram of an evaluation device for a processing unit on a DPU provided in an embodiment of the present disclosure;
[0024] Figure 5 It is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0026] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0027] As used herein, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0028] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0029] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0030] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0031] Specifically, the DPU is a data processor and also a major category of newly developed dedicated processors. It is the third important computing power chip in the data center scenario after the CPU and GPU (Graphics Processing Unit), and provides a computing engine for high-bandwidth, low-latency, and data-intensive computing scenarios. The DPU network card is a network card equipped with a DPU data processor, which provides high-throughput and low-latency data processing capabilities for the server.
[0032] Currently, there is a phenomenon of misestimation or underestimation in roughly evaluating the latency statistics of a certain processing unit on the DPU through the overall latency of the DPU and the workload of a certain processing unit on the DPU, making it difficult to guarantee accuracy, and thus unable to meet the performance evaluation requirements of a certain processing unit on the DPU for various application scenarios.
[0033] In view of the above problems, the evaluation method for the processing unit on the DPU provided by the embodiments of the present disclosure can determine the latency corresponding to different packet processing types of the processing unit based on the data in the hardware entry corresponding to the processing unit, ensuring the accuracy and efficiency of latency calculation, and being able to determine the final latency by averaging the timestamp differences of the same packet processing type based on multiple pieces of data, further ensuring the accuracy of the latency, so as to improve the accuracy of the performance evaluation of the processing unit on the DPU in various application scenarios.
[0034] Specifically, Figure 1 is a schematic flowchart of an evaluation method for a processing unit on the DPU provided by the embodiments of the present disclosure. This method can be executed by an evaluation device for the processing unit on the DPU, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. The DPU includes at least one processing unit, such as Figure 1 shown, and the method includes:
[0035] Step 101: Obtain n pieces of recorded data from the hardware entry corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a packet processing type identifier, the start processing timestamp and the end processing timestamp of each service processing stage corresponding to the packet processing type identifier in the processing unit.
[0036] In the embodiments of the present disclosure, the DPU includes one or more processing units, and each processing unit has a corresponding hardware entry, and the hardware entry is a table entry for storing and managing the processing unit.
[0037] It can be understood that when a packet enters a certain processing unit in the DPU and meets the sampling conditions, the start processing timestamp and the end processing timestamp of each packet processing type corresponding to each service processing stage are recorded as a piece of data and directly stored in the register, and all the data in the register is stored in the hardware entry corresponding to the processing unit at a certain frequency or after the packet processing ends. That is, each processing unit on the DPU pipeline can be configured separately, and each records its own data in its own hardware entry, without interfering with each other and with the minimum impact range.
[0038] In an embodiment of the present disclosure, one or more pieces of data, that is, n pieces of data, can be obtained from the hardware table entries of the processing unit after the sampling is completed or according to a certain evaluation period. It can be understood that the n pieces of data can be data of the same type of message processing or data of multiple types of message processing. Usually, the n pieces of data correspond to multiple different types of message processing.
[0039] In an embodiment of the present disclosure, the data includes a message processing type identifier, and the message processing type identifier can uniquely identify a type of message processing. In an embodiment of the present disclosure, the type of message processing can be determined according to each service processing stage. For example, including three service processing stages "parsing", "encapsulating", and "matching" to determine the message processing type as de-encapsulation. That is, during the process of the processing unit processing the message, the message processing type of the message can be determined according to one or more service processing stages.
[0040] Among them, the service processing stage can also include one or more of processing types such as message reception, message processing service, and message sending. Each piece of data also includes the start processing timestamp and the end processing timestamp of the message processing type identifier corresponding to each service processing stage in the processing unit. That is, each message processing type corresponds to the start processing time and the end processing timestamp of each service processing stage. For example, the start processing timestamp and the end processing timestamp of service A, the start processing timestamp and the end processing timestamp of service B, etc.
[0041] Step 102: Classify the n pieces of data according to the message processing type identifier to obtain m pieces of data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n.
[0042] In an embodiment of the present disclosure, the n pieces of data can include data with different message processing type identifiers. For example, n is 100, the message processing type identifier X corresponds to 50 pieces of data, the message processing type identifier Y corresponds to 20 pieces of data, and the message processing type identifier Z corresponds to 30 pieces of data. Thus, the 100 pieces of data are divided into three parts according to the three different message processing type identifiers. For example, the m pieces of data corresponding to the message processing type identifier X are 50 pieces of data. More specifically, service A has 50 pieces of data.
[0043] Step 103: Calculate based on the start processing timestamp and the end processing timestamp of the message processing type identifier corresponding to each service processing stage in the m pieces of data to obtain the timestamp difference of the message processing type identifier corresponding to each service processing stage.
[0044] In the embodiments of the present disclosure, for each of the m pieces of data, the start processing timestamp and the end processing timestamp corresponding to each service processing stage of the message processing type identifier are obtained, and the timestamp difference corresponding to each service processing stage of the message processing type identifier is obtained. Continuing with the example where the m pieces of data corresponding to the message processing type identifier X are 50 pieces of data, for each piece of data, for example, if the service processing stage is Service A, there are 50 timestamp differences, and there may also be services with less than 50 timestamp differences.
[0045] Step 104: Perform an averaging process based on the timestamp differences of each service processing stage to obtain the target latency of each service processing stage of the message processing type corresponding to each message processing type identifier in the processing unit.
[0046] In the embodiments of the present disclosure, continuing with the previous example, by averaging the 50 timestamp differences, the target latency of Service A in the service processing stage of the message processing type corresponding to the message processing type identifier X can be obtained. Similarly, the target latencies of each service processing stage of the message processing type corresponding to all message processing type identifiers in the processing unit can be obtained.
[0047] The evaluation scheme for the processing unit on the DPU provided by the embodiments of the present disclosure obtains n pieces of recorded data from the hardware entries corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes the message processing type identifier, the start processing timestamp and the end processing timestamp corresponding to each service processing stage of the message processing type identifier in the processing unit; classify the n pieces of data according to the message processing type identifier to obtain m pieces of data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n; calculate based on the start processing timestamp and the end processing timestamp corresponding to each service processing stage of the message processing type identifier in the m pieces of data to obtain the timestamp difference corresponding to each service processing stage of the message processing type identifier; perform an averaging process based on the timestamp differences of each service processing stage to obtain the target latency of each service processing stage of the message processing type corresponding to each message processing type identifier in the processing unit. Thus, it is possible to determine the latency of each service processing stage corresponding to different message processing types of the processing unit based on the data in the hardware entries corresponding to the processing unit, ensuring the accuracy and efficiency of latency calculation, and being able to determine the final latency of each service processing stage based on the averaging of the timestamp differences of the same message processing type in multiple pieces of data, further ensuring the accuracy of the latency, so as to improve the accuracy of the performance evaluation of the processing unit on the DPU in various application scenarios.
[0048] Figure 2 FIG. is a schematic flowchart of another evaluation method for the processing unit on the DPU provided by the embodiments of the present disclosure. The method includes:
[0049] Step 201: When a message enters any processing unit in the DPU, obtain the total number of received messages currently. When the total number of received messages currently is greater than or equal to a preset quantity threshold, obtain the message processing type of the message, and record the start processing timestamp and end processing timestamp corresponding to each service processing stage.
[0050] In the embodiments of the present disclosure, the quantity threshold can be selected and set according to actual needs. When the total number of received messages currently is greater than or equal to the preset quantity threshold, it means that data recording can start. Therefore, the receiving message timestamp, the timestamp when the message starts processing service A, the timestamp when the message ends processing service A, and the timestamp when the message sending ends can be started to be recorded.
[0051] In the embodiments of the present disclosure, the message processing type of the message can be determined according to each service processing stage corresponding to each message, and the start processing timestamp and end processing timestamp of each service processing stage corresponding to the message processing type are recorded.
[0052] Step 202: Generate a piece of data according to the message processing type of the message, the start processing timestamp and end processing timestamp of each service processing stage, and store it in the register of the DPU according to the message processing type identifier. After the processing unit finishes sending the message, obtain all pieces of data from the register and store them in the hardware entry corresponding to the processing unit of the DPU.
[0053] In some embodiments, obtain the total number of sampled data in the register, obtain the preset sampling threshold corresponding to the message. When the total number of samples is greater than or equal to the sampling threshold, stop storing the data corresponding to the message, and control the processing unit to stop recording data for the message.
[0054] Specifically, different messages can be set with different sampling thresholds to further meet personalized needs and ensure processing efficiency and effect. That is to say, the sampling frequency and sampling threshold can be flexibly configured according to actual needs. For example, sample every 100th message, and at most record 512 pieces.
[0055] Specifically, when a message enters a certain processing unit in the DPU and meets the sampling condition (such as the 100th message), record the timestamp information on the current DPU. Record the current message processing type during the processing of this processing unit, and record the current timestamp again when the message finishes processing in this processing unit; due to being sensitive to latency, in order to minimize the time-consuming error caused during sampling as much as possible, the processing of each sampling point should be as fast as possible. The records of these intermediate sampling points are first written into the register of the DPU. After all data sampling related to this message is completed, finally, the data of the sampling points recorded during this sampling process is summarized into one piece of data and recorded in the hardware entry in the DPU.
[0056] Thus, in a single processing unit of the DPU, timestamps and packet processing types of the overall and each service stage are recorded by sampling. Considering the high precision requirement of the DPU for sampled data, it is first written to a register (fast), and finally summarized and recorded (slow) to reduce errors. In addition, each processing unit of the DPU pipeline is configured separately, with the minimum influence range. The sampling frequency, sampling threshold, and sampling stage of each processing unit of the DPU pipeline can be flexibly customized to further meet personalized requirements.
[0057] Step 203: Obtain n pieces of recorded data from the hardware entry corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a packet processing type identifier, start processing timestamps and end processing timestamps of each service processing stage corresponding to the packet processing type identifier in the processing unit.
[0058] Step 204: Classify the n pieces of data according to the packet processing type identifier to obtain m pieces of data corresponding to each packet processing type identifier; where m is a positive integer less than or equal to n.
[0059] Step 205: Calculate based on the start processing timestamps and end processing timestamps of each service processing stage corresponding to the packet processing type identifier in the m pieces of data to obtain the timestamp difference of each service processing stage corresponding to the packet processing type identifier. Then, perform an averaging process on the timestamp differences of each service processing stage to obtain the target latency of each service processing stage corresponding to each packet processing type identifier in the processing unit.
[0060] In some embodiments, obtain the start processing timestamp and end processing timestamp of each service processing stage in each piece of the m pieces of data, calculate the difference between the start processing timestamp and end processing timestamp of each service processing stage to obtain the timestamp difference of each service processing stage, and then perform an averaging process on all the timestamp differences of each service processing stage to obtain the target latency corresponding to each service processing stage in the processing unit.
[0061] Specifically, each processing unit on the DPU pipeline records the packet processing type and the current timestamp information on the DPU system by sampling and records them in the hardware entry. After sampling is completed, the recorded data is extracted from the hardware entry and analyzed, classified according to different packet processing types, and the latency corresponding to each packet processing type is calculated.
[0062] That is to say, after sampling is completed, sampling can be stopped, and the recorded data can be extracted from the hardware table items. Then, the extracted data can be classified according to different message processing types. For each processing type, the difference in the timestamps of each entry and exit of the processing unit is calculated respectively, and the average value is calculated to obtain the average value of the delay, so as to obtain the complete delay statistics of different message processing types on the processing unit on the DPU.
[0063] Therefore, inside each processing unit on the DPU pipeline, you can flexibly record timestamp information before and after different processing according to your own needs, so as to flexibly perform statistical analysis on the latency differences in different stages of different data processing. This can be deployed in a production environment with little impact on business, and can evaluate the latency differences in different message processing by the processing units on the DPU.
[0064] As an example of a scenario, in the processing unit D on the DPU, before a complete round of processing begins, it is determined whether the sampling conditions are currently met: 1) If the sampling conditions are not met, each service processing is performed normally, and the next round of processing continues after the processing is completed; 2) If the sampling conditions are met, the timestamp 1 of the received message is recorded in the register, and the timestamp 2 is recorded in the register after the message is received. Before the message processes m services, the timestamp 3 is recorded in the register, and after the message processes m services, the timestamp 4 is recorded in the register. The processing type of the current message is also recorded in the register, and the timestamp 5 is recorded in the register before the message is sent from the processing unit D. After the message is sent, the timestamp 6 is recorded in the register. Finally, the 6 timestamp information recorded in this sampling process and the corresponding message processing type are summarized into a data record in the hardware table item in the DPU. At this point, the processing of this round of messages in the processing unit D is completed, and the next round of processing continues.
[0065] After the sampling of the above-mentioned processing unit D is completed, the sampling data of the processing unit D can be extracted from the hardware table of the DPU, and then each sampling data is analyzed: complete time: timestamp 6 minus timestamp 1; message reception time: timestamp 2 minus timestamp 1; m service processing time: timestamp 4 minus timestamp 3; message sending time: timestamp 6 minus timestamp 5.
[0066] Then classify each data by the message processing type, and calculate the average value of various time-consuming statistics. For example, there are 500 sampling records for message processing type t1. Add up the time consumption of q business processing (business processing stage) in these 500 records and calculate the average value. The time consumption of q business processing in the case of message processing type t1 in processing unit D is obtained.
[0067] Therefore, after calculating the various time-consuming operations within the processing unit of the DPU, high-precision and latency-sensitive performance data for the processing unit D's handling of different packets can be obtained, and corresponding optimization analysis can also be performed on the time-consuming processes as needed.
[0068] For example, as Figure 3 shown, the packet enters the processing unit, the timestamp when the processing unit receives the packet is recorded, the timestamp before the start of processing 1 is recorded, the timestamp after the end of processing 1 is recorded, until the timestamp before the start of processing p is recorded, the timestamp after the end of processing p is recorded. According to the aforementioned processing flow, the processing type of the packet can be determined, and the timestamp information of each service processing stage can be determined. Finally, the timestamp when the packet is sent out from the processing unit is recorded, and the packet is sent out from the processing unit.
[0069] Thus, inside each processing unit on the DPU pipeline, by sampling, data such as timestamp information at different stages is first recorded into the DPU's registers (with high efficiency), and finally summarized into the DPU's hardware table entries once to reduce errors. After sampling, the sampled data is manually extracted and classified to more accurately calculate the latency statistics of different stages of different data processing in the processing unit.
[0070] Figure 4 FIG. is a schematic structural diagram of an evaluation device for a processing unit on a DPU provided by an embodiment of the present disclosure. This device can be implemented by software and / or hardware and is generally integrated in an electronic device.
[0071] As Figure 4 shown, the DPU includes at least one processing unit, and this device includes:
[0072] An acquisition module 401, configured to obtain n pieces of recorded data from the hardware table entry corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a packet processing type identifier, start processing timestamps and end processing timestamps of each service processing stage corresponding to the packet processing type identifier in the processing unit;
[0073] A classification module 402, configured to classify the n pieces of data according to the packet processing type identifier to obtain m pieces of data corresponding to each packet processing type identifier; where m is a positive integer less than or equal to n;
[0074] A calculation module 403, configured to calculate based on the start processing timestamps and end processing timestamps of each service processing stage corresponding to the packet processing type identifier in the m pieces of data to obtain the timestamp difference of each service processing stage corresponding to the packet processing type identifier;
[0075] The processing module 404 is configured to perform averaging processing based on the time - stamp difference of each service processing stage, so as to obtain the target delay of each service processing stage corresponding to each message processing type identifier in the processing unit.
[0076] Optionally, the device further includes:
[0077] The acquisition quantity module is configured to acquire the total quantity of received messages currently when a message enters any one of the processing units in the DPU;
[0078] The recording module is configured to, when the total quantity of received messages currently is greater than or equal to a preset quantity threshold, acquire the message processing type of the message, and record the start - processing timestamp and end - processing timestamp corresponding to each service processing stage;
[0079] The generating module is configured to generate a piece of data according to the message processing type of the message, the start - processing timestamp and end - processing timestamp of each service processing stage, and store the data in the register of the DPU according to the message processing type identifier.
[0080] Optionally, the device further includes:
[0081] The storage module is configured to, after the processing unit finishes sending the message, acquire all pieces of data from the register and store them in the hardware entry corresponding to the processing unit of the DPU.
[0082] Optionally, the device further includes:
[0083] The sampling quantity acquisition module is configured to acquire the total sampling quantity of the sampled data in the register;
[0084] The threshold acquisition module is configured to acquire the preset sampling threshold corresponding to the message;
[0085] The control module is configured to, when the total sampling quantity is greater than or equal to the sampling threshold, stop storing the data corresponding to the message, and control the processing unit to stop recording data of the message.
[0086] The evaluation device of the processing unit on the DPU provided by the embodiments of the present disclosure can execute the evaluation method of the processing unit on the DPU provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0087] Figure 5 It is a schematic structural diagram of an electronic device including a DPU provided by an embodiment of the present disclosure. Specifically, refer to Figure 5, which shows a schematic structural diagram of the electronic device 500 suitable for implementing the embodiments of the present disclosure. The electronic device 500 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0088] As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 502 or the programs loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0089] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0090] Specifically, according to the embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, it executes the above-mentioned functions defined in the evaluation method of the DPU upper processing unit in the embodiments of the present disclosure.
[0091] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0092] The above-mentioned computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain n pieces of data recorded from the hardware entry corresponding to the processing unit of the DPU; classify the n pieces of data according to the message processing type identifier to obtain m pieces of data corresponding to each message processing type identifier; calculate based on the start processing timestamp and end processing timestamp corresponding to each service processing stage of the message processing type identifier in the m pieces of data to obtain the timestamp difference of each service processing stage corresponding to the message processing type identifier; perform an averaging process based on the timestamp differences of each service processing stage to obtain the target latency of each service processing stage of the message processing type corresponding to each message processing type identifier in the processing unit.
[0093] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0095] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0096] The functions described above herein may be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and the like.
[0097] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0098] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, including:
[0099] a processor;
[0100] a memory for storing executable instructions of the processor;
[0101] The processor is configured to read the executable instructions from the memory and execute the instructions to implement any one of the evaluation methods of the processing unit on the DPU provided by the present disclosure.
[0102] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium storing a computer program for executing any one of the evaluation methods of the processing unit on the DPU provided by the present disclosure.
[0103] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0104] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0105] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An evaluation method for processing units on a DPU, characterized in that The DPU includes at least one processing unit, and the method includes: Obtaining n pieces of recorded data from the hardware entry corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, start processing timestamps and end processing timestamps of each service processing stage corresponding to the message processing type identifier in the processing unit; Classifying the n pieces of data according to the message processing type identifier to obtain m pieces of data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n; Calculating based on the start processing timestamps and end processing timestamps of each service processing stage corresponding to the message processing type identifier in the m pieces of data to obtain the timestamp difference of each service processing stage corresponding to the message processing type identifier; Performing an averaging process based on the timestamp differences of each service processing stage to obtain the target latency of each service processing stage of each message processing type corresponding to the processing unit.
2. The evaluation method of the processing unit on the DPU according to claim 1, wherein The method further includes: When a message enters any one of the processing units in the DPU, obtaining the total number of received messages; When the total number of received messages is greater than or equal to a preset quantity threshold, obtaining the message processing type of the message, and recording the start processing timestamp and end processing timestamp corresponding to each service processing stage; Generating a piece of data according to the message processing type of the message, the start processing timestamps and end processing timestamps of each service processing stage, and storing the data in the register of the DPU according to the message processing type identifier.
3. The evaluation method of the processing unit on the DPU according to claim 2, characterized in that, The method further includes: After the processing unit finishes sending the message, obtaining all pieces of data from the register and storing the data in the hardware entry corresponding to the processing unit of the DPU.
4. The evaluation method of the processing unit on the DPU according to claim 2, wherein The method further includes: Obtaining the total number of sampled data in the register; Obtaining a preset sampling threshold corresponding to the message; When the total number of sampled data is greater than or equal to the sampling threshold, stopping storing the data corresponding to the message, and controlling the processing unit to stop recording data for the message.
5. An evaluation device for a processing unit on a DPU, characterized in that, The DPU includes at least one processing unit, and the apparatus includes: An obtaining module, configured to obtain n pieces of recorded data from the hardware entry corresponding to the processing unit of the DPU; where n is a positive integer greater than or equal to 1, and the data includes a message processing type identifier, start processing timestamps and end processing timestamps of each service processing stage corresponding to the message processing type identifier in the processing unit; A classifying module, configured to classify the n pieces of data according to the message processing type identifier to obtain m pieces of data corresponding to each message processing type identifier; where m is a positive integer less than or equal to n; A calculating module, configured to calculate based on the start processing timestamps and end processing timestamps of each service processing stage corresponding to the message processing type identifier in the m pieces of data to obtain the timestamp difference of each service processing stage corresponding to the message processing type identifier; A processing module, configured to perform averaging processing based on the time - stamp difference of each service processing stage, so as to obtain the target latency of each service processing stage corresponding to each message - processing type identifier in the processing unit.
6. The evaluation device for the processing unit on the DPU according to claim 5, characterized in that, The method further includes: An acquisition quantity module, configured to acquire the total quantity of received messages currently when a message enters any one of the processing units in the DPU; A recording module, configured to, when the total quantity of received messages currently is greater than or equal to a preset quantity threshold, acquire the message - processing type of the message, and record the start - processing time - stamp and end - processing time - stamp corresponding to each service processing stage; A generation module, configured to generate a piece of data according to the message - processing type of the message, the start - processing time - stamp and end - processing time - stamp of each service processing stage according to the message - processing type identifier, and store it in the register of the DPU.
7. The evaluation device for the processing unit on the DPU according to claim 6, characterized in that, The method further includes: A storage module, configured to, after the processing unit finishes sending the message, acquire all pieces of data from the register and store them in the hardware entry corresponding to the processing unit of the DPU.
8. The evaluation device for the processing unit on the DPU according to claim 6, characterized in that, The method further includes: A sampled - quantity acquisition module, configured to acquire the total sampled quantity of the sampled data in the register; A threshold acquisition module, configured to acquire a preset sampling threshold corresponding to the message; A control module, configured to, when the total sampled quantity is greater than or equal to the sampling threshold, stop storing the data corresponding to the message, and control the processing unit to stop recording data for the message.
9. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the evaluation method of the processing unit on the DPU according to any one of claims 1 - 4 above.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the evaluation method of the processing unit on the DPU according to any one of claims 1 - 4 above.
Citation Information
Patent Citations
Method and device for realizing network time synchronization
CN112615694A
Link information tracking method, system and service function
CN114205267A
Link detection method and device, intelligent network card and storage medium
CN116260743A
DPU financial data information analysis system time delay value determination method
CN118590416A
DPU, data transmission method based on DPU, storage medium and electronic equipment
CN118660060A
Cited By
EMMC performance difference test method and system based on three-level hardware timestamps
CN121506223A