Network data monitoring method and related apparatus
Patent Information
- Application Number
- PCT/CN2025/136270
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2025-11-20
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025136270_01102026_PF_FP_ABST
Abstract
Description
A method and related device for monitoring network data
[0001] This application claims priority to Chinese Patent Application No. 202510356555.8, filed on March 25, 2025, entitled "A Method and Apparatus for Monitoring Network Data", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of communications, and in particular to a method and apparatus for monitoring network data. Background Technology
[0003] With the development of technologies such as data centers, cloud computing, and 5G networks, the scale and complexity of networks have increased dramatically. Traditional network data monitoring methods are inadequate for handling large-scale and highly dynamic network environments. In the new network architecture, traffic patterns change frequently and are difficult to predict, and network status has a more significant impact on real-time performance. This necessitates network monitoring technologies that can provide more granular and higher-precision real-time data.
[0004] To achieve fine-grained monitoring of network data, existing technologies propose a network data monitoring method based on inband network telemetry (INT). This method embeds telemetry data directly into data packets, acquiring the status information of each hop in the network in real time. Compared to traditional monitoring methods that collect data outside the data stream, inband network telemetry can embed telemetry information during data stream processing, directly monitoring each data packet.
[0005] However, existing in-band telemetry systems typically rely on the header node creating a telemetry tag header in the packet header. Then, intermediate network interface card (NIC) nodes in the forwarding path add probe information to the telemetry header. Finally, the last-hop NIC node reports the probe data from the telemetry header to the control node for analysis. This causes the telemetry tag header to continuously increase in size during processing, significantly amplifying the data compared to small packet traffic. Simultaneously, intermediate NIC nodes also need to perform packet disassembly and reassembly operations during packet processing, thus impacting their forwarding performance.
[0006] We hope there can be an improvement plan that can better achieve network data sampling, statistics, and information monitoring. Summary of the Invention
[0007] To achieve network data sampling, statistics, and information monitoring, this application provides a network data monitoring method and related apparatus, enabling more granular network data detection. Furthermore, this method eliminates the need for intermediate network interface cards (NICs) to carry additional monitoring information during processing, thus better achieving network data sampling, statistics, and information monitoring, and improving the reliability and real-time performance of data forwarding in data plane monitoring.
[0008] Firstly, this application provides a method for monitoring network data, the network including a first network node and a control node. The first network node includes a data processing unit. The method is applied to the data processing unit and includes: First, the data processing unit can receive a first data stream. The first data stream includes a telemetry head, which includes a period identifier indicating the period for reporting telemetry information to the control node. The first data stream is intended to be processed within the first network node. Then, the data processing unit can insert colored bits into the first data stream according to the telemetry head to obtain a second data stream with colored bits. Then, during the processing of the second data stream by at least one hardware unit in the first network node, the data processing unit can sample at least one hardware unit to obtain telemetry information corresponding to at least one hardware unit. Finally, the data processing unit can send the telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the hardware unit based on the telemetry data.
[0009] In this application, the data processing unit colors the first data stream and inserts the colored bits into it to obtain a second data stream. During the processing of the second data stream by the first network interface card (NIC) node, at least one hardware unit of the first NIC node is sampled to obtain telemetry data at each sampling location. Through this method, the control node can proactively sense subtle changes in the network and accurately reflect the packet loss and latency of at least one hardware unit of the first network node. This achieves fine-grained network data monitoring and improves the reliability and real-time performance of data plane monitoring and data forwarding.
[0010] In some embodiments, after receiving the first data stream, the method may further include: creating a data statistics table corresponding to the odd-numbered period based on the period identifier of the odd-numbered period; and creating a data statistics table corresponding to the even-numbered period based on the period identifier of the even-numbered period.
[0011] In some embodiments, during the processing of the second data stream by at least one hardware unit in the first network node, the method first utilizes at least one hardware unit in the first network node to process the second data stream. Then, during the processing of the second data stream, the at least one hardware unit is sampled to obtain coloring information for the at least one hardware unit. Next, the coloring information for odd-numbered periods is recorded in the data statistics table corresponding to the odd-numbered periods based on the period identifier of the odd-numbered periods, and the coloring information for even-numbered periods is recorded in the data statistics table corresponding to the even-numbered periods based on the period identifier of the even-numbered periods. Finally, based on the data statistics tables corresponding to the odd-numbered periods and the data statistics tables corresponding to the even-numbered periods, the telemetry information of the at least one hardware unit is obtained.
[0012] In this application, the dual-cycle stream identifier facilitates the data processing unit to process the coloring information of odd and even cycles separately. The data statistics table corresponding to the odd cycle is written in the odd cycle, while the data statistics table corresponding to the even cycle is read-only. This realizes that the data statistics tables corresponding to the odd and even cycles can be cyclically overwritten, reducing the on-chip resource overhead of the data processing unit. This method can improve the reliability and real-time performance of data plane monitoring data forwarding.
[0013] In some embodiments, the method can also delete the aforementioned data statistics table that has timed out according to a predetermined aging time for the entries, so as to optimize the memory resources of the first network node.
[0014] In this application, after the data statistics table for odd-numbered periods is uploaded, the entries in the odd-numbered period statistics table can be cleaned up during the reporting process of the data statistics table for even-numbered periods, thereby optimizing the memory resources of the first network node.
[0015] In some embodiments, telemetry information may include a hardware timestamp and the number of packets received corresponding to at least one hardware unit. In this case, the method may include: sending telemetry data to a control node according to a period indicated by a period identifier, so that the control node can perform network data latency statistics and packet loss detection on at least one hardware unit based on the hardware timestamp and the number of packets received.
[0016] In this application, the latency statistics and packet loss detection at the nanosecond level can be achieved by using the hardware timestamp and packet count corresponding to at least one hardware unit, thereby improving the accuracy and performance of network data sampling statistics and information monitoring.
[0017] In some embodiments, at least one hardware unit includes an output port of a first network node for connecting to a second network node. In this case, the method further includes: sending telemetry data from the output port to a control node according to a period indicated by a period identifier, so that the control node can perform network data fault monitoring on the second network node based on the telemetry data.
[0018] In this application, the telemetry data of the output port can be used as the basis for judging the network data monitoring of other network nodes. This method can realize the monitoring of network data between various network nodes, thereby improving the reliability and real-time performance of the data plane forwarding system.
[0019] In some embodiments, the hardware units in the first network node include any one of a parser, a processor, a data cache, an encryption / decryption unit, a decompression unit, and a CRC check unit.
[0020] In some embodiments, after obtaining the coloring information of at least one hardware unit, the method may further send the coloring information to a processor so that the processor can perform network data latency statistics on at least one hardware unit and the second network node based on the coloring information.
[0021] In this application, the task of monitoring latency statistics can be handled by the processor of the first network node, which reduces the computational load on the control node and improves the efficiency of network data detection.
[0022] Secondly, this application proposes a method for monitoring network data, which can be executed by a control node, including: acquiring telemetry data related to a second data stream in a first network node at a predetermined period. The second data stream is obtained by marking and coloring a first data stream sent by a client; and monitoring network data faults in at least one hardware unit of the first network node based on the telemetry data.
[0023] In this application, the control node can actively obtain telemetry data of at least one hardware unit of the first network node, and monitor the network data of the hardware unit based on the telemetry data, thereby realizing fine-grained network data monitoring and improving the network data detection efficiency.
[0024] Thirdly, this application proposes a network data monitoring device, which can be applied to a data processing unit including a coloring module, a sampling module, and a data reporting module. The coloring module receives a first data stream, wherein the first data stream includes a telemetry head, and the telemetry head includes a period identifier, which indicates the period for reporting telemetry information to the control node. The first data stream is intended to be processed in a first network node. The coloring module is also used to insert coloring bits into the first data stream according to the telemetry head to obtain a second data stream after coloring. The sampling module is used to sample at least one hardware unit in the first network node during the processing of the second data stream, obtaining telemetry information corresponding to the at least one hardware unit. The data reporting module is used to send the telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the hardware unit based on the telemetry data.
[0025] In some embodiments, the coloring module is further configured to create a data statistics table corresponding to an odd-numbered period based on the period identifier of the odd-numbered period, and to create a data statistics table corresponding to an even-numbered period based on the period identifier of the even-numbered period.
[0026] In some embodiments, the sampling module is further configured to sample at least one hardware unit during the second data stream processing to obtain coloring information for at least one hardware unit. Then, based on the period identifier of the odd-numbered period, the coloring information for the odd-numbered period is recorded in the data statistics table corresponding to the odd-numbered period, and based on the period identifier of the even-numbered period, the coloring information for the even-numbered period is recorded in the data statistics table corresponding to the even-numbered period. Finally, based on the data statistics tables corresponding to the odd-numbered and even-numbered periods, the telemetry information for at least one hardware unit is obtained.
[0027] In some embodiments, the sampling module is further configured to delete the aforementioned data statistics table that has timed out according to a predetermined table entry aging time, so as to optimize the memory resources of the first network node.
[0028] In some embodiments, the telemetry information includes a hardware timestamp and the number of packets received corresponding to at least one hardware unit. In this case, the data reporting module is further configured to send telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data latency statistics and packet loss detection for at least one hardware unit based on the hardware timestamp and the number of packets received.
[0029] In some embodiments, at least one hardware unit includes an output port of a first network node for connecting to a second network node. In this case, the data reporting module is further configured to send telemetry data from the output port to a control node according to a period indicated by a period identifier, so that the control node can perform network data fault monitoring on the second network node based on the telemetry data.
[0030] In some embodiments, when the second network node is a storage medium, the sampling module is further configured to receive a third data stream returned by the storage medium. This third data stream is intended to be processed in the first network node. Then, during the processing of the third data stream by at least one hardware unit in the first network node, the at least one hardware unit is sampled to obtain telemetry information corresponding to that hardware unit. Finally, the telemetry data is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the at least one hardware unit based on the telemetry data.
[0031] Fourthly, this application proposes a computing device including a data processing unit and a memory. The data processing unit is used to execute instructions stored in the memory, so that the data processing unit performs the network data monitoring method provided by the processor in the first aspect or any possible implementation thereof.
[0032] Fifthly, this application also provides a computer-readable storage medium including computer program instructions, which, when executed by a programmable logic device, can execute the network data monitoring method provided in the first aspect or any possible implementation thereof.
[0033] Sixthly, this application also provides a computer program product, including computer program instructions, which, when executed by a programmable logic device, can execute the network data monitoring method provided by the first aspect or any possible implementation thereof.
[0034] Any of the monitoring devices, computing devices, computer storage media, or computer program products provided above are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects of the corresponding solutions in the corresponding methods provided above, and will not be repeated here. Attached Figure Description
[0035] Figure 1 is a schematic diagram of an existing network data monitoring scenario;
[0036] Figure 2 is a schematic diagram of a network data detection scenario proposed in this application;
[0037] Figure 3 is a schematic diagram of a network data detection scenario inside the server;
[0038] Figure 4 is a flowchart illustrating a network data monitoring method proposed in this application;
[0039] Figure 5 is a schematic diagram of the transmission of the second data stream between the first network node and the second network node.
[0040] Figure 6 is a pipeline diagram of telemetry data reading of the second data stream when the control node processes packet loss detection service;
[0041] Figure 7 is a pipeline diagram of telemetry data reading of the second data stream when the control node processes latency statistics services;
[0042] Figure 8 is a schematic diagram of a packet loss detection scenario during the process of the control node processing data streams and writing them to the disk array;
[0043] Figure 9 is a flowchart illustrating a network data monitoring method proposed in this application;
[0044] Figure 10 is a schematic diagram of the structure of a network data monitoring device proposed in this application;
[0045] Figure 11 is a schematic diagram of the structure of a computing device provided in this application. Detailed Implementation
[0046] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0047] To facilitate understanding of the technical solution of this application, the relevant terms used in this document are explained below.
[0048] A smart network interface controller (NIC) is an advanced network adapter that integrates a dedicated processor, memory, and network interface. It is designed to improve network efficiency in data centers, cloud computing, and high-performance computing environments through hardware acceleration and intelligent features. Compared to traditional NICs, smart NICs incorporate microprocessors or dedicated chips, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), as well as data processing units (DPUs).
[0049] A Data Processing Unit (DPU) is a highly integrated hardware unit with powerful data processing capabilities, designed to efficiently integrate and optimize data processing, transmission, and monitoring tasks in data centers, cloud computing, and related network environments. Leveraging the network interface of a smart network interface card (NIC), the DPU can dynamically schedule network traffic according to preset strategies. For example, in complex network environments, it can monitor the demands of different service flows in real time, prioritizing critical data flows to improve overall network throughput and response speed. Based on these characteristics, the DPU can be closely integrated with in-band network telemetry (INT). For instance, by monitoring data packets hop-by-hop, the DPU accurately acquires performance data of each node in the network, collecting network status information during data processing and providing a basis for network optimization.
[0050] Inband Network Telemetry (INT) is a technology that embeds color bits in real time during network data processing to monitor network status information. Its core lies in using data packets to carry network status data, thereby achieving non-intrusive monitoring and analysis of network behavior.
[0051] Color bits are key identifiers used in network technology to mark traffic characteristics, mainly for network performance monitoring (such as packet loss and latency statistics) and traffic classification.
[0052] Packet loss statistics: The number of lost packets is calculated by comparing the difference between the number of packets entering and leaving the network within the same period (Pi-Pe).
[0053] Delay statistics: Calculate one-way or two-way delay by using the timestamp records of color bit transitions.
[0054] In traditional network architectures, network status information is typically collected using external detection devices, router or switch management interfaces, and traffic sampling.
[0055] External monitoring devices can include traffic collectors, monitoring devices implemented with Simple Network Management Protocol (SNMP), etc. These external monitoring devices are located at the network edge, independent of network data flow, and typically cannot reflect network status in real time. Furthermore, under high traffic conditions, monitoring data may be inaccurate or delayed.
[0056] The management interfaces of routers or switches may include those implemented using network management protocols such as SNMP, Sampled Flow (sFlow), and NetFlow. These management interfaces typically rely on protocol exchanges between devices, are subject to sampling errors, and cannot accurately reflect the state of specific traffic or data packets within the network.
[0057] Traffic sampling can involve periodically sampling network traffic using techniques such as NetFlow and sFlow, and generating flow records to analyze the health of the network architecture. However, traffic sampling methods cannot fully reflect the real-time status of the network, especially when faced with short-term network fluctuations or sudden traffic bursts, where the sampling accuracy may be greatly reduced.
[0058] In summary, traditional methods for monitoring network data often have the following limitations.
[0059] First, there is a significant delay in data acquisition and processing, making it impossible to provide real-time network status.
[0060] Second, because it relies on sampling, the monitoring results cannot cover all traffic, especially in high-traffic environments, where sampling may lose important data packets, resulting in insufficient detection accuracy.
[0061] Third, external monitoring equipment and protocols often require a lot of hardware and bandwidth support, which increases the cost of network management.
[0062] With the development of technologies such as data centers, cloud computing, and 5G networks, the scale and complexity of networks have increased dramatically. Therefore, traditional detection methods are inadequate for handling such large-scale and highly dynamic network environments. In the new network architecture, traffic patterns change frequently and are difficult to predict, and network status has a more significant impact on real-time performance. This necessitates network monitoring technologies that can provide more granular and higher-precision real-time data.
[0063] To achieve fine-grained monitoring of network data, existing technologies propose a network data monitoring method based on in-band network telemetry (INT). For example, Figure 1 illustrates a typical network data monitoring scenario. As shown in Figure 1, the network data monitoring scenario typically includes a telemetry analysis plane and a data plane. The telemetry analysis plane includes a control node, while the data plane typically includes a client, a first-hop network node, at least one intermediate network node, a last-hop network node, and a storage medium. Illustratively, the aforementioned network nodes can be one of a smart gateway, a programmable network interface card (NIC), a switch, or a server. The aforementioned storage medium can be a disk array or a cloud database.
[0064] When the control node needs to perform end-to-end network data monitoring.
[0065] First, the first-hop network node can unpack the data packet sent by the client to obtain the corresponding packet header and packet data. Then, the first-hop network node can use in-band telemetry to insert a telemetry header and the telemetry data sampled by the network node into the packet header. Finally, the first-hop network node can reassemble the new packet header and packet data to obtain the data packet and send it to the intermediate network node.
[0066] The telemetry head includes the configuration information for the aforementioned network data monitoring, as well as basic information such as the device identifier, ingress port number, and egress port number of the network nodes. Telemetry data typically includes real-time network status information.
[0067] Intermediate network nodes can use the same method to unpack data packets, insert the telemetry data 2 sampled by the network node, perform packet reassembly, and send the data packets to the last-hop network node.
[0068] The last-hop network node can also use the same method to unpack the data packet, extract all telemetry data, and encapsulate the monitored telemetry data 1, 2, and 3 with UDP and IP headers according to the packet encapsulation parameters configured by the telemetry header user. The encapsulated telemetry information is then forwarded to the control node. Simultaneously, this network node can also send the data packet to a disk array for storage.
[0069] However, the transmission of these messages between network nodes causes the telemetry tag header to continuously increase in size during processing, significantly enlarging compared to smaller packet traffic. Intermediate network interface card (NIC) nodes also need to perform packet unpacking and repackaging operations during data packet processing, thus affecting their forwarding performance. Furthermore, since the final receiver of the data packets is a disk array without network detection capabilities, it is difficult to detect the telemetry information corresponding to the disk array's back-end data and disk return data.
[0070] To better achieve network data sampling, statistics, and information monitoring, and to improve the reliability and real-time performance of data plane forwarding, this application proposes a network data monitoring method and related apparatus.
[0071] This method involves coloring a first data stream using the data processing unit of a first network node, inserting the colored bits into the first data stream to obtain a second data stream. During the processing of the second data stream by the first network interface card (NIC) node, at least one hardware unit of the first NIC node is sampled to obtain telemetry data at each sampling location. Through this method, the control node can proactively sense subtle changes in the network, accurately reflecting packet loss and latency in at least one hardware unit of the first network node's network. This achieves fine-grained network data monitoring, improving the reliability and real-time performance of data plane monitoring and data forwarding.
[0072] For example, Figure 2 is a schematic diagram of a network data detection scenario proposed in this application. As shown in Figure 2, the network data monitoring scenario typically includes a telemetry analysis plane and a data plane. The telemetry analysis plane includes a control node 100, and the data plane typically includes a client 200, a first-hop network node, at least one intermediate network node, a last-hop network node, and a storage medium.
[0073] As mentioned above, the network nodes can be any of the following: smart gateways, programmable network interface cards (NICs), switches, and servers. Illustratively, the first-hop network node includes a smart gateway 300, at least one intermediate network node includes a switch 400, the last-hop network node includes a server 500, and the storage medium includes a disk array 600.
[0074] Indicatively, the data flow transmission path on the data plane can start from client 200, pass through smart gateway 300, switch 400, server 500, and finally enter the disk array 600.
[0075] In some possible implementations, the data flow in the data plane also includes disk return data from the disk array. The transmission path of the disk return data flow can be starting from the disk array 600, passing through the server 500, the switch 400, the smart gateway 300, and finally entering the client 200.
[0076] On the data plane, client 200 can send data processing requests to smart gateway 300. The data processing request can be one of a read operation request, a write operation request, or a delete operation request.
[0077] Taking a write operation request as an example, the write operation request includes the data that needs to be written to disk array 600.
[0078] Client 200 can send the data to be written to smart gateway 300 in the form of a data stream. The data stream includes multiple data packets.
[0079] Client 200 can add a telemetry header to the header of the aforementioned data stream. The telemetry header includes routing and addressing information for the data stream, as well as network data monitoring configuration information.
[0080] In some possible implementations, to enable periodic reporting of telemetry information, the configuration information for network data monitoring may include a period identifier. This period identifier indicates the period at which telemetry information is reported to the control node. This period can be divided into odd-numbered periods and even-numbered periods, with corresponding period identifiers for each.
[0081] The smart gateway 300 includes a central processing unit (CPU) 310 and a data processing unit 320.
[0082] The central processing unit 310 is one of the core components of the smart gateway 300. In the embodiments of this application, the CPU 310 is used to process the input and output of the aforementioned data streams, run the operating system of the smart gateway 300, perform protocol conversion, manage device connections, and implement various intelligent functions. In actual implementation, the CPU 310 can be one or more of the following cores: application processor (AP), modem, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). Furthermore, the processor 310 can adopt an x86 architecture, an ARM architecture, or a hybrid architecture of the two. In other words, this application does not limit the type and implementation method of the processor.
[0083] The data processing unit 320 is used to implement the dot coloring and data sampling of the above data stream.
[0084] In some possible implementations, the data processing unit 320 may employ Internet Flow Inspection Technology (iFIT) to perform the marking and coloring of the aforementioned data stream, data sampling, and periodic reporting of the inspection data.
[0085] The following sections will introduce iFIT (Internet Frequency Inspection) technology and dot coloring.
[0086] iFIT (Internet Frequency Interpretation) is an in-band inspection technique. Within the iFIT statistical period, network nodes insert feature markers (such as color bits) into data packets to obtain telemetry information. This telemetry information is then reported in real-time to monitoring nodes using telemetry technology, allowing the monitoring nodes to measure network performance metrics based on the telemetry data. These performance metrics may include packet loss rate, latency, and jitter.
[0087] The iFIT statistical period refers to the time interval during which the device counts packets and records timestamps at the measurement point. Within this period, the device collects information such as packet loss and latency of the service flow and reports it to the control node 100 for analysis via telemetry technology.
[0088] Dot coloring is the process of marking features in a message. For example, IFIT marks feature fields by setting the packet loss coloring bit L and the delay coloring bit D to 0 or 1.
[0089] Please refer to Figure 2. First, after the data processing unit 320 receives the data stream (first data stream) sent by the client 300, it can create a data statistics table based on the period identifier in the telemetry head.
[0090] As one possible implementation, the data processing unit 320 can create two data statistics tables corresponding to the odd and even periods of the data stream, respectively, based on the period identifiers of the odd and even periods.
[0091] Then, the data processing unit 320 can determine the input and output identifiers of the data stream based on the period identifier recorded on the telemetry head, determine the coloring bit of the data stream based on the input and output identifiers, and color the coloring bit to obtain the colored data stream (second data stream).
[0092] Then, during the processing of the second data stream, the data processing unit 320 can sample the coloring bits of the second data stream at the sampling position based on the period indicated by the period identifier, obtain the coloring information at that sampling position, and input the coloring information into the aforementioned data statistics table. The coloring information includes the nanosecond-level hardware timestamp and the number of packets received at the sampling position of the data stream.
[0093] By way of example and not limitation, the data processing unit 320 may use a standardized timestamp format (such as IEEE 1588v2) or a format that supports multiple protocol layers (MAC / IP) timestamps, coloring multiple data packets in the data stream. In other words, this application does not limit this.
[0094] In actual implementation, the data processing unit 320 can use atomic counting to count the number of packets in the data stream based on the aforementioned color bits, and obtain the number of packets received in the data stream at the sampling position.
[0095] Next, we will introduce atomic counting through the following content.
[0096] Atomic counting refers to the reading and counting of data packets as an indivisible whole. Atomic counting can be implemented through hardware circuit design to ensure that operations on registers or memory are completed within a single clock cycle, avoiding interruption or interference from other threads or processes.
[0097] If the data processing unit 320 creates two data statistics tables corresponding to odd and even periods.
[0098] In some possible implementations, the data processing unit 320 inputs the coloring information records for odd-numbered periods into the data statistics table corresponding to the odd-numbered periods, and inputs the coloring information records for even-numbered periods into the data statistics table corresponding to the even-numbered periods.
[0099] The sampling locations mentioned above include the data inflow port and data outflow port of the smart gateway 300, as well as the hardware units of the smart gateway 300 (such as the processor 310). These sampling locations can be determined based on the routing, addressing information, and input / output identifiers in the data flow telemetry head.
[0100] The following content will introduce hardware timestamps.
[0101] Hardware timestamps are time stamps directly generated by the data processing unit 320 during data stream transmission and reception, typically used to achieve high-precision time recording at the nanosecond to microsecond level. Their core function is to provide accuracy in the time dimension for network data transmission.
[0102] As one possible implementation, the data processing unit 320 can send the above-mentioned coloring information to the processor 310 so that the processor can determine the delay information of the data stream based on the above-mentioned coloring information.
[0103] In actual implementation, the processor 310 can synchronize with the data processing unit 320 according to the period identifier in order to obtain the hardware timestamp of the above data stream more accurately.
[0104] In some possible implementations, the data processing unit 320 may also obtain the telemetry data 1 of the smart gateway 300 based on the above data statistics table, and send the telemetry data 1 to the control node 100 according to the period identifier, so that the control node 100 can perform behavior analysis and delimit and locate network data transmission faults (such as packet loss and delay) based on the above telemetry data 1.
[0105] As one possible implementation, while the DPU520 performs data statistics, the control node 100 can also actively obtain the aforementioned telemetry data 1 from the data processing unit 320 according to the statistical period. For other network nodes, the control node 100 can also use the same method to obtain the telemetry data sampled by each node, which will not be elaborated here.
[0106] It is worth noting that after the above telemetry data 1 is reported, the smart gateway 300 can also delete the timed-out data statistics table according to the predetermined table aging time, so as to optimize the memory resources of the smart gateway 300.
[0107] Understandably, this occurs when data flows from the smart gateway 300 to the switch 400.
[0108] The switch 400 can also use the same method, utilizing its own processor 410 and data processing unit 420, to realize the above-mentioned data stream dotting and coloring, data sampling, and periodic reporting of telemetry data 2 during the data stream processing process, which will not be elaborated here.
[0109] When data flows from switch 400 to server 500.
[0110] The same approach can be used for server 500 and switch 400 to achieve the above-mentioned data stream marking, data sampling, and periodic reporting of detection data during data stream processing.
[0111] It is worth noting that after the server 500 inputs the data stream to the disk array 600, it can also receive disk return data sent by the disk array 600.
[0112] Next, the following content will introduce the scenario of network data detection inside the server.
[0113] For example, Figure 3 is a schematic diagram of a network data detection scenario inside a server. As shown in Figure 3, server 500 can receive data streams sent by other network nodes. These network nodes may include client 200, smart gateway 300, switch 400, and any of other servers.
[0114] It is worth noting that the server 500 shown in Figure 3 is intended to facilitate understanding of this application. In actual implementation, the network node corresponding to the server 500 can also be any one of the client 200, smart gateway 300, switch 400, and other servers.
[0115] Compared to Figure 2, server 500 includes at least one hardware unit, for example, CPU 510 also includes a latency calculation unit 511, and DPU 520 also includes a parser 521, a disk data cache 522, a CRC check unit 523, and a disk return data cache 524.
[0116] After receiving the aforementioned data stream using DPU 520, server 500 can parse the data stream in parser 521 and create a corresponding data statistics table based on the telemetry header of the data stream. As mentioned earlier, DPU 520 can create data statistics tables corresponding to odd-numbered periods and even-numbered periods respectively.
[0117] Taking an odd-period data statistics table as an example, the data statistics table includes statistical data fields corresponding to period 1, period 3 and period 5.
[0118] After the data statistics table is created, the DPU520 can first perform the first sampling during the data flow from the DPU520 to the CPU510 to obtain the first sampled data, and input the first sampled data into the data field 1 corresponding to period 1 in the odd-period data statistics table.
[0119] Then, after receiving the aforementioned data stream, the processor 510 can determine the storage address of the data stream in the disk array 600 based on the telemetry head, and return the data stream to the lower disk data cache 522 of the DPU 520. At this time, the DPU 520 will perform a second sampling during the processing of the second data stream to obtain the second sampled data, and input the above second sampled data into the data field 2 corresponding to period 1 in the odd-period data statistics table.
[0120] Then, DPU520 can obtain the data stream from the lower disk data cache 522 and send the data stream to the storage address of disk array 600. At this time, DPU520 will perform a third sampling during the processing of the second data stream to obtain the third sampled data, and input the above third sampled data into the data field 3 corresponding to period 1 in the odd-period data statistics table.
[0121] Then, after the disk array 600 completes the storage of the second data stream, it can send the disk return data to the server 500.
[0122] After receiving the disk return data, DPU520 can send the data stream (third data stream) corresponding to the disk return data to the CRC check unit 523 of DPU520 for verification. At this time, DPU520 will perform a fourth sampling during the processing of the third data stream to obtain the fourth sample data, and input the above fourth sample data into the data field 4 corresponding to period 1 in the odd-period data statistics table.
[0123] Then, after the CRC check unit 523 completes the data stream verification, it sends the data stream to the disk return data buffer 524 so that the DPU 520 can send the data stream to other network nodes. At this time, the DPU 520 will perform a fifth sampling during the processing of the third data stream to obtain the fifth sample data, and input the fifth sample data into the data field 5 corresponding to period 1 in the odd-period data statistics table.
[0124] Finally, the DPU520 can obtain the telemetry data 3 of the server 500 based on the above data statistics table, and send the telemetry data 3 to the control node 100 according to the period indicated by the period identifier, so that the control node 100 can perform behavioral analysis and delimit and locate network data transmission faults (such as packet loss and latency) based on the above telemetry data 3.
[0125] It is worth noting that after the telemetry data 3 is uploaded to the control node 100, the control node 100 may experience a network data transmission failure.
[0126] In some possible implementations, the control node 100 can locate and delimit the faults of the processor 510 based on the first sampled data and the second sampled data, and determine the latency and packet loss at the processor 510.
[0127] As one possible implementation, the control node 100 can locate and delimit the faults of the disk array 600 based on the third and fourth sampled data, and determine the latency and packet loss at the disk array 600.
[0128] In some possible implementations, the control node 100 can locate and delimit the faults of the server 500 and other network nodes based on the fifth sampled data, and determine the latency and packet loss during the data flow from the server 500 to other network nodes.
[0129] Furthermore, when other network nodes receive the disk return data sent by server 500, they can also use the same method to detect the data stream corresponding to the disk return data, which will not be elaborated here.
[0130] It is worth noting that at least one hardware unit shown in server 500 in Figure 3 is intended to facilitate understanding of this application. In actual implementation, it may include encryption / decryption unit, decompression unit, etc., for processing the second data stream. It is understood that DPU 520 may also sample the above-mentioned at least one hardware unit during the processing of the second data stream to obtain telemetry information corresponding to the above-mentioned at least one hardware unit. In other words, this application does not limit the implementation method or specific type of the at least one hardware unit.
[0131] Firstly, based on the content described above, the network data monitoring method provided in the embodiments of this application will be introduced. It is understood that this method is proposed based on the content described above, and some or all of the content of this method can be found in the description above.
[0132] For example, Figure 4 is a flowchart of a network data monitoring method proposed in this application. As shown in Figure 4, network data monitoring can be achieved through steps S410 to S440. It is understood that this method can be executed by the DPU in any of the computing devices shown in Figure 2: the smart gateway 300, the switch 400, and the server 500.
[0133] S410: Receive the first data stream.
[0134] Taking server 500 (the first network node) as an example, as mentioned above, server 500 includes CPU 510 and DPU 520.
[0135] The DPU520 can receive data streams (first data streams) from other network nodes, such as clients, switches, servers, and smart gateways. As mentioned earlier, the first data stream includes a telemetry header. This telemetry header includes a period identifier, which indicates the period at which telemetry information is reported to the control node. The first data stream is intended to be processed at the first node.
[0136] Upon receiving the first data stream, the DPU520 can create a data statistics table based on the telemetry header. As mentioned earlier, the telemetry header is recorded in the data stream's header, including the data stream's routing and addressing information, as well as network data monitoring configuration information. This configuration information includes a period identifier. The period identifier indicates the period at which telemetry information is reported to the control node. The routing and addressing information indicates the input / output ports of the first network node.
[0137] In practice, this period can be divided into odd-numbered periods and even-numbered periods. Accordingly, the aforementioned period identifiers include identifiers for both odd-numbered and even-numbered periods.
[0138] In some possible implementations, the data processing unit 320 can create a data statistics table corresponding to an odd period based on the period identifier of the odd period, and create a data statistics table corresponding to an even period based on the period identifier of the even period.
[0139] S420: Based on the telemetry head, insert the colored bits into the first data stream to obtain the second data stream after dotting and coloring.
[0140] After receiving a data stream, the DPU520 can determine the input and output identifiers of the data stream based on the period identifier recorded in the telemetry head, determine the coloring bits of the data stream based on the input and output identifiers, and color the coloring bits to obtain the colored data stream (second data stream).
[0141] S430: During the process of at least one hardware unit in the first network node processing the second data stream, at least one hardware unit is sampled to obtain telemetry information corresponding to at least one hardware unit.
[0142] During the processing of the second data stream by at least one hardware unit in the first network node, the DPU520 can sample the coloring bits of the second data stream based on the period identifier at the sampling position, obtain the coloring information at the sampling position, and record the coloring information in the data statistics table established in step S410.
[0143] In some possible implementations, the above sampling location can be determined based on routing, addressing information and input / output identifiers, including the output port of the first network node and at least one hardware unit inside the first network node related to the second data stream processing.
[0144] The aforementioned hardware units may include one or more of the following: a parser, a processor, a data buffer encryption / decryption unit, a decompression unit, and a CRC check unit, through which the second data stream passes. The aforementioned coloring information includes the hardware timestamp and the number of packets received at the sampling location of the data stream.
[0145] In some possible implementations, when the sampling period is an odd number of periods, the DPU520 can record the above staining information in the data statistics table corresponding to the odd number of periods.
[0146] As one possible implementation, when the sampling period is an even number of periods, the DPU520 can record the above staining information in the data statistics table corresponding to the even number of periods.
[0147] Illustratively, within server 500, the processing path for the second data stream may include parser 521 - processor 510 - lower disk data cache 522.
[0148] When the sampling location is the parser of the first network node.
[0149] After the parser 521 completes the dot coloring, the DPU 520 can send the second data stream from the parser 521 to the processor 510, and sample the colored bits in the second data stream during the processing of the second data stream.
[0150] When the sampling location is the processor 510 of the first network node.
[0151] After determining the target network node (second network node) of the second data stream, the processor 510 can input the second data stream into the lower disk cache 522. The DPU 520 can sample the colored bits in the second data stream during the processing of the second data stream.
[0152] In some possible implementations, when the target network node is a storage medium, such as the disk array 600 shown in Figure 2, the data stream in the data plane (the third data stream) also includes disk return data from the disk array. The sampling location also includes a CRC check unit 523 for verifying the third data stream and a disk return data cache 524 for storing the third data stream.
[0153] In this case, the processing path of the third data stream may include CRC check unit 523 - disk return data cache 524.
[0154] When the sampling location is the CRC check unit of the first network node.
[0155] After receiving the third data stream, the DPU520 can send the third data stream to its CRC check unit 523 for verification. At this time, the DPU520 can sample the colored bits in the third data stream during the processing of the third data stream.
[0156] When the sampling location is the disk of the first network node, the data cache is 524.
[0157] After the CRC check unit 523 completes the verification of the third data stream, the DPU520 can send the third data stream from the CRC check unit 523 to the disk return data buffer 524, and sample the colored bits in the third data stream during the processing of the second data stream.
[0158] S540: Sends telemetry data to the control node according to the period indicated by the period identifier.
[0159] After completing the sampling and statistics of the data statistics table, the DPU520 can obtain the telemetry data of the first network node based on the above data statistics table, and send the telemetry data to the control node 100 according to the period indicated by the period identifier, so that the control node 100 can delimit and locate the network data transmission fault based on the above telemetry data.
[0160] Once the telemetry data is uploaded to control node 100, control node 100 can, in the event of a network data transmission failure, locate and delimit the fault at each sampling location based on the coloring information corresponding to each sampling location carried in the telemetry data, and determine the latency and packet loss during the processing of the second and / or third data streams between multiple network nodes and within each network node.
[0161] In some possible implementations, the data processing unit 520 may also send the above-mentioned coloring information to the processor 510 so that the processor 510 can determine the latency and packet loss during the processing of the second data stream and / or the third data stream between multiple network nodes and within each network node based on the above-mentioned coloring information.
[0162] In one possible implementation, the control node 100 may also actively obtain the aforementioned telemetry data from the data processing unit 320 according to the period indicated by the period identifier while the DPU 520 is performing data statistics.
[0163] As one possible implementation, in order to save on-chip memory resources of DPU520, DPU520 will simultaneously save data statistics tables for two cycles (one odd cycle and one even cycle). Control node 100 needs to read telemetry data after the statistics of the odd cycle stops and before the even cycle begins (i.e. before the data of the previous cycle is overwritten).
[0164] In practice, the DPU520 can report telemetry data to the control node 100 in two ways, depending on the service type of packet loss detection and latency statistics.
[0165] For example, Figure 5 is a schematic diagram of the transmission of the second data stream between the first network node and the second network node. As shown in Figure 5, the first network node sends a second data stream to the second network node through the network. The second data stream includes data packets in two sampling periods, T1 and T2. Period T1 includes data packets identified as 1 to 7 (represented by dark data blocks), corresponding to 7 valid packets in odd-numbered periods. Period T2 includes data packets identified as 1 to 7 (represented by light-colored data blocks), corresponding to 7 valid packets in even-numbered periods. The data packets include coloring information related to a specific sampling position.
[0166] During the processing of the second data stream, the aforementioned data packets may become out of order due to network latency. Illustratively, in the second data stream received by the second network node, the data packets in period T1 are out of order. The order in which the receiver receives the packets in the first period is: data packets identified as 1-6 in period T1, data packet identified as 1 in period T1, data packet identified as 7 in period T2, and data packets identified as 2-7 in period T1.
[0167] To avoid inaccurate statistics due to network latency and out-of-order packets, when the control node 100 processes packet loss detection, a period of time needs to be reserved after the T1 period to ensure that all data packets at time T1 are received completely.
[0168] Control node 100 can perform telemetry data statistics for cycle T1 at the end of cycle T1 and after the aforementioned reserved time period. It is important to note that reading the telemetry data for cycle T1 must be completed before the end of time T2.
[0169] Preferably, the reserved time period can be set to 2 / 3T after the start of period T2. Therefore, the control node 100 can start reading the telemetry data statistics of period T1 from (1+2 / 3)T after the start of period T1. When processing packet loss statistics, the control node 100 can calculate based on the period identifier in the telemetry header, without strictly relying on the control node 100's local time synchronization.
[0170] Furthermore, when the control node 100 processes the latency detection service, it does not need to consider network latency and packet out-of-order delivery. The control node 100 can read the telemetry data of the T1 cycle at any time after the start of T1 and before the end of the T2 cycle.
[0171] Figure 6 is an illustrative diagram illustrating the pipeline for reading telemetry data in the second data stream when the control node processes packet loss detection services. As shown in Figure 6, the reading time for telemetry data related to packet loss detection is 1 / 3T. Here, T represents the iFIT statistical period. When the time exceeds (1+2 / 3)T, the control node 100 and / or DPU520 can determine that the telemetry data statistics of the previous period have been completed and can begin reading the telemetry data of the current period.
[0172] Taking period T1 as an example, the statistical data of period T1 is complete in the last 1 / 3 of time T of period T2, and the statistical data of period T1 can be read in the last 1 / 3 of time T of period T2.
[0173] To ensure that the control node can read complete telemetry data starting from (1+2 / 3)T and before the end of the next cycle, for example, the period for reading packet loss detection data TLoss can be set to ≤1 / 3T to ensure that at least one telemetry data read is performed within the time when the telemetry data is complete.
[0174] As mentioned earlier, after completing the reading of telemetry data, the DPU520 can delete timed-out data statistics tables according to a predetermined entry aging time to optimize the memory resources of the server 500. For example, after completing the reading of telemetry data at time T1, the telemetry data at time T1 in the DPU520's on-chip memory can be deleted immediately.
[0175] Therefore, after time T2 ends, the DPU520 can clean up the statistics table of cycle T1 stored in on-chip memory to release memory resources so as to store the coloring information of the next odd cycle.
[0176] It is understandable that control node 100 can obtain telemetry data for cycles T2 and T3 from the on-chip memory of DPU520 in the same way, which will not be elaborated here.
[0177] It is worth noting that when a data packet of period T1 is received in the second data stream, the control node 100 can record the number of packets and bytes of period T1 based on the timestamp recorded in the packet.
[0178] When the telemetry data reading interval is short, the control node 100 can determine whether duplicate data has been read based on the period identifier in the second data stream.
[0179] If duplicate data is read, the control node 100 can filter the duplicate data based on the aforementioned periodic identifier.
[0180] In some possible implementations, the DPU520 can also send the telemetry data of the above T1, T2, and T3 cycles to the control node 100 in the same way, which will not be elaborated here.
[0181] Figure 7 is an illustrative diagram of the pipeline for reading telemetry data in the second data stream when the control node processes latency statistics services. Taking the T1 cycle shown in Figure 7 as an example, the statistical data for the T1 cycle is complete throughout the entire T2 cycle. When the control node 100 processes latency detection services, it does not need to consider network latency and packet out-of-order delivery.
[0182] As an illustration, in order to ensure that the control node 100 can read complete telemetry data related to the delay detection service, the period for reading the packet loss detection data can be TLoss≤T2. The control node 100 can read the telemetry data of the T1 period at any time after the start of T1 and before the end of the T2 period.
[0183] As mentioned earlier, after completing the reading of telemetry data, the DPU520 can delete timed-out data statistics tables according to a predetermined entry aging time to optimize the memory resources of the server 500. For example, after completing the reading of telemetry data at time T1, the telemetry data at time T1 in the DPU520's on-chip memory can be deleted immediately.
[0184] It is understandable that control node 100 can obtain telemetry data for cycles T2 and T3 from the on-chip memory of DPU520 in the same way, which will not be elaborated here.
[0185] In some possible implementations, the DPU520 can also send the telemetry data of the above T1, T2, and T3 cycles to the control node 100 in the same way, which will not be elaborated here.
[0186] As described in step S440, after the telemetry data is uploaded to the control node 100, the control node 100 can locate and delimit the fault at the sampling location based on the coloring information corresponding to each sampling location carried in the telemetry data, in the event of a network data transmission failure.
[0187] Figure 8 is an illustrative diagram illustrating a scenario of packet loss detection during the process of the control node writing data streams to the disk array. As shown in Figure 8, the second data stream passes through eight sampling positions a to h during the writing process to the disk array.
[0188] The second data stream's writing path to disk array 600 can pass through four sampling locations, a to d. Then, the second data stream can be written to disk array 600 via the Network Quality of Service (QoS) module. The QoS module is used for congestion control of data packets in the second data stream.
[0189] After the disk array 600 completes the writing of the second data stream, it can send the disk return data, i.e., the third data stream, to the sending port of the second data stream. Then, the third data stream can be input to the sampling position e through the network service quality module 700.
[0190] The third data stream can pass through four sampling positions, e to hi, in the processing path returned to the sending port.
[0191] As shown in Figure 8, the control node 100 can analyze the telemetry data of each sampling location to locate and delimit the faults at each sampling location.
[0192] To illustrate, the number of packets received at sampling position a at time T1 is 1000, the number of packets received at sampling position b at time T2 is 1000, the number of packets received at sampling position d at time T3 is 980, and the number of packets received at sampling position e at time T4 is 980.
[0193] Therefore, it can be concluded that during the data drop-off phase of the second data stream, packet loss occurred at sampling position d, with a packet loss count of 20.
[0194] It is worth noting that any of the eight sampling positions a to h mentioned above can be a network node, such as any one of a server, switch, or smart gateway, or any hardware unit of the aforementioned network node, such as any one of a parser, processor data cache, encryption / decryption unit, decompression unit, and CRC check unit. This application does not limit this.
[0195] Using the network data detection method described above, the control node can proactively sense subtle network changes and accurately reflect the packet loss and latency of at least one hardware unit in the first network node's network. Simultaneously, this method allows for cyclical overlay of data statistics tables corresponding to odd and even periods, and by deleting timed-out entries in the data statistics table, it reduces the on-chip resource overhead of the data processing unit. This method achieves fine-grained network data monitoring, improving the reliability and real-time performance of data plane monitoring and data forwarding.
[0196] Secondly, based on the content described above, the network data monitoring method provided in the embodiments of this application will be introduced. It is understood that this method is proposed based on the content described above, and some or all of the contents of this circuit can be found in the description above.
[0197] For example, Figure 9 is a flowchart of a network data monitoring method proposed in this application. As shown in Figure 9, network data monitoring can be achieved through steps S910 to S920. It can be understood that this method can be executed by the control node 100 shown in Figure 2.
[0198] S910: Acquire telemetry data related to the second data stream from the first network node according to a predetermined period. The second data stream is obtained by marking and coloring the first data stream sent by the client.
[0199] S920: Performs network data fault monitoring on at least one hardware unit of the first network node based on telemetry data.
[0200] Through this method, the control node can actively obtain telemetry data of at least one hardware unit of the first network node, and monitor the network data of the hardware unit based on the telemetry data, thereby realizing fine-grained network data monitoring and improving the network data detection efficiency.
[0201] Thirdly, based on the content described above, a network data monitoring device provided in the embodiments of this application will be introduced. It is understood that this device is proposed based on the content described above, and some or all of the circuitry can be found in the description above.
[0202] For example, Figure 10 is a schematic diagram of the structure of a network data monitoring device proposed in this application. As shown in Figure 10, the monitoring device can be applied to a data processing unit including: a coloring module 71, a sampling module 72, and a data reporting module 73.
[0203] The coloring module 71 is used to receive a first data stream, wherein the first data stream includes a telemetry head, the telemetry head includes a period identifier, the period identifier is used to indicate the period of reporting telemetry information to the control node, and the first data stream is intended to be processed in the first network node.
[0204] The coloring module 71 is also used to insert the coloring bits into the first data stream according to the telemetry head to obtain the second data stream after dotting and coloring.
[0205] The sampling module 72 is used to sample at least one hardware unit during the processing of the second data stream by at least one hardware unit in the first network node, so as to obtain telemetry information corresponding to the at least one hardware unit.
[0206] The data reporting module 73 is used to send the telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on at least one hardware unit based on the telemetry data.
[0207] In some possible implementations, the coloring module 71 is also used to create a data statistics table corresponding to the odd period based on the period identifier of the odd period, and to create a data statistics table corresponding to the even period based on the period identifier of the even period.
[0208] As one possible implementation, the sampling module 72 is also used to sample at least one hardware unit during the second data stream processing to obtain the coloring information of at least one hardware unit.
[0209] Then, based on the period identifier of the odd-numbered period, the coloring information of the odd-numbered period is recorded in the data statistics table corresponding to the odd-numbered period, and based on the period identifier of the even-numbered period, the coloring information of the even-numbered period is recorded in the data statistics table corresponding to the even-numbered period.
[0210] Finally, based on the data statistics tables corresponding to odd-numbered periods and even-numbered periods, telemetry information for at least one hardware unit is obtained.
[0211] In some possible implementations, the sampling module 72 is also used to delete the aforementioned data statistics table that has timed out according to a predetermined table entry aging time, so as to optimize the memory resources of the first network node.
[0212] As one possible implementation, the telemetry information includes a hardware timestamp and the number of packets received for at least one hardware unit. In this case, the data reporting module 73 is also used to send telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data latency statistics and packet loss detection for at least one hardware unit based on the hardware timestamp and the number of packets received.
[0213] In some possible implementations, the hardware unit includes an output port of the first network node for connecting to the second network node. In this case, the data reporting module 73 is also used to send telemetry data from the output port to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the second network node based on the telemetry data.
[0214] As one possible implementation, when the second network node is a storage medium, the sampling module 72 is also used to receive a third data stream returned by the storage medium. This third data stream is intended to be processed in the first network node.
[0215] Then, during the processing of the third data stream by at least one hardware unit in the first network node, at least one hardware unit is sampled to obtain telemetry information corresponding to at least one hardware unit.
[0216] Finally, telemetry data is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on at least one hardware unit based on the telemetry data.
[0217] Fourthly, based on the content described above, a computing device provided in the embodiments of this application will be introduced. It is understood that this circuit is proposed based on the content described above, and some or all of the contents of this circuit can be found in the description above.
[0218] For example, FIG11 is a schematic diagram of the structure of a computing device provided in this application. As shown in FIG11, the computing device 80 includes a data processing unit 81 and a memory 82. The data processing unit 81 is used to execute instructions stored in the memory 82 so that the data processing unit 81 performs the network data monitoring method provided by the first aspect or any possible implementation thereof on the processor 82.
[0219] In addition to the methods, apparatus, and electronic devices described above, embodiments of this application may also provide a computer program product, comprising computer program instructions. When executed by a processor, the computer program instructions cause the processor to perform the steps of the methods described in the "Methods" section of this specification. The computer program product can be written in any combination of one or more programming languages to execute the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The computer program code can be in source code form, object code form, executable file, or some intermediate form. The computer program code can be executed entirely on a user's computing device, partially on a user's device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0220] Furthermore, embodiments of this application may also provide a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the methods described in the "Method" section of this specification according to the various embodiments of this disclosure. The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. It should be noted that the content contained in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0221] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0222] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0223] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of the various embodiments of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the specific details described above.
[0224] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0225] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0226] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
[0227] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The specific embodiments described above have further detailed the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
[0228] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for monitoring network data, characterized in that, The network includes a first network node and a control node, the first network node includes a data processing unit, and the method is applied to the data processing unit, the method including: A first data stream is received, the first data stream includes a telemetry head, the telemetry head includes a period identifier, the period identifier is used to indicate the period of reporting telemetry information to the control node, and the first data stream is intended to be processed in the first network node. The colored bits are inserted into the first data stream according to the telemetry head to obtain the second data stream after dotting and coloring. During the process of at least one hardware unit in the first network node processing the second data stream, the at least one hardware unit is sampled to obtain telemetry information corresponding to the at least one hardware unit. The telemetry data is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the hardware unit based on the telemetry data.
2. The method according to claim 1, characterized in that, After receiving the first data stream, the method further includes: Create a data statistics table corresponding to the odd-numbered period based on the period identifier of the odd-numbered period; Create a data statistics table corresponding to the even-numbered period based on the period identifier of the even-numbered period.
3. The method according to claim 2, characterized in that, During the process of at least one hardware unit in the first network node processing the second data stream, sampling is performed on the at least one hardware unit to obtain telemetry information corresponding to the at least one hardware unit, including: The second data stream is processed using at least one hardware unit in the first network node; During the second data stream processing, the at least one hardware unit is sampled to obtain the coloring information of the at least one hardware unit; The coloring information of odd-numbered periods is recorded in the data statistics table corresponding to the odd-numbered periods according to the period identifier of the odd-numbered periods; and the coloring information of even-numbered periods is recorded in the data statistics table corresponding to the even-numbered periods according to the period identifier of the even-numbered periods. Based on the data statistics table corresponding to the odd-numbered periods and the data statistics table corresponding to the even-numbered periods, the telemetry information of the at least one hardware unit is obtained.
4. The method according to any one of claims 2-3, characterized in that, The method further includes: According to the predetermined aging time of the table entries, the timed-out data statistics table is deleted to optimize the memory resources of the first network node.
5. The method according to claim 1, characterized in that, The telemetry information includes a hardware timestamp and the number of packets received corresponding to at least one hardware unit. The step of sending the telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the hardware unit based on the telemetry data, includes: The telemetry data is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data latency statistics and packet loss detection on the at least one hardware unit based on the hardware timestamp and the number of packets received.
6. The method according to claim 1, characterized in that, The at least one hardware unit includes an output port of the first network node, the output port being used to connect to a second network node, and the step of sending the telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the hardware unit based on the telemetry data, further includes: The telemetry data from the output port is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the second network node based on the telemetry data.
7. The method according to claim 6, characterized in that, When the second network node is a storage medium, the method further includes: Receive the third data stream returned by the storage medium; the third data stream is intended to be processed in the first network node; During the process of at least one hardware unit in the first network node processing the third data stream, the at least one hardware unit is sampled to obtain telemetry information corresponding to the at least one hardware unit. The telemetry data is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the at least one hardware unit based on the telemetry data.
8. The method according to claim 1, characterized in that, The hardware unit includes any one of the following: a parser, a processor, a data cache, an encryption / decryption unit, a decompression unit, and a CRC check unit.
9. The method according to any one of claims 3 and 8, characterized in that, The method further includes: The coloring information is sent to the processor so that the processor can perform network data latency statistics on the at least one hardware unit and the second network node based on the coloring information.
10. A method for monitoring network data, characterized in that, The network includes a control node and a first network node, the method is applied to the control node, and the method includes: According to a predetermined cycle, telemetry data related to the second data stream is acquired from the first network node. The second data stream is obtained by marking and coloring the first data stream sent by the client. Based on the telemetry data, at least one hardware unit in the first network node is used to monitor network data faults.
11. A network data monitoring device, characterized in that, The device is applied to a data processing unit, and the device includes: A coloring module is used to receive a first data stream, the first data stream including a telemetry head, the telemetry head including a period identifier, the period identifier being used to indicate the period for reporting telemetry information to the control node, the first data stream being intended to be processed in a first network node; and inserting coloring bits into the first data stream according to the telemetry head to obtain a second data stream after dotting and coloring. The sampling module is used to sample the at least one hardware unit in the first network node during the processing of the second data stream, and to obtain the telemetry information corresponding to the at least one hardware unit. The data reporting module is used to send the telemetry data to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the hardware unit based on the telemetry data.
12. The apparatus according to claim 11, characterized in that, The staining module is also used for: Create a data statistics table corresponding to the odd-numbered period based on the period identifier of the odd-numbered period; Create a data statistics table corresponding to the even-numbered period based on the period identifier of the even-numbered period.
13. The apparatus according to claim 12, characterized in that, The sampling module is also used for: During the second data stream processing, the at least one hardware unit is sampled to obtain the coloring information of the at least one hardware unit; The coloring information of the odd-numbered period is recorded in the data statistics table corresponding to the odd-numbered period according to the period identifier of the odd-numbered period, and the coloring information of the even-numbered period is recorded in the data statistics table corresponding to the even-numbered period according to the period identifier of the even-numbered period. Based on the data statistics table corresponding to the odd-numbered periods and the data statistics table corresponding to the even-numbered periods, the telemetry information of the at least one hardware unit is obtained.
14. The apparatus according to any one of claims 12 and 13, characterized in that, The sampling module is also used for: According to the predetermined aging time of the table entries, the timed-out data statistics table is deleted to optimize the memory resources of the first network node.
15. The apparatus according to claim 11, characterized in that, The telemetry information includes a hardware timestamp and the number of packets received corresponding to at least one hardware unit. The data reporting module is also used for: The telemetry data is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data latency statistics and packet loss detection on the at least one hardware unit based on the hardware timestamp and the number of packets received.
16. The apparatus according to claim 11, characterized in that, The at least one hardware unit includes an output port of the first network node, the output port being used to connect to the second network node, and the data reporting module is further used for: The telemetry data from the output port is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the second network node based on the telemetry data.
17. The apparatus according to claim 11, characterized in that, When the second network node is a storage medium, the sampling module is further used for: Receive the third data stream returned by the storage medium; the third data stream is intended to be processed in the first network node; During the process of at least one hardware unit in the first network node processing the third data stream, the at least one hardware unit is sampled to obtain telemetry information corresponding to the at least one hardware unit. The telemetry data is sent to the control node according to the period indicated by the period identifier, so that the control node can perform network data fault monitoring on the at least one hardware unit based on the telemetry data.
18. A computing device, characterized in that, The computing device includes a data processing unit and a memory, the data processing unit being configured to execute instructions stored in the memory to cause the data processing unit to perform the method according to any one of claims 1-9.
19. A computer-readable storage medium, characterized in that, It includes computer program instructions that, when executed by the data processing unit, cause the data processing unit to perform the method as described in any one of claims 1-9.
20. A computer program product, characterized in that, It includes computer program instructions that, when executed by a data processing unit, cause the data processing unit to perform the method as described in any one of claims 1-9.