Data stream feature acquisition method and device and related equipment

By acquiring the traffic and attribute characteristics of multiple data streams before an emergency occurs and filtering them based on influencing factors, the problem of existing technologies being unable to identify the data streams driving emergencies is solved, enabling more accurate emergency analysis.

CN121644394APending Publication Date: 2026-03-10HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing data stream feature acquisition methods can only collect features of data streams that trigger emergencies, but cannot identify data streams that drive emergencies, resulting in inaccurate analysis.

Method used

Before a sudden event occurs, the traffic and attribute characteristics of multiple data streams are acquired, and the data stream characteristics that are directly or indirectly related to the sudden event are filtered according to the influencing factors of the sudden event.

Benefits of technology

By acquiring data stream characteristics from multiple dimensions, the impact of emergencies can be fully reproduced, improving the accuracy and efficiency of emergency cause analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644394A_ABST
    Figure CN121644394A_ABST
Patent Text Reader

Abstract

The invention provides a data flow feature collection method and device and related equipment, and the method comprises the steps: obtaining the flow features and attribute features related to a plurality of data flows before an emergency, and obtaining the flow features and attribute features related to a plurality of data flows after the emergency. The method comprises the following steps: analyzing an emergency to obtain influence factors and related data corresponding to the emergency, and finally filtering flow characteristics and attribute characteristics related to a plurality of obtained data streams according to the influence factors and the related data. The method comprises the following steps of: acquiring flow characteristics and attribute characteristics of data streams which possibly have a promoting effect on occurrence of emergencies; by implementing the method, the characteristics of the data stream directly related to the occurrence of the emergency and the characteristics of the data stream indirectly related to the occurrence of the emergency before and after the occurrence of the emergency are finally obtained, and the characteristics of the data streams can comprehensively reproduce the influence of the emergency, so that the cause analysis of the emergency is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and in particular to a method, apparatus, and related equipment for acquiring data stream features. Background Technology

[0002] During network transmission, various unforeseen events can occur, including but not limited to packet loss and backpressure events. Packet loss refers to the loss of data packets, which affects data integrity; backpressure events refer to slow data transmission rates, which affect data transmission time. These unforeseen events affect the speed and efficiency of network transmission, thereby impacting the user experience of related network products. For example, when users watch live video on their mobile phones, the occurrence of the above-mentioned unforeseen events may cause delays or stuttering in the live video or audio. Therefore, after the occurrence of such unforeseen events, it is crucial to quickly locate and confirm the cause of the unforeseen events, which is a strong demand from users.

[0003] The occurrence of the aforementioned emergencies is closely related to the data flow during network transmission. For example, a packet loss event is the phenomenon that a data packet is lost at a certain station or in the transmission channel between stations during network transmission; a backpressure event is when the transmission rate of a certain sending port of a certain station exceeds the processing rate of the corresponding receiving port of the next station during network transmission. Therefore, in the event of an emergency, by analyzing the characteristics of the data flow that is closely related to or promotes the occurrence of the emergency, the cause of the emergency can be quickly located and confirmed.

[0004] Current data stream feature acquisition methods collect features of the data stream that triggered the aforementioned sudden event after it has occurred. However, this method only collects the data stream that triggered the event. The data stream that actually caused the sudden event should be one that contributed to its occurrence before the event itself. For example, suppose data stream A experiences packet loss at station 1 during network transmission. Current data stream feature acquisition methods determine the corresponding data stream based on the lost packets, collect its features, and then confirm that data stream A caused the packet loss event. However, in reality, there might also be a data stream B. Data stream B and data stream A both head to the same sending port at site 1, but data stream B's buffer usage is larger than data stream A's, and data stream B arrives at site 1 before data stream A, occupying a large amount of site 1's buffer memory. When data stream A arrives at site 1, the buffer usage at site 1 is already close to saturation, ultimately leading to data packet loss for data stream A. In this packet loss event, although data stream A is the one that experiences packet loss, the reason for data stream A's packet loss is that data stream B occupies a large amount of site 1's buffer memory. Therefore, data stream B is the data stream that caused the packet loss event or played a role in promoting the occurrence of the packet loss event. Summary of the Invention

[0005] This application provides a data stream feature acquisition method, apparatus, and related equipment, which can obtain the features of data streams directly related to the occurrence of a sudden event before and after the event, as well as the features of data streams indirectly related to the occurrence of a sudden event, thereby supporting the analysis of the causes of sudden events.

[0006] In a first aspect, this application provides a data stream feature acquisition method, which includes: acquiring traffic features and attribute features associated with multiple first data streams; acquiring a first event; determining the influencing factors of the first event based on the first event; acquiring first data based on the first event; determining a first traffic feature and a first attribute feature based on the first data and the influencing factors, wherein the traffic feature includes the first traffic feature and the attribute feature includes the first attribute feature; and reporting the first traffic feature and the first attribute feature to an analyzer.

[0007] In the above scheme, the first event is the sudden event, and the influencing factors corresponding to the sudden event can be determined according to the type of the sudden event. For example, if the sudden event occurs at a port, then the influencing factor of the sudden event is the port; the first data is the feature associated with the sudden event. For example, when the sudden event occurs at port 1, the first data can be the features of the data flow at port 1 (including traffic features and attribute features); the features of multiple data flows are filtered according to the influencing factors corresponding to the sudden event and the first data to obtain the features of the data flows that satisfy the influencing factors of the sudden event and / or the first data. For the scheme that only collects the features of the data flows directly related to the occurrence of the sudden event, this application acquires the features of multiple data flows before the sudden event occurs. When the sudden event occurs, the features of these multiple data flows are filtered according to the influencing factors corresponding to the sudden event and the first data to obtain the features of the data flows indirectly related to the occurrence of the sudden event.

[0008] By implementing the above scheme, we can obtain the characteristics of the data streams directly related to the occurrence of the emergency and the data streams indirectly related to the occurrence of the emergency. These data stream characteristics can comprehensively reproduce the impact of the emergency, thereby supporting the analyzer's analysis of the causes of the emergency.

[0009] In one possible implementation of the first aspect, the traffic characteristics include cache-level characteristics, port-level characteristics, queue-level characteristics, and flow-level characteristics. The attribute characteristics include cache identifier id, port id, queue id, and data flow id composed of one or more of the header fields. The header fields include protocol, priority, source network interconnection protocol IP, destination IP, source port number, and destination port number. Among them, cache-level characteristics include cache occupancy / rate and packet loss / rate in the cache; port-level characteristics include port rate, port cache occupancy / rate, and port burst size; queue-level characteristics include enqueue rate, dequeue rate, queue length, queue latency, and packet loss / rate in the queue; and flow-level characteristics include data flow rate, data flow burst size, data flow cache occupancy / rate, and data flow packet loss / rate.

[0010] In the above scheme, the characteristics of the data stream are divided into flow characteristics and attribute characteristics. Flow characteristics are related to the flow rate and size of the data stream. Since the data stream can exist in multiple locations during network transmission, and the flow characteristics corresponding to these locations are different, this scheme further divides the flow characteristics into cache-level characteristics, port-level characteristics, queue-level characteristics, and stream-level characteristics. These represent the flow rate and size of the data stream in the cache, the port, the queue, and the transmission process, respectively. Attribute characteristics are related to the identifier of the data stream. Since the data stream can exist in multiple locations during network transmission, and the attribute characteristics corresponding to these locations are different, this scheme includes cache ID, port ID, queue ID, and data stream ID. These represent the identifier of the cache the data stream passes through, the identifier of the port the data stream passes through, the identifier of the queue the data stream passes through, and the identifier used by the data stream during network transmission, respectively.

[0011] By implementing the above scheme, the characteristics of the data stream are refined. When a sudden event occurs, appropriate characteristics are selected to filter the characteristics of multiple stored data streams according to actual needs. This reduces the proportion of irrelevant data stream characteristics in the final filtered data stream. Furthermore, since the characteristics of the final data stream also have multiple dimensions such as cache level, port level, queue level, and stream level, the analyzer can also analyze the characteristics of the data stream from multiple dimensions, thereby further supporting the analyzer's analysis of the causes of sudden events.

[0012] In one possible implementation of the first aspect, the influencing factors include one or more of the following: cache, port, queue, and data stream.

[0013] In one possible implementation of the first aspect, the first data includes one or more of the following: cache id, port id, queue id, and data stream id associated with the first event.

[0014] In the above scheme, since the traffic and attribute characteristics of the data stream have been divided into cache-level, port-level, queue-level, and stream-level characteristics, the influencing factors and related characteristics (first data) of the sudden event should also include these multiple dimensions of characteristics when a sudden event occurs. For example, if the sudden event occurs at a port, then the influencing factor of the sudden event is the port. If the port ID is 1, then the first data is port ID 1. The characteristics of the multiple data streams are filtered by the influencing factors and the first data, that is, the characteristics of the data stream that satisfies port ID 1 are extracted from the characteristics of the multiple data streams.

[0015] By implementing the above scheme, the influencing factors and associated primary data corresponding to the sudden event are refined. Appropriate influencing factors and primary data are selected in different application scenarios to filter the features of multiple data streams, thereby reducing the number of features of the final data stream and improving the execution efficiency of the method.

[0016] In one possible implementation of the first aspect, determining the first traffic feature and the first attribute feature based on the first data and influencing factors includes: obtaining statistical values ​​of the first features corresponding to multiple first data streams, where the first feature is a feature among the traffic features corresponding to the influencing factors, and the statistical value includes one or more of the maximum value, current value, and cumulative value; filtering the statistical values ​​of the first features corresponding to the multiple first data streams according to a first filtering condition to obtain multiple second features, where the first filtering condition is that the statistical value is greater than a first threshold, and the second feature is one or more of the attribute features corresponding to the first data streams that satisfy the first filtering condition among the statistical values ​​of the first features corresponding to the multiple first data streams; and filtering the traffic features and attribute features associated with the multiple first data streams according to the multiple second features and the first data to obtain the first traffic feature and the first attribute feature, where the first traffic feature and the first attribute feature are the parts of the traffic features and attribute features associated with the multiple first data streams that match the multiple second features and the first data.

[0017] In the above scheme, the occurrence of sudden events is often related to data streams with significant changes in traffic characteristics or large values ​​of traffic characteristics before the sudden event. For example, when the port buffer usage of data stream A suddenly increases, it indicates that data stream A has suddenly occupied a large amount of buffer in that port. If a sudden event similar to packet loss occurs at that port at this time, it can be suspected that the sudden event was caused by the sudden increase in the port buffer usage of data stream A. After obtaining the influencing factors corresponding to the sudden event, the statistical values ​​of the traffic characteristics of the data streams related to the influencing factors are calculated. Then, the statistical values ​​corresponding to the traffic characteristics of multiple data streams are compared with a preset threshold. The attribute characteristics of data streams whose statistical values ​​of traffic characteristics are greater than the preset threshold are obtained. Finally, based on the obtained attribute characteristics of the data streams and the characteristics associated with the occurrence of the sudden event (first data), the characteristics of multiple data streams are filtered to obtain the characteristics of potential data streams that may lead to the occurrence of the sudden event.

[0018] By implementing the above scheme, the traffic characteristics of multiple data streams are processed to filter out the attribute characteristics of data streams that may lead to sudden events. Then, based on the attribute characteristics of these data streams and the characteristics associated with the occurrence of sudden events, the characteristics of the acquired multiple data streams are filtered, thereby reducing the proportion of irrelevant data stream characteristics in the final acquired data stream characteristics and improving the accuracy of the analyzer's analysis of the causes of sudden events.

[0019] In one possible implementation of the first aspect, obtaining traffic features and attribute features associated with multiple first data streams includes: obtaining traffic features and attribute features associated with multiple second data streams; obtaining relevant information for predicting a first event based on the traffic features and attribute features associated with the multiple second data streams, wherein the relevant information for predicting the first event includes one or more of a cache ID, port ID, queue ID, and data stream ID associated with the predicted first event; filtering the traffic features and attribute features associated with the second data streams based on the relevant information for predicting the first event to obtain traffic features and attribute features associated with multiple first data streams; and the traffic features and attribute features associated with the multiple first data streams are traffic features and attribute features associated with the second data streams that match the relevant information for predicting the first event.

[0020] By implementing the above scheme, the location of emergencies can be predicted in advance, and the features of multiple data streams can be initially filtered based on the characteristics associated with the predicted emergencies. This allows for the acquisition of data stream features that match the characteristics associated with the predicted emergencies, thereby reducing the amount of data processing and the overhead of computing resources.

[0021] In one possible implementation of the first aspect, after obtaining the traffic features and attribute features associated with multiple first data streams, the method further includes: filtering the traffic features and attribute features associated with multiple first data streams according to a first abnormal condition, and storing the traffic features and attribute features associated with the first data streams that satisfy the first abnormal condition. The first abnormal condition includes the data stream's cache occupancy / rate being greater than a second threshold, and / or the data stream's burst size being greater than a third threshold.

[0022] In the above scheme, since the occurrence of sudden events is often related to data streams with large cache usage / rate and large burst size, when analyzing the characteristics of data streams, the characteristics of these multiple data streams can be filtered according to the cache usage / rate of the data stream or whether the burst size of the data stream is greater than a preset threshold, thereby reducing the acquisition of features of irrelevant data streams.

[0023] By implementing the above scheme, the features of multiple data streams are initially filtered by setting a preset threshold, and the features of the data streams that meet the preset threshold conditions are obtained, thereby reducing the amount of data processing and the overhead of computing resources.

[0024] In one possible implementation of the first aspect, obtaining traffic features and attribute features associated with multiple first data streams includes: filtering the traffic features and attribute features associated with multiple first data streams according to a second abnormal condition, and storing the traffic features and attribute features associated with the first data streams that satisfy the second abnormal condition. The second abnormal condition includes the data stream's cache occupancy / rate being greater than a fourth threshold, and / or the data stream's burst size being greater than a fifth threshold. The storage areas of the traffic features and attribute features associated with the first data streams that satisfy the first abnormal condition are different from those of the traffic features and attribute features associated with the first data streams that satisfy the second abnormal condition.

[0025] In the above scheme, after obtaining the characteristics of multiple data streams, these characteristics are filtered based on the cache usage / rate of different data streams or the burst size of the data streams. By setting different filtering conditions, characteristics of different types of data streams can be obtained. Selecting the characteristics of different types of data streams according to the actual application scenario can not only reduce the consumption of computing resources, but also improve the accuracy of the analysis results of the causes of sudden events. For example, two thresholds can be set regarding the burst size of the data stream, where threshold A is greater than threshold B. After multiple data stream features are filtered through thresholds A and B respectively, two types of data stream features are obtained. Since threshold A is greater than threshold B, the number of features of the data stream filtered by threshold A will be less than the number of features of the data stream filtered by threshold A. When the analyzer analyzes the features of the data stream, the larger the data volume, the higher the accuracy of the analysis results of the cause of the corresponding burst event, but the greater the computational resource overhead; the smaller the data volume, the lower the accuracy of the analysis results of the cause of the corresponding burst event, but the smaller the computational resource overhead. Therefore, when no burst event has occurred, it is possible to choose to analyze the features of the data stream filtered by threshold A to reduce the computational resource overhead; when a burst event has occurred, it is possible to choose to analyze the features of the data stream filtered by threshold B to improve the accuracy of the analysis results of the cause of the burst event.

[0026] By implementing the above scheme, the features of multiple data streams are filtered according to different filtering conditions to obtain features of different types of data streams. By selecting features of different types of data streams for analysis and processing, a suitable balance is found between the consumption of computing resources and the accuracy of prediction results, thereby meeting the usage needs of various application scenarios.

[0027] In one possible implementation of the first aspect, the first event includes: action-based events and threshold-based events, wherein the action-based events include one or more of the following: packet loss, congestion, explicit congestion notification (ECN) flag, and backpressure flag within the switch; and the threshold-based events include one or more of the following: buffer occupancy / rate within the switch is greater than a sixth threshold, port rate is greater than a seventh threshold, and packet loss / rate in the queue is greater than an eighth threshold.

[0028] In the above scheme, the criteria for judging an emergency (the first event mentioned above) include actual action indicators and pre-set threshold conditions, wherein the threshold conditions can be set according to the user's needs.

[0029] Implementing the above solution expands the definition of emergencies, allowing users to choose appropriate judgment methods based on their needs, thereby enabling the acquisition and analysis of data stream characteristics in different application scenarios. For example, during highly confidential file transfers, it is necessary to ensure that the packet loss rate of the data stream is extremely low. In this case, the packet loss rate of the data stream in the queue can be set to be greater than 0.5% to determine whether network transmission has an anomaly, thus enabling timely updates and maintenance.

[0030] Secondly, this application provides a data stream feature acquisition device, the device comprising: an acquisition unit for acquiring traffic features and attribute features associated with multiple first data streams; an event monitoring unit for acquiring a first event; the event monitoring unit is further configured to determine the influencing factors of the first event based on the first event; and to acquire first data based on the first event; a filtering unit for determining a first traffic feature and a first attribute feature based on the first data and the influencing factors, wherein the traffic feature includes the first traffic feature and the attribute feature includes the first attribute feature; and the filtering unit is further configured to report the first traffic feature and the first attribute feature to an analyzer.

[0031] In one possible implementation of the second aspect, the traffic characteristics include cache-level characteristics, port-level characteristics, queue-level characteristics, and flow-level characteristics. The attribute characteristics include cache identifier id, port id, queue id, and data flow id composed of one or more of the packet header fields. The packet header fields include protocol, priority, source network interconnection protocol IP, destination IP, source port number, and destination port number. Among them, cache-level characteristics include cache occupancy / rate and packet loss / rate in the cache; port-level characteristics include port rate, port cache occupancy / rate, and port burst size; queue-level characteristics include enqueue rate, dequeue rate, queue length, queue latency, and packet loss / rate in the queue; and flow-level characteristics include data flow rate, data flow burst size, data flow cache occupancy / rate, and data flow packet loss / rate.

[0032] In one possible implementation of the second aspect, the influencing factors include one or more of the following: cache, port, queue, and data stream.

[0033] In one possible implementation of the second aspect, the first data includes one or more of the following: cache ID, port ID, queue ID, and data stream ID associated with the first event.

[0034] In one possible implementation of the second aspect, the filtering unit is specifically used to: obtain statistical values ​​of first features corresponding to multiple first data streams, wherein the first features are features in the traffic features corresponding to influencing factors, and the statistical values ​​include one or more of the maximum value, current value, and cumulative value; filter the statistical values ​​of the first features corresponding to multiple first data streams according to a first filtering condition to obtain multiple second features, wherein the first filtering condition is that the statistical value is greater than a first threshold, and the second features are one or more of the attribute features corresponding to the first data streams that satisfy the first filtering condition among the statistical values ​​of the first features corresponding to the multiple first data streams; filter the traffic features and attribute features associated with multiple first data streams according to the multiple second features and the first data to obtain first traffic features and first attribute features, wherein the first traffic features and first attribute features are the parts of the traffic features and attribute features associated with multiple first data streams that match the multiple second features and the first data.

[0035] In one possible implementation of the second aspect, the acquisition unit is specifically used to: acquire traffic features and attribute features associated with multiple second data streams; acquire relevant information for predicting a first event based on the traffic features and attribute features associated with the multiple second data streams, wherein the relevant information for predicting the first event includes one or more of the cache ID, port ID, queue ID, and data stream ID associated with the first event; filter the traffic features and attribute features associated with the second data streams based on the relevant information for predicting the first event, and acquire traffic features and attribute features associated with multiple first data streams; the traffic features and attribute features associated with the multiple first data streams are traffic features and attribute features associated with the second data streams that match the relevant information for predicting the first event.

[0036] In one possible implementation of the second aspect, the acquisition unit is further configured to: filter the traffic features and attribute features associated with multiple first data streams according to a first abnormal condition, and store the traffic features and attribute features associated with the first data streams that satisfy the first abnormal condition, wherein the first abnormal condition includes the data stream's cache occupancy / rate being greater than a second threshold, and / or the data stream's burst size being greater than a third threshold.

[0037] In one possible implementation of the second aspect, the acquisition unit is further configured to: filter the traffic features and attribute features associated with multiple first data streams according to a second abnormal condition, and store the traffic features and attribute features associated with the first data streams that satisfy the second abnormal condition, wherein the second abnormal condition includes the data stream's cache occupancy / rate being greater than a fourth threshold, and / or the data stream's burst size being greater than a fifth threshold; the storage areas of the traffic features and attribute features associated with the first data streams that satisfy the first abnormal condition are different from those of the traffic features and attribute features associated with the first data streams that satisfy the second abnormal condition.

[0038] In one possible implementation of the second aspect, the first event includes: action-based events and threshold-based events, wherein the action-based events include one or more of the following: packet loss, congestion, explicit congestion notification (ECN) flag, and backpressure flag within the switch; and the threshold-based events include one or more of the following: buffer occupancy / rate within the switch is greater than a sixth threshold, port rate is greater than a seventh threshold, and packet loss / rate in the queue is greater than an eighth threshold.

[0039] Thirdly, this application provides a computing device including a processor and a memory, the memory for storing instructions and the processor for executing instructions, so that the computing device implements the data stream feature acquisition method as described in the first aspect and any possible implementation thereof.

[0040] Fourthly, this application provides a computing device cluster, which includes at least one computing device. Each computing device includes a processor and a memory. The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster implements the data flow feature acquisition method in the first aspect and any possible implementation of the first aspect.

[0041] Fifthly, this application provides a computer-readable storage medium storing instructions that are executed by a computing device or a cluster of computing devices to implement the data stream feature acquisition method as described in the first aspect and any possible implementation thereof.

[0042] In a sixth aspect, this application provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to perform the data stream feature acquisition method in the first aspect and any possible implementation of the first aspect.

[0043] The second, third, fourth, fifth, and sixth aspects mentioned above all have various possible designs similar to the first aspect and any possible implementation of the first aspect, and can produce corresponding technical effects, which will not be elaborated here. Attached Figure Description

[0044] Figure 1 This is an architectural diagram of a data stream feature acquisition device provided in this application;

[0045] Figure 2 This is a schematic diagram illustrating the steps of a data stream feature acquisition method provided in this application;

[0046] Figure 3 This is a schematic diagram illustrating the process by which a data stream feature acquisition device, as provided in this application, acquires features of multiple data streams.

[0047] Figure 4 This is a schematic diagram of a data stream upload process provided in this application;

[0048] Figure 5 This is a schematic diagram illustrating the occurrence of an emergency, as provided in this application.

[0049] Figure 6 This is a schematic diagram of the operational logic of a data stream feature acquisition method provided in this application;

[0050] Figure 7 This is a schematic diagram illustrating the analysis results of the characteristics of a data stream provided in this application;

[0051] Figure 8 This is a schematic diagram illustrating the analysis results of another data stream characteristic provided in this application;

[0052] Figure 9 This is a schematic diagram of the hardware structure of a data stream feature acquisition device provided in this application;

[0053] Figure 10 This is a schematic diagram of the hardware structure of another data stream feature acquisition device according to this application;

[0054] Figure 11 This is a schematic diagram of the structure of a computing device cluster provided in this application;

[0055] Figure 12 This is a schematic diagram of another structure of a computing device cluster provided in this application. Detailed Implementation

[0056] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0057] To ensure that the characteristics of the collected data streams accurately reflect the causes of emergencies, this application provides a data stream feature acquisition method, apparatus, and related equipment. The data stream feature acquisition method includes: acquiring traffic and attribute characteristics associated with multiple data streams before the emergency occurs; analyzing the emergency after it occurs to obtain the corresponding influencing factors and related data; and finally filtering the acquired traffic and attribute characteristics of the multiple data streams based on the influencing factors and related data to obtain the traffic and attribute characteristics of data streams that may have a driving effect on the occurrence of the emergency. Compared to directly acquiring the characteristics of the data streams that triggered the emergency, the above data stream feature acquisition method not only acquires the characteristics of data streams directly related to the occurrence of the emergency before and after it occurs, but also acquires the characteristics of data streams indirectly related to the occurrence of the emergency. By analyzing these data stream characteristics, the impact of the emergency can be comprehensively reproduced, thereby supporting the analysis of the causes of the emergency. For example, in the aforementioned background technology, data stream A and data stream B simultaneously head to the same sending port (hereinafter referred to as port 1) of station 1. Since data stream B has a larger buffer usage than data stream A, and data stream B arrives at station 1 first, occupying a large amount of station 1's buffer memory, data stream A experiences packet loss, leading to a packet loss event at station 1. Using the data stream feature acquisition method provided in this application, before the packet loss event occurs, the data stream feature acquisition device acquires and stores the traffic characteristics and attribute characteristics of multiple data streams, including the characteristics of data stream A and data stream B. The attribute characteristics of data stream B include that the destination port is port 1, and the traffic characteristics include the buffer usage of data stream B; the attribute characteristics of data stream A include that the destination port is port 1, and the traffic characteristics include the buffer usage of data stream A. After the packet loss event occurs, the data stream feature acquisition device acquires the influencing factors of the packet loss event (i.e., the packet loss event affects the port) and the first data associated with the occurrence of the packet loss event (data stream A). The system filters the traffic and attribute characteristics of multiple stored data streams using the attribute characteristics (port 1) of data stream A. This process ultimately obtains the traffic and attribute characteristics associated with data stream A and data stream B. Finally, the analyzer analyzes these traffic and attribute characteristics and data stream B, discovering that data stream B arrived at site 1 before data stream A and occupied a large amount of site 1's cache memory, leading to packet loss in data stream A. Therefore, it is confirmed that data stream B occupied a large amount of site 1's cache memory, causing the packet loss event at site 1.

[0058] See Figure 1 , Figure 1 This is an architectural diagram of a data stream feature acquisition device provided in this application, such as... Figure 1 As shown, the architecture includes a data flow feature acquisition device 100, a network transmission path, and an analyzer 200. The data flow feature acquisition device 100 is used to acquire the features of multiple data flows in the network transmission path and send the data flow features to the analyzer 200 in the event of a sudden event. The network transmission path contains multiple data flows used for network transmission. The analyzer 200 is used to locate and confirm the cause of the sudden event based on the features of the data flows.

[0059] The data stream feature acquisition device 100 and analyzer 200 can be deployed on computing devices, including virtual machines, containers, or servers. A virtual machine is a virtualization technology implemented at the computer software level, enabling a single physical computer to create multiple virtual operating systems and application environments, each running its own operating system and applications independently. A container is a lightweight software packaging method used to package an application and its application environment, allowing the application to run in the same way in different environments. Unlike virtual machines, containers do not contain a complete operating system but share the operating system of the physical computer they reside on, making them more lightweight. A server is a general-purpose physical server, including ARM servers or x86 servers.

[0060] The data stream feature acquisition device 100 and the analyzer 200 can also be deployed in a computing device cluster, which includes multiple computing devices as described above.

[0061] The data stream feature acquisition device 100 and analyzer 200 can also be deployed on terminal devices, including computer terminal devices, mobile terminal devices, network terminal devices, point-of-sale (POS) devices, industrial control terminal devices, and virtual terminal devices. Among them, computer terminal devices include personal computers, laptops, and tablets; mobile terminal devices include smartphones, smartwatches, and portable music players; network terminal devices include routers, switches, and modems; point-of-sale (POS) devices include cash registers, card readers, and self-service payment terminals; industrial control terminal devices include sensors, industrial robots, and industrial control panels; and virtual terminal devices include virtual reality (VR) devices.

[0062] This document explains that the data flow feature acquisition device 100 and analyzer 200 described above can be deployed on the same computing device or mobile terminal, or they can be deployed on different computing devices or mobile terminals. For example, the data flow feature acquisition device 100 and analyzer 200 can be deployed on different computing devices in a computing device cluster. Whether the data flow feature acquisition device 100 and analyzer 200 are deployed on the same computing device depends on the specific application environment, and this application does not make specific limitations here.

[0063] The data stream feature acquisition device 100 can be further divided into multiple unit modules, such as... Figure 1 As shown, the data stream feature acquisition device 100 also includes an acquisition unit 101, an event monitoring unit 102, and a filtering unit 103. It should be understood here that... Figure 1 The number and names of the unit modules included in the data stream feature acquisition device 100 are merely examples provided in this application. The data stream feature acquisition device 100 may include more or fewer unit modules, and the names of the unit modules are not limited to [specific examples]. Figure 1 The name of the unit module in the data stream feature acquisition device 100 may be specified. For example, the data stream feature acquisition device 100 may also include a feature calculation unit. The feature calculation unit is used to analyze multiple data streams in the network transmission path to obtain the features of these multiple data streams, and then send the features of the multiple data streams to the acquisition unit 101. For example... Figure 1 The acquisition unit 101 and the event monitoring unit 102 can be combined into a data stream feature storage and transmission unit; Figure 1 The name of the acquisition unit 101 can be changed to the storage unit. The above example is for illustration only and should not be regarded as a specific limitation.

[0064] The aforementioned acquisition unit 101, event monitoring unit 102, and filtering unit 103 can be implemented in software or hardware. The following describes the software and hardware implementation methods of the acquisition unit 101. The software and hardware implementation methods of the event monitoring unit 102 and the filtering unit 103 can refer to the software and hardware implementation methods of the acquisition unit 101.

[0065] When the acquisition unit 101 is implemented by software, the acquisition unit 101 can be code running on the aforementioned computing device or terminal device. That is, the acquisition unit 101 can be code running on a personal computer, smartphone, or server, and the number of computing devices or terminal devices can be one or more. That is, the acquisition unit 101 can also be code running on a cluster of computing devices.

[0066] When the acquisition unit 101 is implemented in hardware, it can be implemented by at least one computing device or terminal device. Alternatively, the acquisition unit 101 can also be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), wherein the PLD includes one or more of complex programmable logic devices (CPLD), field-programmable gate arrays (FPGA), and generic array logic (GAL).

[0067] The functions of the acquisition unit 101, event monitoring unit 102, and filtering unit 103 in the data stream feature acquisition device 100 described above will be explained below.

[0068] The acquisition unit 101 is used to acquire the traffic characteristics and attribute characteristics associated with multiple first data streams.

[0069] The aforementioned first data stream is a data stream, which refers to a continuously transmitted set of data during network transmission, typically consisting of a series of data packets arriving in chronological order. The fundamental characteristics of a data stream include time sequence and dynamism. Time sequence means that the data packets in the data stream are arranged in chronological order, with earlier sent or arrived packets preceding later ones. Dynamism means that the data stream is constantly changing; new data packets may be introduced, and older packets may be discarded or processed, thus the length of the data stream also changes dynamically.

[0070] The traffic characteristics and attribute characteristics associated with the aforementioned first data stream (hereinafter referred to as data stream characteristics) refer to some identifiers and properties of the data stream during network transmission. Among them, attribute characteristics include cache identifier (identitydocument, id), port id, queue id, and data stream id composed of one or more of the header fields. The header fields here include protocol, priority, and source network interconnection protocol (Internet Protocol). Protocol (IP), destination IP, source port number, destination port number, etc. For example, a data flow ID can be a tuple consisting of the source IP and destination IP, or a quintuple consisting of the source IP, destination IP, source port number, destination port number, and protocol. Traffic characteristics can also be divided into buffer-level characteristics, port-level characteristics, queue-level characteristics, and flow-level characteristics. Buffer-level characteristics refer to the characteristics of the data flow in the buffer, specifically including the number of bytes, the number of packets, latency, and packet loss rate, etc. Port-level characteristics refer to the characteristics of the data flow at the source and destination ports, specifically including the data flow rate at the port, the buffer usage at the port, and the burst size at the port, etc. Queue-level characteristics refer to the characteristics of the data flow in the queue, specifically including the enqueue rate, dequeue rate, the length of the data flow in the queue, and the latency generated by passing through the queue, etc. Flow-level characteristics refer to the characteristics inherent in the data flow during network transmission, specifically including the data flow rate, the data flow burst size, the data flow buffer usage rate, and the data flow packet loss rate, etc.

[0071] This explains that the burst size of the data stream refers to a significant increase or decrease in the data flow within a short period of time due to sudden user activity, equipment failure, or network attack. The larger the burst size, the greater the fluctuation in the data flow within a short period of time. The attribute characteristics of the data stream are the same as those of the data packets that constitute the data stream, that is, the protocol, priority, source IP, destination IP, cache ID, source port number, and destination port number contained in the data stream and the data packets that constitute the data stream are the same.

[0072] The acquisition unit 101 acquires traffic characteristics and attribute characteristics associated with multiple first data streams, specifically including: the acquisition unit 101 collects the characteristics and status of each data packet in the network transmission path at a certain time interval; by analyzing and calculating the characteristics and status of each data packet, the characteristics of the data stream corresponding to each data packet are acquired, and the characteristics of the data stream are stored.

[0073] The characteristics of the aforementioned data packets include packet length, packet interval, packet timestamp, packet 5-tuple information, and packet priority. The packet 5-tuple information includes source IP, destination IP, source port number, destination port number, and protocol. The packet status includes whether the packet was dropped, whether the packet was enqueued, and whether the packet was dequeued. It should be noted that the packet timestamp refers to the time information recorded during the network transmission of the data packet, typically including the sending and receiving times.

[0074] The following describes how the acquisition unit 101 acquires the traffic characteristics and attribute characteristics associated with the data stream based on the characteristics and status of the data packet.

[0075] Here, we take the cache-level features in the traffic features as an example to illustrate how the acquisition unit 101 acquires the traffic features associated with the data stream. The features of the data stream in the cache include the number of data packets, the number of bytes, the latency, and the packet loss rate, etc.

[0076] ① Number of data packets remaining in the buffer: Acquisition unit 101 collects the characteristics and status of each data packet in the network transmission path, specifically including the five-tuple information of the data packet, whether the data packet has been dropped, whether the data packet has been enqueued, and whether the data packet has been dequeued; based on the five-tuple information of the data packet, the number of data packets from the same data stream is determined, and then by judging whether the data packet has been dropped, whether the data packet has been enqueued, and whether the data packet has been dequeued; finally, the number of data packets remaining in the buffer is confirmed. For example, based on the five-tuple information of the data packet, it is determined that there are 16 data packets from the same data stream, of which 1 has been dropped, 2 have been enqueued, and 3 have been dequeued, then the number of data packets remaining in the buffer for this data stream is 12 (total number of data packets - number of dropped data packets - number of dequeued data packets).

[0077] ② Number of bytes remaining in the buffer: The acquisition unit 101 collects the characteristics and status of each data packet in the network transmission path, specifically including the five-tuple information of the data packet, the length of the data packet, the interval between data packets, whether the data packet is dropped, whether the data packet is enqueued, and whether the data packet is dequeued; based on the five-tuple information of the data packet, the number of data packets from the same data stream is determined, and then by judging whether the data packet is dropped, whether the data packet is enqueued, and whether the data packet is dequeued, the number of data packets remaining in the buffer is confirmed; finally, the number of bytes remaining in the buffer is determined by combining the length of the data packet and the interval between the data packets. For example, if the number of data packets remaining in the buffer is determined to be 12 based on the five-tuple information of the data packet, and the length of each data packet is 64 bytes and the interval between data packets is 2 bytes, the number of bytes remaining in the buffer is 790 bytes (number of data packets × length of each data packet + (number of data packets - 1) × interval between data packets).

[0078] ③ Data stream latency in the buffer: Acquisition unit 101 collects the characteristics and status of each data packet in the network transmission path, specifically including the packet's five-tuple information, the packet's timestamp, whether the packet was dropped, whether the packet was enqueued, and whether the packet was dequeued. Based on the packet's five-tuple information, it determines the data packets from the same data stream. Then, by analyzing the timestamps of the data packets from the same data stream, it obtains the timestamp of the first data packet to arrive in the buffer and the timestamp of the last data packet to leave the buffer, thereby obtaining the data stream latency in the buffer. For example, based on the packet's five-tuple information, it is determined that the number of data packets staying in the buffer is 12. By analyzing the timestamps of these 12 data packets, the timestamp of the first data packet to arrive in the buffer is 20:21, and the timestamp of the last data packet to leave the buffer is 20:30. Therefore, the latency of this data stream in the buffer is 9s (timestamp of the last data packet to leave the buffer - timestamp of the first data packet to arrive in the buffer).

[0079] ④ Packet loss of the data stream: The acquisition unit 101 acquires the characteristics and status of the data packets, specifically including the five-tuple information of the data packets and whether the data packets were dropped; based on the five-tuple information of the data packets, it determines the number of data packets from the same data stream, and then determines whether the data packets were dropped, thus determining the packet loss of the data stream. For example, if based on the five-tuple information of the data packets, it is determined that there are 16 data packets from the same data stream, of which 1 was dropped, then the packet loss of this data stream is 1 (i.e., the number of dropped data packets).

[0080] Here, we will take the protocol, destination IP, source IP, source port number, and destination port number in the attribute features as examples to illustrate in detail how the acquisition unit 101 acquires the attribute features associated with the data stream.

[0081] The acquisition unit 101 collects the characteristics and status of each data packet in the network transmission path, specifically including the five-tuple information of the data packet; and determines the number of data packets from the same data stream and the attribute characteristics of the data stream based on the five-tuple information of the data packet.

[0082] Optionally, while acquiring the characteristics of multiple data streams, the acquisition unit 101 can also record the statistical values ​​of the traffic characteristics associated with each data stream. These statistical values ​​can be one or more of the maximum value, current value, and cumulative value. By recording the statistical values ​​of the traffic characteristics associated with each data stream, it is possible to quickly determine which data streams exhibited the largest changes in their associated traffic characteristics within a period before and after the occurrence of a sudden event. In other words, the occurrence of a sudden event is often related to the data streams with the largest changes or the largest cumulative changes in characteristics within a period before and after the event.

[0083] The following explanation uses the maximum burst size of the data stream recorded by the acquisition unit 101 as an example. The acquisition unit 101 collects the characteristics and status of each data packet in the network transmission path using a time window approach, specifically including the packet's 5-tuple information, packet length, and packet interval. Within the time window, the number of data packets from the same data stream is determined based on the packet's 5-tuple information. The burst size of each data stream within that time window is determined based on the number of data packets, packet length, and packet interval. Then, the acquisition unit 101 records the burst size of each data stream within that time window as the maximum burst size, and repeats the above steps to obtain the burst size of each data stream in a new time window. The acquisition unit 101 compares the burst size of each data stream in the new time window with its corresponding maximum burst size. When the burst size in the new time window is greater than its corresponding maximum burst size, the maximum burst size of that data stream is updated to the burst size in the new time window. It should be understood that the acquisition unit 101 repeatedly performs the above steps to obtain the maximum burst size of each data stream over a period of time.

[0084] Optionally, the acquisition unit 101 may store only the features of the data stream within a certain time period or only the features of the data stream within a certain space size. It should be understood that the reason the acquisition unit 101 stores only the features of the data stream within a certain time period or a certain space size is due to the dynamic nature of the data stream. Since the features of the data stream are constantly changing, using the features of the data stream with a time interval too large than the sudden event to analyze the sudden event will introduce unnecessary features, thereby reducing the accuracy of subsequent analysis results. Therefore, the data stream feature acquisition device 100 can select the features of the data stream within a certain period before and after the sudden event for processing. For example, the acquisition unit 101 stores the features of each data stream for only 30 seconds. When the storage time of the features of a data stream at a certain moment exceeds 30 seconds, the acquisition unit 101 deletes the features of the data stream at that moment. The acquisition unit 101 stores only 30 features of each data stream. When the acquisition unit 101 has stored 30 features of a data stream and needs to store new features of a data stream, the acquisition unit 101 can delete the features of the data stream with the longest storage time and store the features of the new data stream.

[0085] Optionally, after acquiring the features of multiple data streams, the acquisition unit 101 can also filter these features based on feature thresholds and store the filtered features of the data streams. By setting feature thresholds to filter features of data streams unrelated to the occurrence of sudden events, the amount of subsequent data processing can be reduced, saving computational resources.

[0086] The aforementioned feature threshold refers to a threshold set for the characteristics of a data stream. The acquisition unit 101 will only store the characteristics of data streams that reach the threshold. For example, the feature threshold could be that the burst size of the data stream is greater than TH1, or that the buffer occupancy / rate of the data stream is greater than TH2. The number of feature thresholds can also be one or more. It should be understood here that TH1 and TH2 only indicate that the values ​​are different and have no other meaning.

[0087] Optionally, after acquiring the features of multiple data streams, the acquisition unit 101 can also predict whether a sudden event is about to occur and the possible location of the sudden event based on the features of the data streams and the current operation of the network transmission path. It can then filter the features of the acquired multiple data streams based on one or more of the cache ID, port ID, queue ID, and data stream ID associated with the predicted location of the sudden event, and store the filtered data stream data. For example, if the acquisition unit 101 predicts that the port ID corresponding to the location of the sudden event is 12 and the cache ID is 5 based on the features of the data streams and the current operation of the network transmission path, then after acquiring the features of multiple data streams, the acquisition unit 101 will only store the features of the data stream that passes through port ID 12 and cache ID 5. It should be noted that the acquisition unit 101 can predict the possible location of the sudden event based on the features of the data streams and the current operation of the network transmission path using a prediction model and corresponding algorithm; this application does not limit its specific implementation.

[0088] By predicting in advance whether a sudden event will occur and the location of the event, the features of the data stream are filtered in advance before acquiring the features of the data stream stored in unit 101. This reduces the storage of features of the data stream that are irrelevant to the occurrence of the sudden event, thereby reducing the amount of subsequent data processing and saving computing resources.

[0089] Optionally, after acquiring the features of multiple data streams, the acquisition unit 101 can further filter the features of the data streams according to pre-configured conditions by the user, and store the filtered features of the data streams. For example, if the pre-configured conditions specify that the data stream with port ID 10 and cache ID 2 will pass through the data stream, then after acquiring the features of multiple data streams, the acquisition unit 101 will filter the features of the data streams according to the pre-configured conditions, and will only store the features of the data stream that passes through port ID 10 and cache ID 2.

[0090] It should be understood here that the conditions pre-configured by the user depend entirely on the user's configuration. Therefore, after the acquisition unit 101 filters the features of multiple data streams according to the pre-configured conditions, the features of the data streams obtained are also the features of the data streams specified or of interest to the user.

[0091] Optionally, the acquisition unit 101 can also combine the above-mentioned filtering methods for the characteristics of the data stream to meet more needs. The following is an example of one combination method, and other combination methods can be appropriately derived based on the following content.

[0092] After acquiring the features of multiple data streams, the acquisition unit 101 filters the features of these data streams according to feature thresholds and user-preconfigured conditions, and stores the filtered features. For example, if the feature thresholds are that the burst size of a data stream is greater than TH3 and the buffer occupancy / rate of the data stream is greater than TH4, and the user-preconfigured conditions allow the data stream with port ID 9 and buffer ID 1 to pass through, then the acquisition unit 101, after acquiring the features of multiple data streams, will filter the features of the data streams according to the feature thresholds and user-preconfigured conditions, storing only the features of the data stream with port ID 9, the data stream with buffer ID 1, the data stream burst size greater than TH3, and the data stream buffer occupancy / rate greater than TH4. By setting feature thresholds and user-preconfigured conditions, not only can features of data streams unrelated to the occurrence of burst events be filtered, reducing subsequent data processing volume and saving computing resources, but it can also ensure that the filtered features are those of the data streams of interest specified by the user.

[0093] The event monitoring unit 102 is used to acquire a first event, determine the influencing factors of the first event based on the first event, collect first data based on the first event, and send a notification of the first event to the acquisition unit 101 and send relevant information of the first event to the filtering unit 103.

[0094] The first event refers to the sudden event mentioned in the background art above. A sudden event is an abnormal event that occurs in the network transmission path. Common sudden events include packet loss events and backpressure events. The occurrence of sudden events will affect the stability and performance of network transmission. For a detailed description of the specific impact of sudden events, please refer to the relevant content in the background art above. This application will not elaborate further here.

[0095] The aforementioned first data includes one or more of the following: cache ID, port ID, queue ID, and data stream ID associated with the first event.

[0096] The aforementioned influencing factors include one or more of the following: cache, port, queue, and data stream. The event monitoring unit 102 internally stores a relationship table between the first event and its influencing factors. Once the event monitoring unit 102 determines the type of the first event, it can obtain the corresponding influencing factors based on this relationship table. For example, if the relationship table includes packet loss events and the corresponding influencing factors are cache, port, and queue, then after the event monitoring unit 102 determines that the first event is a packet loss event, it determines that the influencing factors of the first event include cache, port, and queue based on the relationship table.

[0097] The relevant information for the aforementioned first event includes the first data and the influencing factors of the first event.

[0098] The aforementioned first event notification is used to remind the acquisition unit 101 to send the characteristics of the stored data stream to the filtering unit 103. It should be understood that the aforementioned emergency event notification does not require the acquisition unit 101 to immediately send the characteristics of the stored data stream to the filtering unit 103. Since the characteristics of the data stream for a period of time after the emergency event occurs are also closely related to the occurrence of the emergency event, the acquisition unit 101 may send the characteristics of the stored data stream to the filtering unit 103 after receiving the emergency event notification at a certain interval.

[0099] The specific method by which the event monitoring unit 102 obtains the first event includes: the event monitoring unit 102 determines whether the first event has occurred based on whether the data packets in the network transmission path contain a marker for the first event. For example, it determines whether the network transmission path is congested by judging whether the data packets contain an Explicit Congestion Notification (ECN) marker, and it determines whether a backpressure event has occurred in the network transmission path by using a backpressure marker.

[0100] The ECN tag mentioned above is a mechanism used in Transmission Control Protocol / Internet Protocol (TCP / IP) networks to notify of network congestion. When a network device (such as a router) detects congestion on the network transmission path, it can add an ECN tag to the ECN field of the IP header of data packets in the network transmission path. The ECN tag has two states: set and not set. When not set, it indicates that the network transmission path is unobstructed; when set, it indicates that the network transmission path is congested.

[0101] The aforementioned backpressure flag is a data flow control mechanism, typically used in network transmission. When the rate at which data flows to the receiver exceeds the receiver's processing capacity, i.e., when the receiver cannot continue to receive more data, the receiver can send information with a backpressure flag to the sender through specific protocol information, TCP windows, or other mechanisms to notify the sender to temporarily stop sending data in order to avoid data packet loss.

[0102] Optionally, the event monitoring unit 102 can also determine whether a first event has occurred in the network transmission path based on user-defined conditions. For example, the user can set corresponding thresholds, including port rate thresholds and port burst size thresholds. After the event monitoring unit 102 detects that the port rate of a certain station in the network transmission path is lower than the port rate threshold or the port burst size exceeds the port burst size threshold, the event monitoring unit 102 determines that a first event has occurred at that station location.

[0103] Optionally, the event monitoring unit 102 may also be located outside the data stream feature acquisition device 100, that is, the event monitoring unit 102 is a component of other devices. After the other devices obtain the relevant information of the first event, they send the notification of the first event to the acquisition unit 101 in the data stream feature acquisition device 100, and send the relevant information of the first event to the filtering unit 103 in the data stream feature acquisition device 100.

[0104] The filtering unit 103 is used to determine the first flow characteristics and the first attribute characteristics based on the first data and influencing factors, and to report the first flow characteristics and the first attribute characteristics to the analyzer.

[0105] For details regarding the aforementioned first data and influencing factors, please refer to the relevant description in the event monitoring unit 102 mentioned above; this application will not elaborate further here.

[0106] The aforementioned first flow characteristic and first attribute characteristic are characteristics of the data flow that may contribute to the occurrence of the first event.

[0107] The specific method by which the filtering unit 103 filters the features of the stored data stream according to the filtering conditions includes: the filtering unit 103 updates the information in the influencing factor record table according to the first data and influencing factors; filtering the features of multiple data streams in the acquisition unit 101 based on the information in the influencing factor record table, and filtering and recording the features of the data streams that meet the information in the influencing factor record table; the features of the data stream finally recorded are the first flow feature and the first attribute feature; after acquiring the first flow feature and the first attribute feature, the filtering unit 103 transmits the feature to the analyzer 200. It should be understood here that the influencing factor record table exists in the data stream feature acquisition device 100, and the filtering unit 103 has the ability to access and read / write the influencing factor record table. The filtering unit 103 filters the features of the data stream by reading the information in the influencing factor record table and according to the content of the influencing factor record table.

[0108] Optionally, the filtering unit 103 can also update the influencing factor record table with the user-configured filtering conditions (such as the user-configured port ID, priority ID, and queue ID), thereby enabling the filtering unit 103 to filter out the features of the data streams that the user is interested in. For example, the user can send the user-configured filtering conditions to the filtering unit 103 wirelessly, and after receiving the user-configured filtering conditions, the filtering unit 103 updates them to the influencing factor record table.

[0109] Optionally, after updating the port ID, priority ID, and queue ID associated with the occurrence location of the first event (hereinafter referred to as the filtering condition of the first event) and the user-configured port ID, priority ID, and queue ID (hereinafter referred to as the filtering condition of the user) to the influencing factor record table, the user can also pre-set the relationship between the two filtering conditions in the influencing factor record table, including setting the characteristics of the data flow to simultaneously meet the filtering conditions of the first event and the filtering conditions configured by the user, or setting the characteristics of the data flow to only meet one of the filtering conditions of the first event or the filtering conditions configured by the user.

[0110] When the characteristics of the data stream need to simultaneously satisfy the filtering conditions of the first event and the filtering conditions configured by the user, the filtering unit 103 filters the characteristics of the data stream according to the above filtering conditions. The advantage of doing so is that the characteristics of the filtered data stream satisfy both the location where the first event occurs and the location that the user is interested in.

[0111] When the characteristics of the data stream only need to satisfy one of the filtering conditions of the first event or the filtering conditions configured by the user, the filtering unit 103 selects the characteristics of the data stream that satisfy the filtering conditions of the first event and the filtering conditions configured by the user according to the above filtering conditions. The advantage of doing so is that it takes into account the location where the first event itself occurs, and does not miss the location that the user is interested in.

[0112] Optionally, the filtering unit 103 can also receive statistical values ​​of traffic characteristics associated with multiple data streams sent by the acquisition unit 101. For a detailed description of the statistical values ​​of traffic characteristics, please refer to the corresponding description in the acquisition unit 101 above. This application will not elaborate further here. After acquiring the influencing factors of the first event, the filtering unit 103 determines the traffic characteristics corresponding to those influencing factors. Then, based on the determined statistical values ​​of the traffic characteristics, it sorts the characteristics of multiple data streams and updates the attribute characteristics corresponding to the top three data streams in the sequence to the aforementioned influencing factor record table. For example, if the filtering unit 103 acquires that the influencing factor of the first event is a data stream, and this influencing factor corresponds to the flow-level characteristic (hereinafter referred to as the burst of the data stream) in the traffic characteristics, then... Taking size as an example, the filtering unit 103 sorts the features of multiple data streams according to the statistical values ​​corresponding to the burst size of the data streams, and determines that the top 3 data streams include data stream 1, data stream 2 and data stream 3. Among them, the attribute features of data stream 1 include data stream id 1, cache id 4, port id 17 and queue id 5; the attribute features of data stream 2 include data stream id 2, cache id 4, port id 18 and queue id 15; the attribute features of data stream 3 include data stream id 3, cache id 4, port id 19 and queue id 16. The filtering unit 103 updates the attribute features corresponding to the above data stream 1, data stream 2 and data stream 3 to the influencing factor record table. It should be understood here that the above scheme is only an example provided by this application and should not be regarded as a specific limitation. For example, the above filtering unit 103 can also take the attribute features corresponding to the top 2 data streams in the sequence and update them to the above influencing factor record table. It can also set a special threshold. The filtering unit 103 selects the attribute features of the data stream and updates them to the influencing factor record table based on whether the statistical value of the determined flow feature is greater than the threshold.

[0113] For details on how analyzer 200 analyzes the characteristics of the data stream, please refer to the following: Figure 7 as well as Figure 8 The relevant content in this application will not be described in detail here.

[0114] In summary, this application provides a data stream feature acquisition device. This device acquires and stores features of multiple data streams in a network transmission path through an acquisition unit 101; an event monitoring unit 102 monitors for the occurrence of any sudden events in the network transmission path; and in the event of a sudden event, sends a notification to the acquisition unit 101 and relevant information about the sudden event to a filtering unit 103; after receiving the notification, the acquisition unit 101 sends the stored features of the multiple data streams to the filtering unit 103; and after receiving the features of the multiple data streams and the relevant information about the sudden event, the filtering unit 103 filters the features of the multiple data streams based on the relevant information about the sudden event, thereby acquiring features of data streams that may contribute to the occurrence of the sudden event. Compared to directly collecting the characteristics of data streams that trigger emergencies, the aforementioned data stream characteristic acquisition device stores the characteristics of multiple data streams that may cause an emergency before the emergency occurs through the acquisition unit 101. After the emergency occurs, the filtering unit 103 filters the stored data stream characteristics according to the filtering conditions, and finally obtains the characteristics of data streams directly related to the occurrence of the emergency before and after the emergency, as well as the characteristics of data streams indirectly related to the occurrence of the emergency. These data stream characteristics can comprehensively reproduce the impact of the emergency, thereby supporting the analysis of the cause of the emergency.

[0115] The structure and implementation of the data stream feature acquisition device provided in this application have been introduced above. The following section will combine... Figures 2 to 5 The present application introduces a data stream feature acquisition method, and the data stream feature acquisition device 100 described above can implement the following data stream feature acquisition method.

[0116] See Figure 2 , Figure 2 This is a schematic diagram illustrating the steps of a data stream feature acquisition method provided in this application, as shown below. Figure 2 As shown, this data stream feature acquisition method includes:

[0117] S201; The data stream feature acquisition device acquires the traffic features and attribute features associated with multiple first data streams.

[0118] This step can be implemented by the acquisition unit 101 in the data stream feature acquisition device 100 described above.

[0119] The first data stream, traffic characteristics, and attribute characteristics mentioned above can be found in the above... Figure 1 The relevant description of the acquisition unit 101 is not elaborated here.

[0120] The data flow feature acquisition device acquires traffic features and attribute features associated with multiple first data flows, specifically including: the data flow feature acquisition device acquires the features and status of each data packet in the network transmission path at certain time intervals; by analyzing and calculating the features and status of each data packet, it acquires the features of the data flow corresponding to each data packet; and stores the features of the data flow.

[0121] For details regarding the characteristics and status of the aforementioned data packets, please refer to the relevant description at the acquisition unit 101. This application will not elaborate further here.

[0122] The following is combined Figure 3 This section details the process by which the data stream feature acquisition device acquires the traffic features and attribute features associated with multiple first data streams. Figure 3 This is a schematic diagram illustrating the process by which a data stream feature acquisition device, as provided in this application, acquires features from multiple data streams. Figure 3 As shown, the process by which the data stream feature acquisition device acquires the traffic features and attribute features associated with multiple first data streams may include the following steps:

[0123] S301: The data stream feature acquisition device acquires the features of the data stream based on the features and state of the data packets.

[0124] For details on the process of obtaining data stream features based on packet characteristics and packet state, please refer to the relevant description at the acquisition unit 101 above, which will not be elaborated further here.

[0125] Optionally, while acquiring the characteristics of multiple data streams, the data stream feature acquisition device can also record statistical values ​​of the traffic characteristics associated with each data stream. These statistical values ​​can be one or more of the following: maximum value, current value, and cumulative value. By recording the statistical values ​​of the traffic characteristics associated with each data stream, it is possible to quickly determine which data streams have experienced significant changes in their associated traffic characteristics within a period before and after an event. In other words, the occurrence of an event is often related to the data streams with the largest changes or the largest cumulative changes in their characteristics within a period before and after the event.

[0126] S302: The data stream feature acquisition device determines whether the features of the data stream meet the screening conditions.

[0127] After acquiring the features of a data stream, the data stream feature acquisition device can first perform preliminary screening of these features, and then store the screened features. By eliminating features of data streams that are less correlated with the occurrence of sudden events, the amount of data processing in subsequent steps is reduced, thereby saving overall computing resources. For example, packet loss events and backpressure events are often associated with data streams with large burst sizes or large cache usage. After acquiring the features of the data stream, the data stream feature acquisition device filters the features based on thresholds corresponding to burst size and cache usage, and then stores the screened features.

[0128] Optionally, the data stream feature acquisition device can filter the features of the data stream based on feature thresholds. For details regarding feature thresholds, please refer to the relevant description in the acquisition unit 101 described above; this application will not elaborate further here.

[0129] Optionally, after acquiring the characteristics of the data stream, the data stream feature acquisition device can also predict whether a sudden event is about to occur and its possible location based on the acquired data stream characteristics and the current operation of the network transmission path. Furthermore, it can filter the acquired data stream features based on one or more of the cache ID, port ID, queue ID, and data stream ID associated with the predicted sudden event. By predicting whether a sudden event will occur and its location in advance, and by filtering the data stream features before the data stream feature acquisition device stores the features, the amount of subsequent data processing can be significantly reduced, saving computational resources.

[0130] Optionally, after acquiring the features of the data stream, the data stream feature acquisition device can also filter the features of the data stream according to pre-configured conditions by the user. It should be understood that the pre-configured conditions depend entirely on the user's settings, and the filtered data stream features are also those of the data stream that the user specifies, is interested in, and wants to understand.

[0131] Optionally, after acquiring the features of the data stream, the data stream feature acquisition device can combine the aforementioned various feature filtering methods before filtering the data stream features, thereby meeting more diverse needs. For example, by simultaneously setting feature thresholds and user-preset conditions, it can not only filter features of data streams unrelated to the occurrence of emergencies, reducing subsequent data processing volume and saving computing resources, but also ensure that the filtered data stream features are those of the data streams of interest specified by the user.

[0132] S303: The data stream feature acquisition device encapsulates the features of the data stream and the timestamp into a message.

[0133] The aforementioned message is the basic unit of communication between software and components, used for information transfer between programs.

[0134] After filtering the features of the data stream according to the filtering criteria, the data stream feature acquisition device encapsulates the features and timestamps of the data stream into a message, so that the data stream feature acquisition device can send the features of the data stream and store them in the memory. It should be understood that the above-mentioned encapsulation of the timestamp into the message is to indicate the storage time of the data stream features, which facilitates the subsequent discarding of data stream features.

[0135] The aforementioned memory can be located inside or outside the data stream feature acquisition device; this application does not limit the specific location of the memory. The memory can be implemented as a First-Input-First-Out (FIFO) memory or a ring buffer. Both FIFO and ring buffers store a certain amount of data; when the stored data reaches its storage limit, newly stored data overwrites the oldest data.

[0136] S304: The memory determines whether its storage space is full. If the memory determines that the storage space is full, it discards the oldest message and stores the newest message; if the memory determines that the storage space is not full, it stores the newest message.

[0137] By configuring the memory to only store a certain amount of data streams, we can ensure that the amount of data processed later is not too large. Furthermore, since the data streams most correlated with the occurrence of sudden events are those within a certain period before and after the event, setting a certain storage space for the memory can also filter out some data streams with low correlation to the occurrence of sudden events.

[0138] Optionally, the data stream feature acquisition device can store only the features of the data stream within a certain time period. This is because of the dynamic nature of data streams; since data stream features are constantly changing, using features from a data stream with a large time interval before or after the event to analyze the event would introduce unnecessary features, reducing the accuracy of subsequent analysis results. Therefore, the data stream feature acquisition device only needs to process the features of the data stream within a certain period before and after the event. For example, the memory can be set to only store data stream features for 30 seconds; if the difference between the timestamp corresponding to a stored message and the current time exceeds 30 seconds, the memory will delete the stored message.

[0139] Optionally, multiple memories can be configured, with different memories used to store features of different types of data streams. By configuring multiple memories, the data stream feature acquisition device can filter out appropriate data stream features under different sudden events. Taking a memory comprising coarse-grained and fine-grained memories as an example, the coarse-grained memory stores coarse-grained data stream features. The data stream feature acquisition device consumes fewer computational resources when processing coarse-grained data stream features, thus it can periodically send these features to the analyzer to predict the location and probability of sudden events even when no sudden event has occurred. The fine-grained memory stores fine-grained data stream features. While the data stream feature acquisition device consumes more computational resources when processing these features, the analyzer's results based on these features are more accurate. Therefore, the data stream feature acquisition device can send these fine-grained data stream features to the analyzer to analyze the cause of sudden events when they occur. Here, coarse-grained data stream characteristics refer to the characteristics of the data stream calculated at a time granularity, including the average, maximum, and minimum values ​​of the data stream characteristics over a period of time; fine-grained characteristics are the characteristics of the data stream without any processing, that is, the characteristics of the data stream at each moment. After acquiring the characteristics of the data stream, the data stream feature acquisition device obtains coarse-grained data stream characteristics by backing up the data stream characteristics and processing the backed-up data stream characteristics; subsequently, the data stream feature acquisition device sends the coarse-grained data stream characteristics to a coarse-grained memory and the fine-grained data stream characteristics to a fine-grained memory.

[0140] S202: The data stream feature acquisition device acquires the first event, determines the influencing factors of the first event based on the first event, and collects the first data based on the first event.

[0141] This step can be implemented by the event monitoring unit 102 in the aforementioned data stream feature acquisition device 100.

[0142] For details regarding the aforementioned first event, influencing factors, and first data, please refer to the above. Figure 1 The relevant descriptions of the event monitoring unit 102 are not elaborated here.

[0143] The specific methods by which the data flow feature acquisition device acquires the first event include: the data flow feature acquisition device determines whether the first event has occurred based on whether the data packets in the network transmission path contain a marker for the first event. For example, it determines whether the network transmission path is congested by judging whether the data packets contain an Explicit Congestion Notification (ECN) marker and whether a backpressure event has occurred in the network transmission path by using a backpressure marker. After determining that the first event has occurred, the data flow feature acquisition device determines the influencing factors of the first event based on the type of the first event and acquires the first data corresponding to the first event.

[0144] Optionally, the data flow feature acquisition device can also determine whether a first event has occurred in the network transmission path based on user-defined conditions. For example, the user can set corresponding thresholds, including port rate thresholds and port burst size thresholds. After the data flow feature acquisition device detects that the port rate at a certain station in the network transmission path is lower than the port rate threshold or the port burst size exceeds the port burst size threshold, the data flow feature acquisition device determines that a first event has occurred at that station location.

[0145] S203: The data flow feature acquisition device determines the first flow feature and the first attribute feature based on the first data and influencing factors, and reports the first flow feature and the first attribute feature to the analyzer.

[0146] This step can be implemented by the filtering unit 103 in the aforementioned data stream feature acquisition device 100.

[0147] The data stream feature acquisition device determines the first flow feature and the first attribute feature based on the first data and influencing factors, and reports the first flow feature and the first attribute feature to the analyzer in the following specific ways: the data stream feature acquisition device updates the information in the influencing factor record table based on the first data and influencing factors; filters the features of multiple stored data streams based on the information in the influencing factor record table, and selects and records the features of data streams that meet the information in the influencing factor record table; the features of the data streams finally recorded are the first flow feature and the first attribute feature; after acquiring the first flow feature and the first attribute feature, the data stream feature acquisition device transmits the first flow feature and the first attribute feature to the analyzer.

[0148] The following is combined Figure 4 This section details the process by which the data stream feature acquisition device determines the first flow characteristics and the first attribute characteristics based on the first data and influencing factors, and then reports these characteristics to the analyzer. Figure 4 This is a schematic diagram of a data stream upload process provided in this application, such as... Figure 4 As shown, the data flow feature acquisition device determines the first flow feature and the first attribute feature based on the first data and influencing factors, and reports the first flow feature and the first attribute feature to the analyzer, which may include the following steps:

[0149] S401: The data stream feature acquisition device judges the features of the first data stream based on the influencing factors, obtains all the second features that meet the requirements, and updates the second features to the influencing factor record table.

[0150] After acquiring the influencing factors of the first event, the data stream feature acquisition device determines the traffic characteristics corresponding to the influencing factors. Then, it sorts the characteristics of multiple data streams according to the statistical values ​​of the determined traffic characteristics. Finally, the data stream feature acquisition device selects the attribute characteristics (i.e., the second characteristics mentioned above) of the data stream based on whether the statistical values ​​of the determined traffic characteristics are greater than a threshold and updates them to the influencing factor record table.

[0151] S402: The data stream feature acquisition device updates the influencing factor record table based on the first data.

[0152] The aforementioned first data includes one or more of the following: port ID, priority ID, queue ID, and data stream ID associated with the location where the first event occurred. For a detailed description of the first data, please refer to the above. Figure 1 The contents of the intermediate filter unit 103 will not be described in detail here.

[0153] The first data is explained below using the port ID associated with the location of the first event. The data stream ID, queue ID, and priority ID associated with the location of the first event can be appropriately derived from the following content.

[0154] See Figure 5 , Figure 5 This application provides a schematic diagram illustrating the occurrence of an emergency, such as... Figure 5 As shown, in Figure 5The site has two receiving ports (receiving port 1 and receiving port 2), two sending ports (sending port 1 and sending port 2, where sending port 1 has an ID of 1011 and sending port 2 has an ID of 1012), and a buffer memory. Three data streams (data stream 1, data stream 2, and data stream 3) pass through the site. Data streams 1 and 2 enter the site's buffer memory through receiving port 1 and exit through sending port 1; data stream 3 enters the site's buffer memory through receiving port 2 and exits through sending port 2. Note that the buffer memory corresponding to different sending port IDs is different; therefore, the buffer memory for data stream 3 is not the same as the buffer memory for data streams 1 and 2.

[0155] When data stream 2 experiences packet loss at the station's sending port 1 (in Figure 5 In the diagram, the lost data packets in data stream 2 are represented by dashed circles. At this time, a packet loss event occurs at the sending port 1 of the station. The port ID associated with the location of the packet loss event is 1011 (the ID of sending port 1).

[0156] The data stream feature acquisition device updates the filtering conditions in the influencing factor record table based on one or more of the information such as port ID, priority ID, and queue ID associated with the location of the sudden event.

[0157] Optionally, users can configure corresponding filtering conditions (such as port ID, priority ID, and queue ID) in advance in the influencing factor record table, so that the data stream feature acquisition device can filter out the features of the data stream that the user is interested in.

[0158] Optionally, after the data stream feature acquisition device updates the port ID, priority ID, queue ID, and other information associated with the location of the aforementioned incident (hereinafter referred to as the incident filtering conditions) and the user-configured port ID, priority ID, and queue ID (hereinafter referred to as the user-configured filtering conditions) to the influencing factor record table, the user can also set the relationship between these two filtering conditions. This includes setting the data stream features to simultaneously meet both the incident filtering conditions and the user-configured filtering conditions, or setting the data stream features to only meet one of the incident filtering conditions or the user-configured filtering conditions.

[0159] When the characteristics of a data stream need to simultaneously satisfy both the filtering conditions for sudden events and the filtering conditions configured by the user, the data stream feature acquisition device filters the data stream features based on the above filtering conditions. The features of the data stream that meet the filtering conditions for sudden events and the filtering conditions configured by the user are the intersection. The advantage of doing this is that the features of the filtered data stream satisfy both the location where the sudden event occurs and the location that the user is interested in.

[0160] When the data stream features only need to satisfy one of the filtering conditions for sudden events or the filtering conditions configured by the user, the data stream feature acquisition device selects the features of the data stream based on the above filtering conditions as the union of the filtering conditions for sudden events and the filtering conditions configured by the user. The advantage of doing this is that it takes into account both the location where the sudden event itself occurs and does not miss the location that the user is interested in.

[0161] S403: The data stream feature acquisition device filters out the features of the data stream that meet the filtering conditions in the influencing factor record table from the stored messages.

[0162] After the influencing factor record table is updated and the stored message is received from the memory, the data stream feature acquisition device filters the received message according to the filtering conditions in the influencing factor record table, and selects and records the features of the data stream that meet the filtering conditions in the influencing factor record table.

[0163] S404: The data stream feature acquisition device sends messages that meet the filtering conditions in the influencing factor record table to the analyzer.

[0164] After filtering the data stream features according to the filtering conditions in the influencing factor record table, the data stream features acquire the first flow feature and the first attribute feature (i.e., messages that meet the filtering conditions in the influencing factor record table). The data stream feature acquisition device sends the first flow feature and the first attribute feature to the analyzer. The analyzer analyzes the features of the data stream to obtain the cause of the sudden event.

[0165] For the specific steps of the above analyzer in obtaining the cause of the sudden event based on the characteristics of the data stream, please refer to the following: Figure 7 as well as Figure 8 The relevant content in this application will not be described in detail here.

[0166] S405: Update of the data stream feature acquisition device rollback influencing factor record table.

[0167] After the data stream feature acquisition device sends the first flow feature and the first attribute feature to the analyzer, in order to facilitate the next update of the filtering conditions in the influencing factor record table by the data stream feature acquisition device, the data stream feature acquisition device needs to roll back the filtering conditions in the influencing factor record table or clear the records in the influencing factor record table.

[0168] In summary, the overall operational logic of the data stream feature acquisition method provided in this application can be found in the following description. Figure 6 The relevant content in [the document / document].

[0169] See Figure 6 , Figure 6 This is a schematic diagram illustrating the operational logic of a data stream feature acquisition method provided in this application, such as... Figure 6 As shown, the operational logic of the data stream feature acquisition method provided in this application includes:

[0170] ① Calculation of data stream characteristics: The data stream characteristic acquisition device calculates and obtains the characteristics of multiple data streams based on the characteristics and state of the data packets.

[0171] ② Data stream feature storage: The data stream feature acquisition device stores the features of multiple data streams it acquires in the memory, waiting for notification of the first event.

[0172] ③ Updating the feature record table of data streams: The data stream feature acquisition device acquires and records the statistical values ​​of traffic features associated with multiple data streams, and updates the statistical values ​​corresponding to each data stream in the feature record table of data streams according to the set time window.

[0173] ④ Event monitoring: The data stream feature acquisition device determines whether the first event has occurred in the network transmission path based on information such as the status of the data packets, the cache status on the network transmission path, and the device status. If the first event has occurred in the network transmission path, the first event notification is sent to the memory to obtain the features of multiple data streams stored in the memory; and obtains relevant information about the first event, including the influencing factors of the first event and the first data.

[0174] ⑤ Acquisition of the second feature: The data stream feature acquisition device acquires the second feature based on the influencing factors of the first event and the statistical values ​​corresponding to multiple data streams recorded in the data stream feature record table.

[0175] ⑥ Updating the Influencing Factor Record Table: The data stream feature acquisition device updates the filtering conditions in the influencing factor record table based on the first data and the second feature.

[0176] ⑦ Filtering based on filtering conditions: The characteristics of multiple data streams obtained from the memory are filtered according to the filtering conditions in the influencing factor record table to obtain the first flow characteristics and the first attribute characteristics. Finally, the first flow characteristics and the first attribute characteristics are sent to the analyzer to locate and confirm the cause of the first event.

[0177] In summary, this application provides a data stream feature acquisition method. This method uses a data stream feature acquisition device to acquire and store the traffic and attribute features associated with multiple data streams before a sudden event occurs. After the sudden event occurs, the data stream feature acquisition device analyzes the event to obtain the corresponding influencing factors and related data. Finally, based on the influencing factors and related data, the acquired traffic and attribute features of the multiple data streams are filtered to obtain the traffic and attribute features of data streams that may have a driving effect on the occurrence of the sudden event. By implementing the above method, the characteristics of data streams directly related to the occurrence of the sudden event and those indirectly related to it are finally obtained. These data stream features can comprehensively reproduce the impact of the sudden event, thereby supporting the analysis of the causes of the sudden event.

[0178] See Figure 7 , Figure 7 This is a schematic diagram illustrating the analysis results of a data stream characteristic provided in this application, such as... Figure 7 As shown in the figure, this analysis uses the burst size of the data stream to determine the causes of burst events. Figure 7 The diagram includes coordinate axes corresponding to two data streams (data stream 1 and data stream 2). The horizontal axis represents time, and the vertical axis represents the burst size. The length of the shaded area on the horizontal axis represents the burst duration of the data stream, which refers to the duration of the data stream burst.

[0179] from Figure 7 It can be observed that at the moment of the packet loss event, data stream 2 experienced a significant burst in duration, indicating a substantial change in its data flow. Furthermore, no other data streams showed similar changes at the location of the packet loss event. Therefore, it can be confirmed that data stream 2 experienced packet loss, leading to the packet loss event. However, from... Figure 7 It can also be found that before the packet loss event occurred, data stream 1 showed a significant burst of duration, indicating that the data traffic of data stream 1 changed significantly. Considering that the packet loss event occurred at the destination port, it can be determined that the data traffic of data stream 1 suddenly increased dramatically in a short period of time, occupying a large amount of cache memory in front of the destination port, which led to the packet loss phenomenon in data stream 2, and thus caused the packet loss event to occur.

[0180] See Figure 8 , Figure 8 This is a schematic diagram illustrating the analysis results of another data stream characteristic provided in this application, such as... Figure 8 As shown in the figure, this analysis uses data stream buffer usage to determine the causes of sudden events. Figure 8The table includes coordinate axes corresponding to two data streams (data stream 1 and data stream 2), where the horizontal axis represents time and the vertical axis represents the cumulative number of bytes of the data stream in the cache memory.

[0181] from Figure 8 It can be observed that at the moment of the packet loss event, the number of bytes accumulated in the cache memory for data stream 2 stopped increasing, indicating that data stream 2 was unable to store data packets in the cache memory, meaning that data stream 2 experienced data packet loss, thus leading to the packet loss event. However, from... Figure 8 It can also be found that before the packet loss event occurred, data stream 1 had accumulated a large number of bytes in the cache memory. At the time of the packet loss event, the number of bytes accumulated in the cache memory of data stream 1 no longer changed, indicating that data stream 1 had occupied a large amount of cache memory, which caused the data packets of data stream 2 to be unable to be stored in the cache memory in time, resulting in the loss of data packets in data stream 2, which in turn led to the packet loss event.

[0182] The above text combines Figures 2 to 6 The data stream feature acquisition method provided in this application is described in detail below, and will be combined with... Figure 9 and Figure 10 This describes the specific hardware structure corresponding to the data stream feature acquisition device provided in this application.

[0183] Figure 9 This is a schematic diagram of the hardware structure of a data stream feature acquisition device provided in this application. Figure 9 The computing device 900 shown can perform the corresponding steps executed by the data stream feature acquisition device in the method of the above embodiments.

[0184] Furthermore, the computing device 900 includes a processor 901, a storage unit 902, a storage medium 903, and a communication interface 904. The processor 901, the storage unit 902, the storage medium 903, and the communication interface 904 communicate via a bus 905, and also via other means such as wireless transmission.

[0185] Processor 901 comprises multiple general-purpose processors, such as CPUs, NPUs, or a combination of CPUs and hardware chips. The aforementioned hardware chips are application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The aforementioned PLDs are complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), data processing units (DPUs), systems-on-chips (SoCs), or any combination thereof. Processor 901 executes various types of digital storage instructions, such as software or firmware programs stored in storage unit 902, enabling computing device 900 to provide a wide range of services.

[0186] In a specific implementation, as one embodiment, the processor 901 includes one or more CPUs, for example... Figure 9 CPU0 and CPU1 are shown in the diagram.

[0187] In a specific implementation, as one example, the computing device 900 also includes multiple processors, for example... Figure 9 The processors 901 and 906 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor refers to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0188] Storage unit 902 is used to store program code, and its execution is controlled by processor 901 to perform the above-mentioned tasks. Figures 2 to 6 The processing steps of the data stream feature acquisition method in any embodiment are described. The program code includes one or more software units. The one or more software units mentioned above are... Figure 1 In this embodiment, the acquisition unit 101, event monitoring unit 102, and filtering unit 103 are specifically configured to perform... Figure 2 Step S201 in the embodiment and Figure 3 The optional step in the process, the event monitoring unit 102 is used to perform Figure 2 Step S202 in the embodiment and Figure 4 In the optional steps, the filtering unit 103 is used to perform Figure 2 Step S202 in the embodiment and Figure 4 The optional steps in this process will not be elaborated here.

[0189] Storage unit 902 includes read-only memory and random access memory, and provides instructions and data to processor 901. Storage unit 902 also includes non-volatile random access memory. Storage unit 902 is volatile memory or non-volatile memory, or a combination of both. The non-volatile memory is read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory is random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It can also refer to hard disks, USB flash drives, flash memory, SD cards, Memory Sticks, etc., where hard disks include hard disk drives (HDDs), solid-state drives (SSDs), mechanical hard disks (HDDs), etc., and this application does not specifically limit the types used.

[0190] Storage medium 903 is a carrier for storing data, such as hard disk, USB flash drive, flash memory, SD card, memory stick, etc. The hard disk can be a hard disk drive (HDD), solid state disk (SSD), mechanical hard disk (HDD), etc. This application does not make specific limitations.

[0191] The communication interface 904 is a wired interface (e.g., an Ethernet interface), an internal interface (e.g., a Peripheral Component Interconnect express (PCIe) bus interface), a wired interface (e.g., an Ethernet interface), or a wireless interface (e.g., a cellular network interface or a wireless LAN interface) for communicating with other servers or units.

[0192] The 905 bus is a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, Unified Bus (Ubus or UB), Compute Express Link (CXL), Cache Coherent Interconnect for Accelerators (CCIX), etc. The 905 bus is divided into address bus, data bus, and control bus.

[0193] In addition to the data bus, bus 905 also includes the power bus, control bus, and status signal bus. However, for clarity, all buses are labeled as bus 905 in the diagram.

[0194] It needs to be explained that, Figure 9 This is merely one possible implementation of an embodiment of this application. In practical applications, the computing device 900 may include more or fewer components, and this is not a limitation. For content not shown or described in the embodiments of this application, please refer to the foregoing. Figures 1 to 8 The relevant descriptions in the embodiments will not be repeated here.

[0195] Understandable, Figure 9The computing device 900 shown is merely a simplified design of a data stream feature acquisition device. In practical applications, the data stream feature acquisition device can also include any number of communication interfaces, processors, or storage units.

[0196] Figure 10 This is a schematic diagram of the hardware structure of another data stream feature acquisition device according to this application. Figure 10 The computing device 1000 shown can execute the corresponding steps performed by the data stream feature acquisition device in the method of the above embodiments.

[0197] like Figure 10 As shown, the computing device 1000 includes: a main control board 1010, a switching network board 1020, an interface board 1030, and an interface board 1040. The main control board 1010, interface board 1030, interface board 1040, and switching network board 1020 communicate with each other via a system bus connected to the system backplane. The main control board 1010 performs system management, equipment maintenance, and protocol processing functions; the switching network board 1020 performs data exchange between the interface boards (also called line cards or service boards); and the interface boards 1030 and 1040 provide various service interfaces (e.g., POS interface, GE interface, ATM interface, etc.) and forward data packets. In one possible implementation, the computing device 1000 is a controller, a network management device, or a server.

[0198] The interface board 1030 may include a central processing unit 1031, a forwarding table entry memory 1034, a physical interface card 1033, and a network processor 1032. The central processing unit 1031 is used to control and manage the interface board and communicate with the central processing unit 1031 on the main control board; the forwarding table entry memory 1034 is used to store forwarding table entries; the physical interface card 1033 is used to acquire and transmit data stream characteristics; and the network processor 1032 is used to control the acquisition and transmission rate of data stream characteristics by the physical interface card 1033 according to the forwarding table entries.

[0199] Specifically, the physical interface card 1033 is used to send data stream characteristics (including first attribute characteristics and first traffic characteristics) to the analyzer. Specifically, the central processing unit 1031 is used to control the network processor 1032 to send data stream characteristics to the analyzer via the physical interface card 1033.

[0200] Optionally, the central processing unit 1031 sends the characteristics of the stored data stream to the central processing unit 1011, the central processing unit 1011 filters the characteristics of the stored data stream, and the central processing unit 1011 sends the filtered characteristics of the data stream to the central processing unit 1031. The central processing unit 1031 controls the network processor 1032 to send the filtered characteristics of the data stream to the analyzer via the physical interface card 1033.

[0201] It should be understood that the actions on interface board 1040 in this embodiment are consistent with the actions on interface board 1030, and for the sake of simplicity, they will not be described again. It should be understood that the computing device 1000 in this embodiment may correspond to the functions and / or various steps implemented in the above method embodiments, and will not be described again here.

[0202] Furthermore, it should be noted that there may be one or more main control boards, including a primary and a backup main control board. There may also be one or more interface boards; the stronger the data processing capability of the data flow feature acquisition device, the more interface boards it can provide. Each interface board may also have one or more physical interface cards. There may be no switching network board, or one or more; multiple boards can share the load for redundancy and backup. In a centralized forwarding architecture, the data flow feature acquisition device may not need a switching network board, as the interface boards handle the processing of the entire system's business data. In a distributed forwarding architecture, the data flow feature acquisition device can have at least one switching network board, enabling data exchange between multiple interface boards and providing high-capacity data exchange and processing capabilities. Therefore, the data access and processing capabilities of a distributed architecture data flow feature acquisition device are greater than those of a centralized architecture device. The specific architecture adopted depends on the specific network deployment scenario, and no limitations are made here.

[0203] Figure 11 This is a schematic diagram of a computing device cluster provided in this application, which includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0204] like Figure 11 As shown, the computing device cluster includes at least one computing device 1100. The memory 1103 of one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the data stream feature acquisition method.

[0205] In some possible implementations, the memory 1103 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data stream feature acquisition method. In other words, a combination of one or more computing devices 1100 can jointly execute the instructions for executing the data stream feature acquisition method.

[0206] It should be noted that the memory 1103 in different computing devices 1100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the resource migration system. That is, the instructions stored in the memory 1103 of different computing devices 1100 can implement the acquisition unit 101, the event monitoring unit 102, and the filtering unit 103. Specifically, the acquisition unit 101 is used to execute... Figure 2 Step S201 in the embodiment and Figure 3 The optional step in the process, the event monitoring unit 102 is used to perform Figure 2 Step S202 in the embodiment and Figure 4 In the optional steps, the filtering unit 103 is used to perform Figure 2 Step S203 in the embodiment and Figure 4 The optional steps in this process will not be elaborated here.

[0207] The computing device 1100 includes a processor 1101, a communication interface 1102, a memory 1103, and a bus 1104. Further descriptions of the processor 1101, communication interface 1102, memory 1103, and bus 1104 can be found in [reference needed]. Figure 9 The descriptions of the processor 901, storage unit 902, storage medium 903, communication interface 904, and bus 905 in the embodiments will not be repeated here.

[0208] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc., as described below. Figure 12 One possible implementation method is shown.

[0209] Figure 12 This is another schematic diagram of a computing device cluster provided in this application, such as... Figure 12 As shown, the two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this possible implementation, the memory 1103 in computing device 1100A stores instructions for implementing the acquisition unit 101 and the event monitoring unit 102. Meanwhile, the memory 1103 in computing device 1100B stores instructions for implementing the filtering unit 103.

[0210] It should be understood that Figure 12 The functions of computing device 1100A shown can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.

[0211] It needs to be explained that, Figure 12 The implementation shown may be implemented when the processing power of the computing device 1100A is insufficient, or when the storage space of the computing device 1100A is insufficient, or in other business scenarios. This application does not make any specific limitations.

[0212] This application also provides another type of computing device cluster. The interconnection relationships between the computing devices in this computing device cluster can be similarly referenced. Figure 11 and Figure 12 The connection method of the computing device cluster. The difference is that the memory 1103 of one or more computing devices 1100 in the computing device cluster can store the same instructions for executing the data stream feature acquisition method.

[0213] In some possible implementations, the memory 1103 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data stream feature acquisition method. In other words, a combination of one or more computing devices 1100 can jointly execute the instructions for executing the data stream feature acquisition method.

[0214] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a data stream feature acquisition method.

[0215] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives). The computer-readable storage medium includes instructions that instruct the computing device to perform a data stream feature acquisition method.

[0216] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes a plurality of computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0217] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of data stream feature collection, the method comprising: The method comprises: obtaining traffic features and attribute features associated with a plurality of first data streams; obtaining a first event; determining an influence factor of the first event according to the first event; collecting first data according to the first event; determining first traffic features and first attribute features according to the first data and the influence factor, wherein the traffic features comprise first traffic features, and the attribute features comprise first attribute features; reporting the first traffic features and the first attribute features to an analyzer.

2. The method of claim 1, wherein, The traffic features comprise cache level features, port level features, queue level features and flow level features, and the attribute features comprise cache identifiers (IDs), port IDs, queue IDs and data stream IDs composed of one or more of message header fields, wherein the message header fields comprise protocols, priorities, source Internet Protocol (IP) addresses, destination IP addresses, source port numbers and destination port numbers, and wherein the cache level features comprise cache occupancy / occupancy rate and cache packet loss / loss rate; the port level features comprise port rate, port cache occupancy / occupancy rate and port burst size; the queue level features comprise enqueue rate, dequeue rate, queue length, queue delay and queue packet loss / loss rate; and the flow level features comprise data stream rate, data stream burst size, data stream cache occupancy / occupancy rate and data stream packet loss / loss rate.

3. The method of claim 2, wherein, The influence factor comprises one or more of a cache, a port, a queue and a data stream.

4. The method according to claim 2 or 3, characterized in that, The first data comprises one or more of cache IDs, port IDs, queue IDs and data stream IDs associated with the first event.

5. The method according to any one of claims 2 to 4, characterized in that, The determination of the first traffic features and the first attribute features according to the first data and the influence factor comprises: obtaining statistical values of first features corresponding to the plurality of first data streams, wherein the first features are features corresponding to the influence factor in the traffic features, and the statistical values comprise one or more of maximum value, current value and cumulative value; filtering the statistical values of the first features corresponding to the plurality of first data streams according to a first filtering condition to obtain a plurality of second features, wherein the first filtering condition is that the statistical values are greater than a first threshold, and the second features are one or more of attribute features corresponding to the first data streams in the statistical values of the first features corresponding to the plurality of first data streams that satisfy the first filtering condition; filtering the traffic features and the attribute features associated with the plurality of first data streams according to the plurality of second features and the first data to obtain the first traffic features and the first attribute features, wherein the first traffic features and the first attribute features are parts of the traffic features and the attribute features associated with the plurality of first data streams that match the plurality of second features and the first data.

6. The method according to any one of claims 2 to 5, characterized in that, The obtaining of the traffic features and the attribute features associated with the plurality of first data streams comprises: obtaining traffic features and attribute features associated with a plurality of second data streams; According to the traffic features and attribute features associated with the second data streams, obtain relevant information of a predicted first event, the relevant information of the predicted first event including one or more of a cache id, a port id, a queue id, and a data stream id associated with the predicted first event; According to the relevant information of the predicted first event, filter the traffic features and attribute features associated with the second data streams, and obtain traffic features and attribute features associated with a plurality of first data streams; The traffic features and attribute features associated with the plurality of first data streams are traffic features and attribute features associated with the second data streams that match the relevant information of the predicted first event.

7. The method according to any one of claims 2 to 6, characterized in that, After obtaining the traffic features and attribute features associated with the plurality of first data streams, the method further includes: According to a first abnormal condition, filter the traffic features and attribute features associated with the plurality of first data streams, and store the traffic features and attribute features associated with the first data streams that satisfy the first abnormal condition, the first abnormal condition including that a cache occupancy rate of a data stream is greater than a second threshold value and / or a burst size of the data stream is greater than a third threshold value.

8. The method of claim 7, wherein, The obtaining of the traffic features and attribute features associated with the plurality of first data streams includes: According to a second abnormal condition, filter the traffic features and attribute features associated with the plurality of first data streams, and store the traffic features and attribute features associated with the first data streams that satisfy the second abnormal condition, the second abnormal condition including that the cache occupancy rate of the data stream is greater than a fourth threshold value and / or the burst size of the data stream is greater than a fifth threshold value; The traffic features and attribute features associated with the first data streams that satisfy the first abnormal condition are stored in a different storage area from the traffic features and attribute features associated with the first data streams that satisfy the second abnormal condition.

9. The method according to any one of claims 1 to 8, characterized in that, The first event includes an action-based event and a threshold-based event, wherein The action-based event includes one or more of a packet loss, congestion, an explicit congestion notification (ECN) mark, and a back pressure mark in a switch; The threshold-based event includes one or more of a cache occupancy rate greater than a sixth threshold value, a port rate greater than a seventh threshold value, and a queue packet loss rate greater than an eighth threshold value in a switch.

10. A data stream feature collection apparatus, characterized by, The apparatus includes: An obtaining unit configured to obtain traffic features and attribute features associated with a plurality of first data streams; An event monitoring unit configured to obtain a first event; The event monitoring unit is further configured to determine an influence factor of the first event according to the first event, and collect first data according to the first event; A filtering unit configured to determine first traffic features and first attribute features according to the first data and the influence factor, the traffic features including first traffic features, and the attribute features including first attribute features; The filtering unit is further configured to report the first traffic features and the first attribute features to an analyzer.

11. The apparatus of claim 10, wherein, The flow features include a cache level feature, a port level feature, a queue level feature and a flow level feature, the attribute features include a cache identifier id, a port id, a queue id and a data flow id composed of one or more of message header fields, the message header fields include a protocol, a priority, a source Internet Protocol IP, a destination IP, a source port number and a destination port number, wherein The cache level feature includes a cache occupancy / rate, a packet loss amount / rate in the cache; The port level feature includes a port rate, a cache occupancy / rate of the port, a burst size of the port; The queue level feature includes an enqueuing rate, a dequeuing rate, a length in the queue, a time delay in the queue, a packet loss amount / rate in the queue; The flow level feature includes a data flow rate, a data flow burst size, a cache occupancy / rate of the data flow, a packet loss amount / rate of the data flow.

12. The apparatus of claim 11, wherein, The influence factors include one or more of a cache, a port, a queue and a data flow.

13. The apparatus of claim 11 or 12, wherein, The first data includes one or more of a cache id, a port id, a queue id and a data flow id associated with the first event.

14. The apparatus of any one of claims 11 to 13, wherein, The filtering unit is specifically configured to: obtain statistical values of first features corresponding to a plurality of first data flows, the first features being features corresponding to the influence factors in the flow features, and the statistical values including one or more of a maximum value, a current value and a cumulative value; filter the statistical values of the first features corresponding to the plurality of first data flows according to a first filtering condition, obtain a plurality of second features, the first filtering condition being that the statistical values are greater than a first threshold, and the second features being one or more of attribute features of the first data flows corresponding to the statistical values of the first features corresponding to the plurality of first data flows that satisfy the first filtering condition; filter the flow features and the attribute features associated with the plurality of first data flows according to the plurality of second features and the first data, to obtain first flow features and first attribute features, the first flow features and the first attribute features being parts of the flow features and the attribute features associated with the plurality of first data flows that match the plurality of second features and the first data.

15. The apparatus of any one of claims 11 to 14, wherein, The obtaining unit is specifically configured to: obtain flow features and attribute features associated with a plurality of second data flows; obtain relevant information of a predicted first event according to the flow features and the attribute features associated with the plurality of second data flows, the relevant information of the predicted first event including one or more of cache ids, port ids, queue ids and data flow ids associated with the predicted first event; filter the flow features and the attribute features associated with the second data flows according to the relevant information of the predicted first event, to obtain flow features and attribute features associated with a plurality of first data flows; The flow features and the attribute features associated with the plurality of first data flows are flow features and attribute features associated with the second data flows that match the relevant information of the predicted first event.

16. The apparatus of any one of claims 11 to 15, wherein, The obtaining unit is further configured to: The traffic feature and the attribute feature associated with the first data flow satisfying the first abnormal condition are stored.

17. The apparatus of claim 16, wherein, The acquisition unit is further configured to: The traffic feature and the attribute feature associated with the first data flow satisfying the second abnormal condition are stored. The traffic feature and the attribute feature associated with the first data flow satisfying the first abnormal condition are stored in a storage area different from that of the traffic feature and the attribute feature associated with the first data flow satisfying the second abnormal condition.

18. The apparatus of any one of claims 10 to 17, wherein, The first event includes an action-based event and a threshold-based event, wherein The action-based event includes one or more of the following: packet loss, congestion, Explicit Congestion Notification (ECN) marking, and backpressure marking in the switch; The threshold-based event includes one or more of the following: buffer occupancy / occupancy rate greater than a sixth threshold, port rate greater than a seventh threshold, and queue packet loss / loss rate greater than an eighth threshold.

19. A computing device, comprising: The computing device includes a processor and a memory, the memory being configured to store instructions, and the processor being configured to execute the instructions to cause the computing device to implement the data flow feature collection method according to any one of claims 1 to 9.

20. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, each of the at least one computing device including a processor and a memory, and the processor of the at least one computing device being configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to implement the data flow feature collection method according to any one of claims 1 to 9.

21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and the instructions are executed by a computing device or a computing device cluster to implement the data flow feature collection method according to any one of claims 1 to 9.