A data acquisition method, device, equipment, medium and computer product
By grouping data into characteristic flow data in the 5G core network UPF traffic detection technology and using a variable step-size moving analysis window for data collection, the problem of redundant data caused by full data collection is solved, and efficient traffic collection and resource allocation are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE ZIJIN INNOVATION INST CO LTD
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-03
AI Technical Summary
Existing 5G core network UPF traffic detection technology relies on full data collection, resulting in a large amount of redundant data that increases the system burden and makes it impossible to efficiently collect valuable feature traffic.
By acquiring statistical information of the data to be collected, grouping it into characteristic flow data, and using the analysis window to move with variable step size for data collection, resources are dynamically allocated to avoid redundant data, and differentiated collection is performed for different traffic volumes.
It effectively avoids the increase of redundant data caused by full data collection, realizes dynamic resource allocation for different traffic and efficient characteristic traffic collection, and improves the efficiency of collection and analysis.
Smart Images

Figure CN122054217B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of flow acquisition technology, and in particular to a data acquisition method, apparatus, equipment, medium and computer product. Background Technology
[0002] With the rapid development of 5G networks, traffic detection of the 5G core network's User Plane Function (UPF) is becoming increasingly important as a crucial link in ensuring stable network operation, optimizing resource allocation, and achieving high-quality services. As 5G applications expand deeper into fields such as the Industrial Internet, intelligent transportation, and telemedicine, unprecedentedly stringent requirements are being placed on the real-time performance, accuracy, and efficiency of network traffic detection. Existing 5G core network UPF traffic detection technologies have revealed numerous problems that urgently need to be addressed.
[0003] Currently, most 5G core network UPF traffic detection relies on traditional hardware-based detection methods. These technologies often require the deployment of additional hardware probes or dedicated detection equipment to collect and analyze network traffic. Regarding traffic collection, existing technologies typically employ a full-data collection approach, indiscriminately collecting all traffic from the network. While this method can acquire comprehensive traffic data, in practical applications, the large amount of redundant data not only increases the burden of data storage and transmission but also makes it extremely difficult to extract valuable characteristic traffic from massive amounts of data, failing to achieve efficient collection of valuable characteristic traffic. Summary of the Invention
[0004] The purpose of this invention is to provide a data acquisition method, apparatus, device, medium, and computer product to solve the problem that the existing technology usually adopts a full-volume acquisition method for traffic acquisition, which leads to a large amount of redundant data increasing the system burden, and cannot dynamically allocate acquisition and parsing resources for different traffic.
[0005] To achieve the above objectives, embodiments of the present invention provide a data acquisition method, comprising:
[0006] Acquire the data to be collected and the statistical information of the data to be collected; wherein, the statistical information is the information obtained by statistically classifying the data to be collected according to preset elements;
[0007] The data to be collected is grouped according to the statistical information, and at least one feature stream data and an analysis window corresponding to each feature stream data in the at least one feature stream data are obtained; wherein, each group of data to be collected corresponds to one feature stream data;
[0008] Simultaneously, each of the analysis windows is moved one step length from the first moving time node, and data of each of the feature stream data is collected in the time period corresponding to the analysis window to obtain the collected data; wherein, the first moving time node is the starting time point of the first moving time period among multiple moving time periods; the first step length is obtained based on the at least one feature stream data.
[0009] Optionally, the method, wherein acquiring data for each of the feature stream data within the time period corresponding to the analysis window, includes:
[0010] Determine the temporal characteristics of the feature stream data within the time period corresponding to the analysis window;
[0011] Based on the time characteristics, a data acquisition method is determined, and the characteristic stream data is acquired using the determined data acquisition method to obtain the acquired data.
[0012] Optionally, the method, wherein the time feature of the feature stream data is a continuous feature, determines a data acquisition method based on the time feature, and acquires the feature stream data using the determined data acquisition method to obtain the acquired data, includes:
[0013] Based on the first step length and the feature curve of the feature stream data, a segmentation value is obtained; wherein, the segmentation value is an integer;
[0014] The analysis window is divided into at least one sub-window with a number equal to the division value;
[0015] The feature stream data is collected as the number of frames within the time period corresponding to each sub-window.
[0016] Optionally, the method, wherein the time feature of the feature stream data is a single-point feature, determines a data acquisition method based on the time feature, and acquires the feature stream data using the determined data acquisition method to obtain the acquired data, includes:
[0017] The feature stream data is collected as all data within the time period corresponding to the analysis window.
[0018] Optionally, in the method, after acquiring the data for each of the feature stream data within the time period corresponding to the analysis window, the method further includes:
[0019] Obtain the feature variance of a portion of the feature stream data for each of the aforementioned feature stream data within the time period corresponding to the analysis window;
[0020] Based on the feature variance and the first step length of each feature stream data, the current step length of each feature stream data is obtained;
[0021] The maximum value of the current step size corresponding to each of the at least one feature stream data is taken as the second step size;
[0022] The analysis windows corresponding to the at least one feature stream data are simultaneously moved by the second step size from the second moving time node; wherein, the second moving time node is the starting time point of the second moving time period; the second moving time period is the moving time period that is adjacent to the first moving time period and located after the first moving time period among the plurality of moving time periods.
[0023] Optionally, the method, wherein obtaining the current step size of each feature stream data based on the feature variance and the first step size of each feature stream data includes:
[0024] If the feature variance is greater than a first variance threshold, the first step size is reduced based on the feature variance and the first variance threshold to determine the current step size;
[0025] If the feature variance is less than the second variance threshold, the first step size is increased based on the feature variance and the second variance threshold to determine the current step size; wherein the second variance threshold is less than the first variance threshold.
[0026] If the feature variance is less than or equal to the first variance threshold, and if the feature variance is greater than or equal to the second variance threshold, the first step length is taken as the current step length.
[0027] Optionally, in the method, after acquiring the data for each of the feature stream data within the time period corresponding to the analysis window, the method further includes:
[0028] If the acquisition of the first feature stream data is completed, the acquisition action within the time period corresponding to the analysis window corresponding to the first feature stream data is stopped.
[0029] Optionally, the method further includes, before acquiring the collected data for each of the feature stream data within the time period corresponding to the analysis window:
[0030] If it is determined that there is newly added second feature stream data, the analysis window corresponding to the second feature stream is increased.
[0031] To achieve the above objectives, embodiments of the present invention provide a data acquisition device, comprising:
[0032] The first acquisition module is used to acquire the data to be collected and the statistical information of the data to be collected; wherein, the statistical information is the information obtained by statistically classifying the data to be collected according to preset elements;
[0033] The second acquisition module is used to group the data to be collected according to the statistical information, and acquire at least one feature stream data and an analysis window corresponding to each feature stream data in the at least one feature stream data; wherein, each group of the data to be collected corresponds to one feature stream data;
[0034] The third acquisition module is used to simultaneously move each of the analysis windows by a step length from the first moving time node, and collect data of each of the feature stream data in the time period corresponding to the analysis window, thereby acquiring the collected data; wherein, the first moving time node is the starting time point of the first moving time period among multiple moving time periods; the step length is obtained based on the at least one feature stream data.
[0035] To achieve the above objectives, embodiments of the present invention provide a network device, including: a processor, a memory, and a program or instructions stored in the memory and executable on the processor; wherein, when the processor executes the program or instructions, it implements the data acquisition method described above.
[0036] To achieve the above objectives, embodiments of the present invention provide a readable storage medium having a program or instructions stored thereon, wherein the program or instructions, when executed by a processor, implement the steps in the data acquisition method described above.
[0037] To achieve the above objectives, embodiments of the present invention provide a computer program product, which includes computer instructions that, when executed by a processor, implement the steps of the data acquisition method described above.
[0038] The beneficial effects of the above-described technical solution of the present invention are as follows:
[0039] In this embodiment of the invention, feature stream data is obtained by grouping the data to be collected based on statistical information. Each feature stream data corresponds to an analysis window. After moving multiple analysis windows by one step size simultaneously, data of each feature stream data within its corresponding analysis window is collected to obtain the collected data. The step size of the analysis window is calculated based on the feature stream data. By moving the analysis window with a varying step size according to the feature stream data, data is collected for different feature stream data within the time period corresponding to the analysis window. This avoids the burden on the system caused by a large amount of redundant data due to full data collection and allows for dynamic allocation of collection and parsing resources for different traffic volumes. Attached Figure Description
[0040] Figure 1 This is a schematic diagram of the data acquisition method described in an embodiment of the present invention;
[0041] Figure 2 This is a schematic diagram of the data acquisition system corresponding to the data acquisition method described in the embodiments of the present invention;
[0042] Figure 3 This is a schematic diagram of the data grouping to be collected in the data acquisition method described in this embodiment of the invention;
[0043] Figure 4 This is a schematic diagram of the analysis window of the data acquisition method described in an embodiment of the present invention;
[0044] Figure 5 This is a flowchart of the data acquisition method described in an embodiment of the present invention;
[0045] Figure 6 This is a flowchart illustrating the data acquisition process of the data acquisition method described in this embodiment of the invention.
[0046] Figure 7 This is a schematic diagram of the data acquisition device described in an embodiment of the present invention. Detailed Implementation
[0047] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0048] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0049] In various embodiments of the present invention, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0050] In addition, the terms "system" and "network" are often used interchangeably in this article.
[0051] In the embodiments provided by this invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined based on A. However, it should also be understood that determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0052] For ease of understanding, the following describes some aspects of the embodiments of the present invention:
[0053] like Figure 1 As shown, an embodiment of the present invention provides a data acquisition method, which includes:
[0054] Step S10: Obtain the data to be collected and the statistical information of the data to be collected; wherein, the statistical information is the information obtained after statistically classifying the data to be collected according to preset elements;
[0055] It should be noted that, Figure 2 The data acquisition system corresponding to the data acquisition method described in this embodiment of the invention mainly includes a User Plane Function (UPF) and a traffic collector. In step 201, NetFlow statistics are provided; the UPF interfaces with the NetFlow protocol to provide the traffic collector with traffic-related NetFlow statistics (i.e., the statistics). In step 202, port traffic mirroring is provided; the UPF provides port traffic mirroring (i.e., the data to be collected). The traffic collector includes a sliding window module and a traffic acquisition window. Based on the NetFlow statistics provided by the UPF, the traffic collector, through its own sliding window module, controls multiple traffic acquisition windows to perform more detailed traffic acquisition of the port traffic mirroring (i.e., the data to be collected), achieving a more streamlined and accurate traffic acquisition process. The NetFlow protocol is a network traffic monitoring and analysis technology that uses "flow" as the basic unit to collect and analyze network traffic information. A flow is defined as a set of data packets within a certain time period that all five elements (the source Internet Protocol (IP) address, destination IP address, source port number, destination port number, and protocol type) are identical. By statistically analyzing these flows, various characteristics and behaviors of network traffic can be obtained. In this scheme, it primarily provides statistical information to the traffic collector for calculating relevant control parameters of the sliding window.
[0056] Step S20: Group the data to be collected according to the statistical information, and obtain at least one feature stream data and an analysis window corresponding to each feature stream data in the at least one feature stream data; wherein, each group of the data to be collected corresponds to one feature stream data;
[0057] It should be noted that, as Figure 2As shown, before step 203, the main function of the NetFlow statistics processing submodule in the sliding window module is to extract features from the statistical information provided by UPF, providing a foundation for subsequent sliding window grouping, parameter calculation, and dynamic control. Figure 3 As shown, traffic data can be viewed as a continuously distributed dataset in the time domain. The main features extracted are as follows:
[0058] a. Traffic characteristics: Traffic with the same five-tuple (i.e., source IP address, destination IP address, source port number, destination port number, and protocol type) is considered to have the same characteristics and is identified by the hash result of the five-tuple;
[0059] b. Continuous flow: A characteristic flow that occurs continuously within a certain time period, such as... Figure 3 Feature flow 1, feature flow 2, and feature flow n in the data;
[0060] c. Single-point flow: A characteristic flow that occurs only once at relatively long intervals, such as... Figure 3 Single-point flow in;
[0061] d. Start and end points of a continuous flow: The starting and ending points of a characteristic flow in the time domain, such as... Figure 3 The starting points of feature flow 1 and feature flow n in the time domain;
[0062] e. Continuous Stream Feature Fitting: The packet sizes of discrete data frames in the characteristic stream are fitted to a characteristic curve of the stream using the least squares algorithm, thus reflecting the temporal trend of the stream's flow rate. Figure 3 The characteristic curves corresponding to characteristic flow 1, characteristic flow 2 and characteristic flow n in the figure are respectively.
[0063] The data to be collected is grouped according to the aforementioned flow characteristics, and features are extracted from each group of flow data to classify them into continuous flow and single-point flow. The start and end points of the continuous flow are obtained, and continuous flow feature fitting is performed on the continuous flow to obtain at least one feature flow data. A value is set for each of the feature flow data, such as... Figure 4 The data acquisition window shown is the analysis window. Figure 5As shown, in step 501, NetFlow statistics are received, i.e., the statistics are acquired; in step 502, a hash operation is performed to group the feature flows, with X groups, i.e., the data to be collected is identified by the hash result of a quintuple according to the flow feature (a) of the main features mentioned above; in step 503, connected flows are filtered, and the start and end points of continuous flows with different features are identified, i.e., the data to be collected is processed according to the three features (b), (c), and (d) of the main features mentioned above; in step 504, the continuous flow feature curve is fitted using least squares, i.e., the data to be collected is processed according to feature (e) of the main features mentioned above; in step 505, X collection and analysis windows are opened according to the feature grouping results, similarly, in Figure 2 In step 203, sliding window grouping, that is, grouping the data to be collected according to the statistical information, and obtaining the analysis window corresponding to each feature stream data.
[0064] Step S30: Simultaneously, each of the analysis windows is moved one step length from the first moving time node, and data of each of the feature stream data is collected in the time period corresponding to the analysis window to obtain the collected data; wherein, the first moving time node is the starting time point of the first moving time period among multiple moving time periods; the first step length is obtained based on the at least one feature stream data.
[0065] It should be noted that the size, starting point, speed, and sliding direction of the analysis window are consistent for multiple feature stream data. In, for example... Figure 4 In the time domain, the analysis window moves a step size corresponding to the starting time of the first moving segment in multiple moving time periods to the starting time of the next moving time period. The analysis window obtains a unified step size based on the feature stream data within the time periods corresponding to all the analysis windows, so that the analysis window uses this step size in the next moving process. After the analysis window moves, it collects data on the feature stream data within the time periods corresponding to the window, that is, it moves each analysis window from the first moving time node simultaneously, and collects the data of each feature stream data within the time period corresponding to the analysis window based on the first step size obtained from at least one feature stream data, thereby obtaining the collected data. The working principle of the variable sliding window acquisition algorithm is as follows: Figure 4As shown, the acquisition window slides forward at a certain speed over the time domain with a certain length, and traffic data is acquired within the sliding window range. During the forward movement, the forward movement speed of the sliding window (i.e., the number of data frames the window moves forward) is controlled by combining the distribution characteristics of different feature flows in the time domain (i.e., the fitted curve in the main feature e) with a certain algorithm (i.e., the curve variance within the window). Each forward movement samples a set number of feature flow data according to different acquisition methods, thereby realizing different acquisition control for different feature flow data.
[0066] In this embodiment, feature stream data is obtained by grouping the data to be collected based on statistical information. Each feature stream data corresponds to an analysis window. After moving multiple analysis windows by one step size simultaneously, the data of each feature stream data within its corresponding analysis window is collected, resulting in collected data. The step size of the analysis window is calculated based on the feature stream data. By moving the analysis window with a varying step size according to the feature stream data, data is collected for different feature stream data within the time period corresponding to the analysis window. This avoids the burden on the system caused by a large amount of redundant data due to full data collection and allows for dynamic allocation of collection and parsing resources for different traffic volumes.
[0067] Optionally, the method, wherein step S30 includes:
[0068] Determine the temporal characteristics of the feature stream data within the time period corresponding to the analysis window;
[0069] Based on the time characteristics, a data acquisition method is determined, and the characteristic stream data is acquired using the determined data acquisition method to obtain the acquired data.
[0070] In this embodiment, the temporal characteristic of the feature stream data is either a single-point stream or a continuous stream. Different data acquisition methods are set according to the temporal characteristic of the feature stream data, and data is acquired from the feature stream data within a time period corresponding to at least one analysis window of the feature stream data. For example... Figure 2 As shown, in step 206, the sliding window module acquires and controls the flow acquisition window, acquiring the corresponding data frames of the corresponding feature flow data through the feature flow 1 acquisition window, the feature flow 2 acquisition window, and the feature flow n acquisition window, respectively. For example... Figure 6 As shown, in step 602, it is determined whether the current feature acquisition window is a continuous flow. Different acquisition methods are used in steps 603 and 606 respectively. Discrete acquisition within the window is used for continuous flow, while full acquisition within the window is used for single-point flow. Simultaneously, a sliding window algorithm is used to control the frequency of discrete acquisition. This effectively reduces redundant data in large flows, improves the loss of small-scale feature flows, and enhances the effectiveness of flow acquisition and analysis.
[0071] Optionally, the method, wherein the time feature of the feature stream data is a continuous feature, determines a data acquisition method based on the time feature, and acquires the feature stream data using the determined data acquisition method to obtain the acquired data, includes:
[0072] Based on the first step length and the feature curve of the feature stream data, a segmentation value is obtained; wherein, the segmentation value is an integer;
[0073] The analysis window is divided into at least one sub-window with a number equal to the division value;
[0074] The feature stream data is collected as the number of frames within the time period corresponding to each sub-window.
[0075] In this embodiment, such as Figure 6 As shown, if the result of determining whether the current feature acquisition window is a continuous stream in step 602 is yes, that is, the time feature of the feature stream data is a continuous feature, then step 603 is performed to calculate the current window's own step size Ln based on the feature curve. Then, in step 604, the sliding window is divided into equal sub-windows of Lmax / Ln, where Lmax is the step size of the analysis window's movement in the next moving time period (in step 601, the current maximum step size Lmax is obtained). That is, based on the first step size and the feature curve of the feature stream data, a segmentation value is obtained, and the analysis window is divided into at least one sub-window according to the segmentation value. Note that Lmax / Ln may be a decimal and needs to be rounded up to the nearest integer as the segmentation value. In step 605, one frame of successfully matched feature data is taken from each sub-window, that is, one frame of data from the feature stream data within the time period corresponding to each sub-window is collected as the collected data. The preset number of frames can be one frame or can be adjusted according to requirements.
[0076] Optionally, the method, wherein the time feature of the feature stream data is a single-point feature, determines a data acquisition method based on the time feature, and acquires the feature stream data using the determined data acquisition method to obtain the acquired data, includes:
[0077] The feature stream data is collected as all data within the time period corresponding to the analysis window.
[0078] In this embodiment, such as Figure 6As shown, if the result of determining whether the current feature acquisition window is a continuous stream in step 602 is no, that is, the time feature of the feature stream data is a single-point feature, then step 606 is performed to acquire all data frames in the sliding window where all features are successfully matched, that is, to acquire all the data of the feature stream data within the time period corresponding to the analysis window as the acquired data.
[0079] Optionally, the method, after step S30, further includes:
[0080] Obtain the feature variance of a portion of the feature stream data for each of the aforementioned feature stream data within the time period corresponding to the analysis window;
[0081] Based on the feature variance and the first step length of each feature stream data, the current step length of each feature stream data is obtained;
[0082] The maximum value of the current step size corresponding to each of the at least one feature stream data is taken as the second step size;
[0083] The analysis windows corresponding to the at least one feature stream data are simultaneously moved by the second step size from the second moving time node; wherein, the second moving time node is the starting time point of the second moving time period; the second moving time period is the moving time period that is adjacent to the first moving time period and located after the first moving time period among the plurality of moving time periods.
[0084] In this embodiment, such as Figure 5 As shown, in step 506, the window baseline step size L (i.e., the number of data frames the window moves forward at one time), the lower variance threshold T1 (i.e., the minimum allowable variance of the characteristic flow within the window), and the upper variance threshold T2 (i.e., the maximum allowable variance of the characteristic flow within the window) are determined; similarly, in Figure 2 In step 204, the sliding window parameters are calculated, including: the window baseline step size L (i.e., the number of data frames the window moves forward at one time), the lower variance threshold T1 (i.e., the minimum allowable variance of the characteristic flow within the window), and the upper variance threshold T2 (i.e., the maximum allowable variance of the characteristic flow within the window). In step 507, the maximum window migration step size Lmax is calculated, that is, in step S30, each of the analysis windows is moved by the first step size simultaneously. Similarly, in Figure 2In step 205, the window is dynamically controlled, and each analysis window moves simultaneously according to the distance of the first step length. After step S30, the step length of the analysis window is recalculated, that is, the second step length is calculated. In step 509, each acquisition window acquires data based on the step length Ln (i.e., the current step length) calculated according to its own characteristic curve. That is, for feature stream n, the variance Vn (i.e., the feature variance) within the time period corresponding to the analysis window is acquired. Based on the feature variance Vn and the first step length Lmax, the current step length Ln corresponding to feature stream n is calculated. The initial value of this Ln is uniformly set as the window reference step length L. In step 510, it is determined whether all data acquisition has been completed. In step 511, if it is not completed, a new maximum step length Lmax = max(L1, L2, ..., Ln) is obtained. In step 512, Lmax is adjusted according to the current maximum step length. That is, the maximum value of the current step length Ln corresponding to at least one feature stream data is selected as the second step length Lmax.
[0085] Optionally, the method, wherein obtaining the current step size of each feature stream data based on the feature variance and the first step size of each feature stream data includes:
[0086] If the feature variance is greater than a first variance threshold, the first step size is reduced based on the feature variance and the first variance threshold to determine the current step size;
[0087] If the feature variance is less than the second variance threshold, the first step size is increased based on the feature variance and the second variance threshold to determine the current step size; wherein the second variance threshold is less than the first variance threshold.
[0088] If the feature variance is less than or equal to the first variance threshold, and if the feature variance is greater than or equal to the second variance threshold, the first step length is taken as the current step length.
[0089] In this embodiment, the process of adjusting the first step length Lmax according to the feature variance Vn to obtain the current step length Ln is as follows:
[0090] When the variance (i.e., the feature variance Vn) is greater than the upper variance threshold T2 (i.e., the first variance threshold), the data exhibits large fluctuations, requiring an increase in the sampling rate and a decrease in the window step size. Therefore, the following adjustments are made:
[0091] Ln = Lmax - k*(Vn-T2) (where Lmax is the length of the first step), and k is a constant coefficient greater than 0;
[0092] When the variance (i.e., the feature variance Vn) is less than the lower variance threshold T1 (i.e., the second variance threshold), the data fluctuation is small, and the sampling rate needs to be reduced and the window step size increased. Therefore, the following adjustments are made:
[0093] Ln = Lmax + k*(T1-Vn)(In this formula, Lmax is the length of the first step, and k is a constant coefficient greater than 0;
[0094] When the variance of a certain feature stream data is between the first variance threshold and the second variance threshold, the first step length is taken as the current step length corresponding to that feature stream data. Wherein, the second variance threshold T1 is less than the first variance threshold T2.
[0095] Optionally, in the method, after acquiring the data for each of the feature stream data within the time period corresponding to the analysis window, the method further includes:
[0096] If the acquisition of the first feature stream data is completed, the acquisition action within the time period corresponding to the analysis window corresponding to the first feature stream data is stopped.
[0097] In this embodiment, the endpoint of the first feature stream data exists within the time period corresponding to the analysis window corresponding to the first feature stream data. Based on this, it is determined that the acquisition of the first feature stream data is completed. Therefore, the acquisition action within the time period corresponding to the analysis window corresponding to the first feature stream data is stopped, and the corresponding analysis window is closed.
[0098] Optionally, the method further includes, before acquiring the collected data for each of the feature stream data within the time period corresponding to the analysis window:
[0099] If it is determined that there is newly added second feature stream data, the analysis window corresponding to the second feature stream is increased.
[0100] In this embodiment, in step 508, the hash operation is performed to obtain the features of all data within the window. That is, if it is determined that there is newly added second feature stream data, the analysis window corresponding to the second feature stream is increased.
[0101] It should be noted that the overall process of the data acquisition method described in this invention is as follows: The traffic collector receives statistical information from the NetFlow protocol (i.e., the statistical information), performs a hash operation using the 5-tuple of the data frame, groups the traffic by characteristics, and determines the number of groups X. Based on the continuity of the characteristic traffic distribution, the characteristic traffic is divided into continuous flow and single-point flow. For continuous flow, the packet size of each discrete data frame is taken, and a least squares algorithm is used for fitting to form the characteristic curve of the characteristic flow. X acquisition windows (i.e., the analysis windows) are created for acquisition. The initial baseline step size L, the lower variance threshold T1 (i.e., the second variance threshold), and the upper variance threshold T2 (i.e., the first variance threshold) are set for the window. The sliding window is moved forward by Lmax steps (i.e., the first step size). The features of newly added data frames in the window are calculated by hash operation. Traffic is acquired in each characteristic window. It is determined whether the characteristic flow data acquisition is complete. If it is complete, the corresponding acquisition window is closed. If it is not complete, the step size Ln (i.e., the current step size) of each characteristic flow in the latest sliding window is calculated according to the above formula. The maximum step size is adjusted to the new maximum value Lmax of Ln (i.e., the second step size), the window is moved forward by Lmax (i.e., the second step size), and the above steps are repeated.
[0102] This invention addresses the 5G core network UPF traffic acquisition scenario and designs a 5G core network UPF traffic acquisition method based on a variable sliding window. The data acquisition system corresponding to this method includes NetFlow statistics, a traffic collector, and a sliding window module. The NetFlow statistics are provided by the UPF to the traffic collector, serving as the data basis for the collector's sliding window algorithm acquisition and dynamic control. The NetFlow statistics processing module within the sliding window divides the traffic into continuous flows and single-point flows with certain characteristics in the time domain. Then, using a least-squares fitting algorithm, it fits the characteristic curve of the continuous flow to characterize the trend of its traffic magnitude. The sliding window module divides the feature window based on the processing results of the information processing module. The window slides to acquire data with a certain step size. During data acquisition, the sliding step size is dynamically adjusted based on the characteristic curves of each feature flow, thereby achieving differentiated acquisition of each feature flow.
[0103] The traffic feature segmentation method proposed in this invention can effectively characterize the traffic trends of different feature flows. It dynamically adjusts the sliding window step size based on the characteristic trends of different flows. The variance of the feature curve within the window is used to characterize traffic fluctuations. Large fluctuations can be mitigated by decreasing the step size to increase the sampling rate, while small fluctuations can be mitigated by increasing the step size to decrease the sampling rate. This effectively reduces redundant data and lowers the processing pressure on the device in high-traffic scenarios. Simultaneously, it can effectively identify single-point flows and improve the handling of small-scale feature flows.
[0104] like Figure 7 As shown, to achieve the above objectives, embodiments of the present invention provide a data acquisition device, comprising:
[0105] The first acquisition module 701 is used to acquire the data to be collected and the statistical information of the data to be collected; wherein, the statistical information is the information obtained by statistically classifying the data to be collected according to preset elements;
[0106] The second acquisition module 702 is used to group the data to be collected according to the statistical information, and acquire at least one feature stream data and an analysis window corresponding to each feature stream data in the at least one feature stream data; wherein, each group of the data to be collected corresponds to one feature stream data;
[0107] The third acquisition module 703 is used to simultaneously move each of the analysis windows by a step length from the first moving time node, and collect data of each of the feature stream data in the time period corresponding to the analysis window, thereby acquiring the collected data; wherein, the first moving time node is the starting time point of the first moving time period among multiple moving time periods; the step length is obtained based on the at least one feature stream data.
[0108] Optionally, in the aforementioned apparatus, the third acquisition module 703 includes:
[0109] The first determining unit is used to determine the temporal characteristics of the feature stream data within the time period corresponding to the analysis window;
[0110] The first acquisition unit is used to determine the data acquisition method based on the time characteristics, and to acquire the feature stream data using the determined data acquisition method to acquire the acquired data.
[0111] Optionally, in the aforementioned apparatus, the temporal feature of the feature stream data is a continuous feature, and the first acquisition unit includes:
[0112] The first acquisition component is used to acquire a segmentation value based on the first step length and the feature curve of the feature stream data; wherein the segmentation value is an integer;
[0113] A first processing component is configured to divide the analysis window into at least one sub-window with a number equal to the division value;
[0114] The second processing component is used to collect the feature stream data for a preset number of frames within the time period corresponding to each sub-window as the collected data.
[0115] Optionally, in the aforementioned apparatus, the temporal feature of the feature stream data is a single-point feature, and the first acquisition unit includes:
[0116] The third processing component is used to collect all data of the feature stream data within the time period corresponding to the analysis window as the collected data.
[0117] Optionally, the device further includes:
[0118] The fourth acquisition module is used to acquire the feature variance of a portion of the feature stream data for each feature stream data in the time period corresponding to the analysis window;
[0119] The fifth acquisition module is used to acquire the current step size of each feature stream data according to the feature variance and the first step size of each feature stream data.
[0120] The first processing module is used to take the maximum value of the current step size corresponding to the at least one feature stream data as the second step size;
[0121] The second processing module is used to simultaneously move the analysis windows corresponding to the at least one feature stream data by the second step size starting from the second moving time node; wherein, the second moving time node is the starting time point of the second moving time period; the second moving time period is the moving time period that is adjacent to the first moving time period and located after the first moving time period among the plurality of moving time periods.
[0122] Optionally, in the aforementioned apparatus, the fifth acquisition module includes:
[0123] The second determining unit is configured to determine the current step size by reducing the first step size based on the feature variance and the first variance threshold when the feature variance is greater than the first variance threshold.
[0124] The third determining unit is configured to determine the current step size by increasing the first step size based on the feature variance and the second variance threshold when the feature variance is less than the second variance threshold; wherein the second variance threshold is less than the first variance threshold.
[0125] The first processing unit is configured to use the first step size as the current step size when the feature variance is less than or equal to the first variance threshold and when the feature variance is greater than or equal to the second variance threshold.
[0126] Optionally, the device further includes:
[0127] The third processing module is used to stop the acquisition action within the time period corresponding to the analysis window corresponding to the first feature stream data when it is determined that the acquisition of the first feature stream data has been completed.
[0128] Optionally, the device further includes:
[0129] The fourth processing module is used to increase the analysis window corresponding to the second feature stream when it is determined that there is newly added second feature stream data.
[0130] It should be noted that the apparatus provided in this embodiment of the invention can implement all the method steps implemented in the above method embodiment and can achieve the same technical effect. Therefore, the parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0131] To achieve the above objectives, embodiments of the present invention provide a network device, including: a processor, a memory, and a program or instructions stored in the memory and executable on the processor; wherein, when the processor executes the program or instructions, it implements the data acquisition method described above.
[0132] To achieve the above objectives, embodiments of the present invention provide a readable storage medium having a program or instructions stored thereon, wherein the program or instructions, when executed by a processor, implement the steps in the data acquisition method described above.
[0133] To achieve the above objectives, embodiments of the present invention provide a computer program product, which includes computer instructions that, when executed by a processor, implement the steps of the data acquisition method described above.
[0134] It should be further noted that the terminals described in this specification include, but are not limited to, smartphones, tablets, etc., and many of the functional components described are referred to as modules in order to emphasize the independence of their implementation.
[0135] In this embodiment of the invention, the module can be implemented in software so that it can be executed by various types of processors. For example, an identified executable code module may include one or more physical or logical blocks of computer instructions, which may be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but may include different instructions stored in different locations, which, when logically combined, constitute the module and achieve the module's intended purpose.
[0136] In practice, an executable code module can be a single instruction or many instructions, and can even be distributed across multiple different code segments, different programs, and across multiple memory devices. Similarly, operational data can be identified within the module and can be implemented in any suitable form and organized within any suitable data structure. This operational data can be collected as a single dataset or distributed across different locations (including different storage devices), and can exist, at least in part, solely as electronic signals within the system or network.
[0137] When a module can be implemented using software, considering the current level of hardware technology, modules that can be implemented in software can be implemented using hardware circuits by those skilled in the art to achieve the corresponding functions, without considering cost. These hardware circuits include conventional very-large-scale integrated circuits (VLSI) or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules can also be implemented using programmable hardware devices, such as field-programmable gate arrays, programmable array logic, and programmable logic devices.
[0138] The exemplary embodiments described above are with reference to the accompanying drawings. Many different forms and embodiments are feasible without departing from the spirit and teachings of the invention. Therefore, the invention should not be construed as limiting the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided to make the invention complete and convey the scope of the invention to those skilled in the art. In these drawings, component dimensions and relative dimensions may be exaggerated for clarity. The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. As used herein, unless clearly indicated otherwise, the singular forms “a,” “an,” and “the” are intended to include all such forms. It will be further understood that the terms “comprising” and / or “including”, when used in this specification, indicate the presence of the stated features, integers, steps, operations, components, and / or elements, but do not exclude the presence or addition of one or more other features, integers, steps, operations, components, and / or groups thereof. Unless otherwise indicated, when stated, a range of values includes the upper and lower limits of the range and any subranges in between.
[0139] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A data acquisition method, characterized by, include: Acquire the data to be collected and the statistical information of the data to be collected; wherein, the statistical information is the information obtained by statistically classifying the data to be collected according to preset elements; The data to be collected is grouped according to the statistical information, and at least one feature stream data and an analysis window corresponding to each feature stream data in the at least one feature stream data are obtained; wherein, each group of data to be collected corresponds to one feature stream data; Simultaneously, each of the aforementioned analysis windows is moved one step length from the first moving time node, and data of each of the aforementioned feature stream data is collected within the time period corresponding to the analysis window to obtain the collected data; wherein, the first moving time node is the starting time point of the first moving time period among multiple moving time periods; the first step length is obtained based on the at least one feature stream data; After acquiring the data for each of the aforementioned feature stream data within the time period corresponding to the analysis window, the method further includes: Obtain the feature variance of a portion of the feature stream data for each of the aforementioned feature stream data within the time period corresponding to the analysis window; Based on the feature variance and the first step length of each feature stream data, the current step length of each feature stream data is obtained; The maximum value of the current step size corresponding to each of the at least one feature stream data is taken as the second step size; The analysis windows corresponding to the at least one feature stream data are simultaneously moved by the second step size from the second moving time node; wherein, the second moving time node is the starting time point of the second moving time period; the second moving time period is the moving time period that is adjacent to the first moving time period and located after the first moving time period among the plurality of moving time periods.
2. The method of claim 1, wherein, Collect data for each of the aforementioned feature stream data within the time period corresponding to the analysis window, and obtain the collected data, including: Determine the temporal characteristics of the feature stream data within the time period corresponding to the analysis window; Based on the time characteristics, a data acquisition method is determined, and the characteristic stream data is acquired using the determined data acquisition method to obtain the acquired data.
3. The method according to claim 2, characterized in that, The time feature of the feature stream data is a continuous feature. A data acquisition method is determined based on this time feature, and the feature stream data is acquired using the determined data acquisition method to obtain the acquired data, including: Based on the first step length and the feature curve of the feature stream data, a segmentation value is obtained; wherein, the segmentation value is an integer; The analysis window is divided into at least one sub-window with a number equal to the division value; The feature stream data is collected as the number of frames within the time period corresponding to each sub-window.
4. The method according to claim 2, characterized in that, The time feature of the feature stream data is a single-point feature. A data acquisition method is determined based on the time feature, and the feature stream data is acquired using the determined data acquisition method to obtain the acquired data, including: The feature stream data is collected as all data within the time period corresponding to the analysis window.
5. The method according to claim 1, characterized in that, Based on the feature variance and the first step length of each feature stream data, the current step length of each feature stream data is obtained, including: If the feature variance is greater than a first variance threshold, the first step size is reduced based on the feature variance and the first variance threshold to determine the current step size; If the feature variance is less than the second variance threshold, the first step size is increased based on the feature variance and the second variance threshold to determine the current step size; wherein the second variance threshold is less than the first variance threshold. If the feature variance is less than or equal to the first variance threshold, and if the feature variance is greater than or equal to the second variance threshold, the first step length is taken as the current step length.
6. The method according to claim 1, characterized in that, After acquiring the data for each of the aforementioned feature stream data within the time period corresponding to the analysis window, the method further includes: If the acquisition of the first feature stream data is completed, the acquisition action within the time period corresponding to the analysis window corresponding to the first feature stream data is stopped.
7. The method according to claim 1, characterized in that, Before acquiring the collected data for each of the aforementioned feature stream data within the time period corresponding to the analysis window, the method further includes: If it is determined that there is newly added second feature stream data, the analysis window corresponding to the second feature stream is increased.
8. A data acquisition device, characterized in that, include: The first acquisition module is used to acquire the data to be collected and the statistical information of the data to be collected; wherein, the statistical information is the information obtained by statistically classifying the data to be collected according to preset elements; The second acquisition module is used to group the data to be collected according to the statistical information, and acquire at least one feature stream data and an analysis window corresponding to each feature stream data in the at least one feature stream data; wherein, each group of the data to be collected corresponds to one feature stream data; The third acquisition module is used to simultaneously move each of the analysis windows by a step length from the first moving time node, and collect data of each of the feature stream data in the time period corresponding to the analysis window, thereby acquiring the collected data; wherein, the first moving time node is the starting time point of the first moving time period among multiple moving time periods; the step length is obtained based on the at least one feature stream data; The device further includes: The fourth acquisition module is used to acquire the feature variance of a portion of the feature stream data for each feature stream data in the time period corresponding to the analysis window; The fifth acquisition module is used to acquire the current step size of each feature stream data according to the feature variance and the first step size of each feature stream data. The first processing module is used to take the maximum value of the current step size corresponding to the at least one feature stream data as the second step size; The second processing module is used to simultaneously move the analysis windows corresponding to the at least one feature stream data by the second step size starting from the second moving time node; wherein, the second moving time node is the starting time point of the second moving time period; the second moving time period is the moving time period that is adjacent to the first moving time period and located after the first moving time period among the plurality of moving time periods.
9. A network device, comprising: A processor, a memory, and a program or instructions stored in the memory and executable on the processor; characterized in that, when the processor executes the program or instructions, it implements the data acquisition method as described in any one of claims 1-7.
10. A readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instructions are executed by the processor, they implement the steps of the data acquisition method as described in any one of claims 1-7.
11. A computer program product, characterized in that, It includes computer instructions, which, when executed by a processor, implement the steps of the data acquisition method as described in any one of claims 1-7.
Citation Information
Patent Citations
CN105376105A
CN117768366A