Real-time DNS gateway poor-quality delimiting method based on waveform fitting and message queue
Through the method based on waveform fitting and message queue, the Kafka queue is used to pass DNS monitoring delay data, and network element waveform fitting is used and searched to generate quality difference events, which solves the problem of capturing network element waveforms from global perspective in DNS delay analysis, and improves delimiting efficiency and real-time monitoring capabilities.
Patent Information
- Application Number
- CN202510370132.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to efficiently capture the connection between network element waveforms from a global perspective, resulting in the flexibility and inefficiency of DNS delay analysis, especially in the analysis of large-scale network elements.
Using a method based on waveform fitting and message queue, DNS monitoring delay data is passed through Kafka message queue, network element waveform fitting is used and checked to generate quality difference events and perform data synchronization, eliminate the impact of superior network element degradation, and realize real-time quality difference delimitation.
It improves the efficiency of delimiting the quality of DNS gateways, and can capture the connection between network element waveforms from a global perspective, reduces the overhead of repeated operation, and realizes real-time monitoring and analysis.
Smart Images

Figure CN120281744A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of DNS gateway monitoring delay and quality difference delimitation, and specifically to a real-time DNS gateway quality difference delimitation method based on waveform fitting and message queue. Background Art
[0002] The DNS (Domain Name System) gateway monitoring delay refers to the time interval measured from when a DNS query request is sent by a client to when a response is received. This time interval reflects the speed at which the DNS server processes requests and the network transmission speed, and is one of the important indicators for evaluating DNS performance and reliability.
[0003] The DNS gateway monitoring delay consists of the following four parts:
[0004] ① Query sending time: The time required for the client to send a DNS request packet to the DNS gateway or recursive resolver.
[0005] ② DNS gateway processing time: The time when the DNS gateway receives the request, parses it, forwards it to the authoritative DNS server (if necessary), and waits for the return result.
[0006] ③ Response receiving time: The time required for the DNS gateway to return the resolved IP address response to the client.
[0007] ④ Network transmission delay: It includes the time when all involved data packets are transmitted on the network during the process from ① to ③, and is affected by factors such as physical distance, routing selection, and bandwidth.
[0008] Monitoring DNS delay is necessary. A low DNS resolution delay can speed up web page loading and improve the user's speed experience of accessing Internet resources. The fluctuation of DNS delay can reflect potential faults of network elements and evaluate the security and stability of the system. By continuously monitoring DNS delay, it can help administrators quickly discover potential problems, such as DNS server overload, unstable network connection, etc. In addition, an abnormally high delay may indicate the existence of DDoS attacks or other malicious activities, and timely monitoring helps to ensure the stable operation of the system. Means such as active probing or passive listening can be used to obtain DNS delay data. For example, use command-line tools such as dig and nslookup or dedicated monitoring software to send test requests to a specified DNS server regularly and record the response time. Capture actual DNS requests and responses in network traffic to obtain their delay situations.
[0009] However, the commonly used DNS analysis methods are very simple. We can only analyze the overall delay of a single network element from a statistical perspective, using methods such as mean, variance or percentile, and it is difficult to establish the connection between network elements of different levels, such as Bras, OLT, and PON port. If you want to analyze the DNS delay of network elements in a certain period of time, you can only set a sliding time window of fixed size, calculate the statistical distribution such as the average value of the observations in the window by moving average, or use exponential weighted moving average to distinguish the importance of observations at different times, and then analyze the statistical distribution of a specific time period. Such a method obviously lacks flexibility, and the analysis results are all based on a sliding time window of fixed size. If you want to obtain analysis results for different time periods, you must set multiple sliding time windows and make judgments in sequence, but the efficiency is low. If the scope of analysis is expanded to a larger range of network elements, it will incur huge time overhead.
[0010] Therefore, how to capture the connection between network element waveforms from a global perspective and improve the demarcation efficiency is a technical problem that needs to be solved urgently. Summary of the invention
[0011] The technical task of the present invention is to provide a real-time DNS gateway quality difference delimitation method based on waveform fitting and message queue to solve the problem of how to capture the connection between network element waveforms from a global perspective and improve the delimitation efficiency.
[0012] The technical task of the present invention is achieved in the following way: a real-time DNS gateway quality difference delimitation method based on waveform fitting and message queue, the method is as follows:
[0013] Data collection: Gather DNS monitoring latency data from the underlying data using SQL;
[0014] Obtain latency data: put the DNS latency monitoring data into the Kafka message queue, and the data reading program consumes the Kafka message to obtain the DNS latency data of the network element;
[0015] Data preprocessing: fill in missing values in the DNS delay data of the network element and convert the data format of the DNS delay data of the network element;
[0016] Waveform fitting: Perform waveform fitting based on the union-find method on the DNS delay data of the network elements to obtain a list of network elements with similar waveforms;
[0017] Generate poor quality events: determine whether network elements with similar waveforms are degraded, and determine whether the degraded network elements are degraded together, eliminate the poor quality of the upper network elements, and generate poor quality events;
[0018] Data synchronization: synchronize the data of the network elements involved in the poor quality event;
[0019] Poor quality input: Input poor quality events into the result table and start the poor quality demarcation for the next batch of network elements.
[0020] Preferably, data collection is to summarize data from multiple data sources through SQL statements and insert the result data into the input table, as follows:
[0021] Clear the data corresponding to the date in the input table;
[0022] Generate timestamp data every 1 minute;
[0023] Group the underlying network element data tables by network element name and filter out the network element data corresponding to the date; among them, when filtering, select the data record with the largest number of attached users as the network element data of the corresponding network element; among them, the network element data includes geographical location information and network performance index data;
[0024] Connect the timestamp data and the network element data, calculate the average value of the network performance index to fill in the missing network performance index data, and insert the processed result data into the input table.
[0025] Preferably, obtain the delay data as follows:
[0026] Put the DNS delay data summarized every minute into the Kafka message queue, and each network element has a unique Topic message queue;
[0027] Read the data in the kafka message queue through a data reading program based on the python language;
[0028] DNS delay demarcation: The data reading program continuously consumes the message queue of the Topic corresponding to each network element; each time it consumes, the data reading program obtains the DNS delay data of the previous n time periods; among them, n is a variable time period window, and n is an integer greater than or equal to 3.
[0029] More preferably, the input of the data reading program is the kafka message queue and the output is the DNS delay data; as follows:
[0030] After installing the kafka-python third-party library, instantiate the KafkaConsumer class in the data reading program;
[0031] During the instantiation process, set the IP address of the Kafka broker and the Topic name parameter;
[0032] Consume the Kafka message queue through the KafkaConsumer class to return the corresponding data.
[0033] Preferably, the data reading program obtains the DNS delay data for the previous n time periods as follows:
[0034] When instantiating the KafkaConsumer in the data reading program, set auto.offset.reset to earliest;
[0035] Subtract n time periods from the current timestamp to obtain the timestamp of the start time; since Kafka uses millisecond-level timestamps, both the current timestamp and n time periods are converted to milliseconds during the calculation;
[0036] Calculate the offset corresponding to the start time, move the consumer to the offset position, start consuming the message queue, and obtain the DNS delay data.
[0037] As a preference, the data preprocessing is as follows:
[0038] Create a time code table containing n time periods;
[0039] Perform a left join on the time code table and the network element data;
[0040] For the fields of time series and geographical location information, fill them with the same values as in the context, and for the network performance metrics, fill them with 0; at the same time, adopt a cautious strategy that when no data is collected from the device, it is considered that there is no quality problem with the network element;
[0041] After filling the missing data, convert the data format of each field, and convert them to int64 type, str type, and float64 type respectively according to the field type.
[0042] As a preference, the waveform fitting is as follows:
[0043] Obtain the out_device containing the network element names from the input device_dict, combine all the network element names in pairs, and the function combinations() can return non-repeated combinations out_pairs;
[0044] When calculating the similarity, both the Pearson similarity and the cosine similarity are used: if the delay data of any network element is the same value within the time period, the cosine similarity is used for waveform fitting; otherwise, the Pearson similarity is used for waveform fitting;
[0045] After all the network element combinations in out_pairs are calculated, return the fitted network element list simila_rdevi.
[0046] As a preference, generating quality problem events is as follows:
[0047] Eliminate the deterioration of subordinate network elements caused by faults in superior network elements;
[0048] Judge whether each network element in the network element list deteriorates: Use the 3σ rule to judge deterioration, that is, calculate a deterioration threshold through threshold = mean + σ × std: If the network element is higher than the threshold at any time point, it is considered to have deteriorated; Traverse all network elements to judge the deterioration of each network element within the time period; where, threshold represents the quality difference threshold; mean represents the average DNS detection delay within the statistical time period; std represents the standard deviation of DNS detection delay within the statistical time period; σ takes a value of 2;
[0049] According to the deterioration conditions of all network elements within the time period, judge whether there are multiple network elements deteriorating simultaneously within the time period:
[0050] If there is simultaneous deterioration, it means a quality difference event is formed.
[0051] More preferably, the judgment of the quality difference event category is as follows:
[0052] If the number of network elements involved in the quality difference event is the same as the number of network elements under the superior network element, it means that the deterioration of the subordinate network elements is caused by the fault of the superior network element, and record the event category according to the different network element levels;
[0053] If the number of network elements involved in the quality difference event is less than the number of network elements under the superior network element and greater than 1, it means that multiple subordinate network elements deteriorate simultaneously, and record the event category;
[0054] After judging the quality difference event category, record the start date, start time, end time, and network element name of the deterioration event at the same time for making the top-level conclusion.
[0055] Preferably, data synchronization is to synchronize the input of the network elements involved in each quality difference event to another data table, and make a bar chart by querying the data table;
[0056] Among them, the table structure of the data table is exactly the same as the fields of the data in the Kafka message queue;
[0057] One quality difference event corresponds to one bar chart, and multiple network elements in the quality difference event correspond to several sub-charts in the bar chart, that is, each sub-chart is a bar chart of the DNS monitoring experiment of a network element; Multiple network elements within one quality difference event produce a fit on the waveform, and several sub-charts within one bar chart present similar graphs.
[0058] The real-time DNS gateway quality difference delimitation method based on waveform fitting and message queue of the present invention has the following advantages:
[0059] (1) The present invention uses a Kafka message queue to transmit DNS monitoring delay data of network elements at all levels. After the DNS delay data of a batch of network elements is summarized, the delimiter program consumes the Kafka message queue. When performing waveform fitting, all network elements are paired pairwise to generate two sets of network elements. The Pearson correlation coefficient is calculated for each set of network elements, and the sets with a Pearson correlation coefficient less than the threshold are discarded. The remaining sets are merged using the union-find set. If the intersection of two sets is not empty, they are merged until there are no sets that can be merged. The sets of fitted network elements are judged for common degradation, and the time periods when all network elements simultaneously exhibit quality degradation are calculated. When judging common degradation, the degradation events of superior network elements need to be excluded. Finally, quality degradation events are generated and stored in the database, and then the Kafka message queue is continuously consumed until the delimitation of degradation for all network elements is completed, achieving the capture of the connection between network element waveforms from a global perspective and improving the delimitation efficiency.
[0060] (2) The present invention uses a Kafka message queue to perform real-time transmission of DNS monitoring delay data and uses waveform fitting based on the union-find set, which can obtain different sets. The network elements within each set exhibit relatively similar waveforms. Compared with pairwise comparison of network elements in sequence for clustering, it can capture the connection between network element waveforms from a global perspective.
[0061] (3) The present invention first pairs all network elements pairwise. Suppose there are n network elements, and at this time, sets will be generated, and each set contains two different network elements. The Pearson correlation coefficient is calculated for each set of network elements, and the sets with a Pearson correlation coefficient less than the threshold are discarded. The remaining sets are merged using the union-find set. If the intersection of two sets is not empty, they are merged until there are no sets that can be merged. Using waveform fitting based on the union-find set can obtain different sets, and the network elements within each set exhibit relatively similar waveforms. Compared with pairwise comparison of network elements in sequence for clustering, it can capture the connection between network element waveforms from a global perspective.
[0062] (4) Traditional anomaly recognition or quality degradation delimitation often waits until all network element data is available before starting the delimitation of quality degradation. For example, for data of Bras, OLT, or PON ports in a first-level area, the time overhead of SQL aggregation and writing data to the database and the program reading data from the database accounts for the largest proportion of the time-consuming in the entire process and is a repetitive operation. However, the present invention can perform real-time quality degradation delimitation through the Kfaka message queue. As long as the delay data aggregation of a batch of network elements is completed, the data is put into the Kafka message queue, and the data reading program can consume and obtain the data for delimitation, repeating in a cycle, which can greatly improve the efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The present invention will be further described below in conjunction with the accompanying drawings.
[0064] Attached Figure 1 is a flowchart of a real-time DNS gateway quality difference delimitation method based on waveform fitting and message queue;
[0065] Attached Figure 2 is a schematic diagram of the 3σ rule;
[0066] Attached Figure 3 is a schematic diagram of a bar chart. Specific embodiments
[0067] The real-time DNS gateway quality difference delimitation method based on waveform fitting and message queue of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0068] Embodiment:
[0069] As shown in the attached Figure 1 This embodiment provides a real-time DNS gateway quality difference delimitation method based on waveform fitting and message queue, and the method is as follows:
[0070] S1. Data collection: Aggregate DNS monitoring delay data starting from the underlying data through SQL;
[0071] S2. Obtain delay data: Put the DNS monitoring delay data into the Kafka message queue, and the data reading program consumes the Kafka message to obtain the DNS delay data of the network element;
[0072] S3. Data preprocessing: Fill in the missing values in the DNS delay data of the network element and convert the data format of the DNS delay data of the network element;
[0073] S4. Waveform fitting: Perform waveform fitting based on the union-find set on the DNS delay data of the network element to obtain a list of network elements with similar waveforms;
[0074] S5. Generate quality difference events: Judge whether the network elements with similar waveforms deteriorate, and judge whether the deteriorating network elements deteriorate together, eliminate the quality difference of the superior network element, and generate quality difference events;
[0075] S6. Data synchronization: Synchronize the data of the network elements involved in the quality difference events;
[0076] S7. Quality difference input: Input the quality difference events into the result table and start the quality difference delimitation of the next batch of network elements.
[0077] The input data of the delimiter program contains multiple levels of network element data, and each level of network element contains data from multiple sources. Taking the PON port as an example, it includes time series (daily granularity, hourly granularity, etc.), geographical location information (primary region, secondary region, tertiary region, etc.), and network performance indicators (gateway DNS monitoring delay). As shown in Table 1, the data collection in step S1 of this embodiment is to summarize data from multiple data sources through SQL statements and insert the result data into the input table, specifically as follows:
[0078] S101. Clear the data corresponding to the date in the input table;
[0079] S102. Since the delimiting period is 1-minute granularity, it is necessary to generate timestamp data every 1 minute; for example, 20250120000000, 20250120000001;
[0080] S103. Group the underlying network element data tables by network element name (primary region, secondary region, tertiary region, etc.), and filter out the network element data corresponding to the date; among them, since the underlying data is not necessarily completely standardized and there may be duplicates in the data collected by the device, when filtering, select the data record with the largest number of attached users as the network element data of the corresponding network element; among them, the network element data includes geographical location information and network performance indicator data;
[0081] S104. Connect the timestamp data and the network element data. Since there may be missing network performance indicators collected by the device, it is necessary to calculate the average value of the network performance indicators to fill in the missing network performance indicator data, and insert the processed result data into the input table.
[0082] Table 1 Table structure of the PON port input table
[0083]
[0084] The specific process of obtaining delay data in step S2 of this embodiment is as follows:
[0085] S201. Put the DNS delay data summarized every minute into the Kafka message queue, and each network element has a unique Topic message queue;
[0086] S202. Read the data in the kafka message queue through a data reading program based on the python language;
[0087] S203. DNS Latency Bounding: The data reading program continuously consumes the message queues of the Topics corresponding to each network element; each time it consumes, the data reading program obtains the DNS latency data for the previous n time periods; where n is a variable time period window, and to ensure the fitting effect, n is an integer greater than or equal to 3. After experiments with the four candidate values of 3, 5, 10, and 15, when n is set to 10, a good rhythm can be maintained between the bounding program and data aggregation.
[0088] The input of the data reading program in this embodiment is the kafka message queue, and the output is the DNS latency data; specifically: after installing the third-party library kafka-python, instantiate the KafkaConsumer class in the data reading program; during the instantiation process, set the IP address of the Kafka broker and the Topic name parameter; consume the Kafka message queue through the KafkaConsumer class to return the corresponding data.
[0089] The data reading program in this embodiment obtains the DNS latency data for the previous n time periods as follows:
[0090] First, when instantiating the KafkaConsumer in the data reading program, set auto.offset.reset to earliest;
[0091] Second, subtract n time periods from the current timestamp to obtain the timestamp of the start time; since Kafka uses millisecond-level timestamps, when calculating, convert both the current timestamp and n time periods to milliseconds;
[0092] Finally, calculate the offset corresponding to the start time, move the consumer to the offset position, start consuming the message queue, and obtain the DNS latency data.
[0093] Since the underlying device fails to successfully collect data at a certain time point during data collection, it may result in a network element having data in the previous time period but no data in the current time period, and filling is required. The filling operation performed at this time is different from the filling operation in step S1. Step S1 fills a certain field with missing values, while here it fills an entire missing data entry. This requires filling all fields, including time series, geographical location information, and network performance metrics.
[0094] To fill in the missing data, first, create a time code table containing n time periods; second, perform a left join on the time code table and the network element data; finally, for the fields of time series and geographical location information, fill them with the same values as in the context, and for the network performance metrics, fill them with 0; at the same time, adopt a cautious strategy. When no data is collected by the device, it is considered that there is no quality problem with the network element.
[0095] After filling in the missing data, convert the data format of each field, and convert it to int64 type, str type, and float64 type respectively according to the field type. Convert the fields time_key, time_hour, time_day, city_key, country_key, bras_key, olt_key, pon_key to int64 type, convert the fields city_name, country_name, bras_name, olt_name, pon_name to str type, and convert the field soft_dns_resp_delay to float64 type.
[0096] The waveform fitting in step S4 of this embodiment is specifically as follows:
[0097] S401. Obtain out_device containing the network element names from the input device_dict, combine all the network element names in pairs, and the function combinations() can return non-repeated combinations out_pairs;
[0098] S402. When calculating the similarity, use both Pearson similarity and cosine similarity: if the delay data of any network element is the same value within the time period, use cosine similarity for waveform fitting; otherwise, use Pearson similarity for waveform fitting;
[0099] S403. When all the network element combinations in out_pairs are calculated, return the fitted network element list similar_devices.
[0100]
[0101] Algorithm 1 only calculates the list of similar network elements for fitting, such as: [[a, b], [e, f], [c, e]]. Next, it is necessary to merge the similar_device. Union-Find is a data structure suitable for handling the merging and querying of disjoint sets. Using Union-Find, the merging of multiple network element lists can be achieved. If the intersection of two network element lists is not empty, the sets represented by these two network element lists are merged until there is no need for further merging. For the previous example, after merging, [[a, b], [c, e, f]] can be obtained.
[0102] The generation of quality difference events in step S5 of this embodiment is specifically as follows:
[0103] S501. Eliminate the deterioration of subordinate network elements caused by the failure of superior network elements; since this solution performs delimitation in the order from top to bottom, that is, primary area → secondary area → tertiary area → Bras → Olt → Pon. Therefore, the deterioration of subordinate network elements caused by the failure of superior network elements does not need to output quality difference events repeatedly. Taking Bras as an example, if the primary area, secondary area, and tertiary area to which network element A belongs have already had quality difference events during a certain event period, then the DNS monitoring delay of network element A during this time period is set to 0.
[0104] S502. Judge whether each network element in the network element list deteriorates: Use the 3σ rule to judge deterioration, that is, calculate a deterioration threshold through threshold = mean + σ × std: If the network element is higher than the threshold at any time point, it is considered to have deteriorated; traverse all network elements to judge the deterioration situation of each network element within the time period; where, threshold represents the quality difference threshold; mean represents the average value of DNS detection delay within the statistical time period; std represents the standard deviation of DNS detection delay within the statistical time period; σ takes a value of 2, that is, data outside the 95% range belongs to outliers, as shown in the appendix Figure 2 as shown;
[0105] S503. According to the deterioration situation of all network elements within the time period, judge whether there are multiple network elements deteriorating simultaneously within the time period:
[0106] If there is simultaneous deterioration, it means that a quality difference event is formed;
[0107] If the number of network elements involved in the quality difference event is the same as the number of network elements under the superior network element, it means that the deterioration of subordinate network elements is caused by the failure of the superior network element. According to the different levels of network elements, record the event category as primary area failure, secondary area failure, tertiary area failure, BRAS failure, OLT failure;
[0108] If the number of network elements involved in the poor quality event is less than the number of network elements under the superior network element and greater than 1, it indicates that multiple subordinate network elements deteriorate simultaneously. Record the event category as level 1 regional fault, level 2 regional fault, level 3 regional fault, inter-region B fault, inter-BO fault, multi-pon port fault;
[0109] After determining the poor quality event category, record the start date, start time, end time, and network element name of the deterioration event simultaneously to produce the top-level conclusion. For example: Level 1 regional name / Level 2 regional name / Level 3 regional name / Bras name: olt1, olt2 deteriorated on January 5, 2025, and the deterioration curves are highly fitted. The deterioration interval is: (01:40, January 5, 2025, 01:46, January 5, 2025).
[0110] In step S6 of this embodiment, since the front end needs to present the bar chart corresponding to the poor quality event, it is necessary to synchronize the input of the network elements involved in each poor quality event to another data table. The front end needs to query the data in these tables to produce the bar chart, as shown in the appendix Figure 3 shown.;
[0111] Among them, the table structure of the data table is exactly the same as the fields of the data in the Kafka message queue;
[0112] One poor quality event corresponds to one bar chart. Multiple network elements in the poor quality event correspond to several sub-charts in the bar chart, that is, each sub-chart is a bar chart of the DNS monitoring experiment of a network element; multiple network elements in one poor quality event produce waveform fitting, and several sub-charts in one bar chart present similar graphics.
[0113] In step S7 of this embodiment, in addition to displaying the bar chart, the front end also needs to present the details of the poor quality event. Therefore, input the poor quality event result data produced in step S5 into the result table. The table structure of the result table is shown in Table 2. If there are still network elements that have not been demarcated, start the poor quality demarcation of the next batch of network elements, otherwise the program ends.
[0114] Table 2 Table structure of the result table
[0115] Serial number Field name Field type Field explanation 1 time_key number(12) Time - minute 2 city_key number(10) City identifier 3 city_name varchar2(256) City name 4 country_key number(10) District identifier 5 country_name varchar2(256) District name 6 bras_key number(19) Attributed bras identifier 7 bras_name varchar2(256) Attributed bras name 8 olt_key number(19) Attributed OLT identifier 9 olt_name varchar2(256) Attributed OLT name 10 pon_port varchar2(256) Waveform fitting PON port 11 dia_result_id number(19) Delimiting conclusion identifier 12 dia_result varchar2(4000) Delimiting conclusion
[0116] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues, characterized in that The method is as follows: Data collection: Aggregate DNS monitoring latency data starting from the underlying data through SQL; Obtain latency data: Put the DNS monitoring latency data into the Kafka message queue, and the data reading program consumes the Kafka message to obtain the DNS latency data of the network element; Data preprocessing: Fill in the missing values in the DNS latency data of the network element and convert the data format of the DNS latency data of the network element; Waveform fitting: Perform waveform fitting based on the union-find set on the DNS latency data of the network element to obtain a list of network elements with similar waveforms; Generate poor quality events: Determine whether the network elements with similar waveforms deteriorate, and determine whether the deteriorated network elements deteriorate together, eliminate the poor quality of the upper-level network elements, and generate poor quality events; Data synchronization: Synchronize the data of the network elements involved in the poor quality events; Poor quality input: Input the poor quality events into the result table and start the poor quality delimitation of the next batch of network elements.
2. The real-time DNS gateway quality degradation demarcation method based on waveform fitting and message queue according to claim 1, wherein Data collection is to summarize data from multiple data sources through SQL statements and insert the result data into the input table, as follows: Clear the data corresponding to the date in the input table; Generate timestamp data every 1 minute; Group the underlying network element data tables by network element name and filter out the network element data corresponding to the date; among them, when filtering, select the data record with the largest number of attached users as the network element data of the corresponding network element; among them, the network element data includes geographical location information and network performance index data; Connect the timestamp data and the network element data, calculate the average value of the network performance index to fill in the missing network performance index data, and insert the processed result data into the input table.
3. The method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues according to claim 1, wherein Obtain latency data as follows: Put the DNS latency data summarized every minute into the Kafka message queue, and each network element has a unique Topic message queue; Read the data in the kafka message queue through a data reading program based on the python language; DNS latency delimitation: The data reading program continuously consumes the message queue corresponding to each network element; each time it consumes, the data reading program obtains the DNS latency data of the previous n time periods; among them, n is a variable time period window, and n is an integer greater than or equal to 3.
4. The method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues according to claim 3, wherein The input of the data reading program is the kafka message queue, and the output is the DNS latency data, as follows: After installing the kafka-python third-party library, instantiate the KafkaConsumer class in the data reading program; During the instantiation process, set the IP address of the Kafka broker and the Topic name parameter; Consume the Kafka message queue through the KafkaConsumer class to return the corresponding data.
5. The method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues according to claim 3, wherein The data reading program obtains the DNS latency data of the previous n time periods as follows: When instantiating the KafkaConsumer in the data reading program, set auto.offset.reset to earliest; Subtract n time periods from the current timestamp to obtain the timestamp of the start time. Since Kafka uses millisecond-level timestamps, both the current timestamp and n time periods are converted to milliseconds during the calculation. Calculate the offset corresponding to the start time, move the consumer to the offset position, start consuming the message queue, and obtain DNS delay data.
6. The real-time DNS gateway quality degradation delimitation method based on waveform fitting and message queue according to claim 1, characterized in that The data preprocessing is as follows: Create a time code table containing n time periods. Perform a left join on the time code table and the network element data. For the fields of time series and geographical location information, fill them with the same values as the context. For network performance metrics, fill them with 0. At the same time, adopt a cautious strategy. When no data is collected from the device, it is considered that there is no quality degradation in the network element. After filling in the missing data, convert the data format of each field, and convert them to int64 type, str type, and float64 type respectively according to the field type.
7. The method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues according to claim 1, characterized in that, The waveform fitting is as follows: Obtain the out_device containing the network element names from the input device_dict, combine all the network element names in pairs, and the function combinations() can return non-repeated combinations out_pairs. When calculating the similarity, both Pearson similarity and cosine similarity are used: If the delay data of any network element is the same value within the time period, cosine similarity is used for waveform fitting; otherwise, Pearson similarity is used for waveform fitting. When all the network element combinations in out_pairs are calculated, return the fitted network element list similar_devices.
8. The method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues according to claim 1, characterized in that, Generating quality degradation events is as follows: Eliminate the degradation of subordinate network elements caused by the failure of the superior network element. Judge whether each network element in the network element list is degraded: Use the 3σ rule to judge degradation, that is, calculate a degradation threshold through threshold = mean + σ × std: If a network element is higher than the threshold at any time point, it is considered degraded; traverse all network elements to judge the degradation situation of each network element within the time period. Among them, threshold represents the quality degradation threshold; mean represents the average DNS detection delay within the statistical time period; std represents the standard deviation of DNS detection delay within the statistical time period; σ takes the value of 2. According to the degradation situation of all network elements within the time period, judge whether there are multiple network elements degrading simultaneously within the time period: If there is simultaneous degradation, it means that a quality degradation event is formed.
9. The method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues according to claim 8, wherein The judgment of the quality degradation event category is as follows: If the number of network elements involved in the quality degradation event is the same as the number of network elements under the superior network element, it means that the inferior network elements are degraded due to the failure of the superior network element, and record the event category according to the different network element levels. If the number of network elements involved in the quality degradation event is less than the number of network elements under the superior network element and greater than 1, it means that multiple subordinate network elements are degraded simultaneously, and record the event category. After judging the quality degradation event category, record the start date, start time, end time, and network element name of the degradation event at the same time to make the top-level conclusion.
10. The method for determining poor quality boundaries of a real-time DNS gateway based on waveform fitting and message queues according to claim 1, wherein Data synchronization is to synchronize the inputs of network elements involved in each poor quality event to another data table, and a bar chart is made by querying the data table; Among them, the table structure of the data table is exactly the same as the fields of the data in the Kafka message queue; One poor quality event corresponds to one bar chart. Multiple network elements in a poor quality event correspond to several subcharts in the bar chart, that is, each subchart is a bar chart of the DNS monitoring experiment of a network element; Multiple network elements within a poor quality event produce a waveform fit, and several subcharts within a bar chart present similar graphs.