Voice call data analysis system based on cloud computing

Through cloud computing-based distributed collection and dynamic compression technology, the data loss and delay problems of traditional voice call data analysis systems in high-concurrency scenarios are solved, the accuracy of path stability assessment and anomaly detection is achieved, and data access efficiency and analysis quality are improved.

CN120751428AActive Publication Date: 2025-10-03RUGAO JIAYI INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511223442.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Traditional voice call data analysis systems are prone to data loss and processing delays in high-concurrency scenarios. Path stability assessments cannot distinguish between dynamic network fluctuations and persistent failures. Anomaly detection relies on single-dimensional statistics and has weak capabilities for identifying complex failure patterns. Static storage architectures cannot be elastically expanded, resulting in data redundancy and query delays.

Method used

Distributed collection nodes based on cloud computing are used to obtain call link parameters, and the sliding window difference algorithm is combined to calculate the path stability. The distributed K-means algorithm is used to cluster abnormal trajectories, dynamically compress feature vectors, and construct a cloud call behavior feature matrix to achieve elastic storage and query.

Benefits of technology

Accurately identify path stability segments, reduce data redundancy, improve anomaly detection accuracy, optimize storage structure, improve the efficiency of accessing massive data, and form an end-to-end quality analysis closed loop.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751428A_ABST
    Figure CN120751428A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of service signaling, in particular to a voice call data analysis system based on cloud computing, which comprises a signaling acquisition module, a stability evaluation module, a dynamic compression module, an abnormal track clustering module and a cloud feature library module. According to the method, link multi-dimensional parameters are collected through distributed nodes and stored in a time slice mode, a continuous three-hop rate difference is calculated through a sliding difference value, path stable segmentation is accurately recognized, time delay in a mean value compression screening segment is generated to generate feature vectors to reduce redundancy, cloud K-means integrates state duration, timeout rate and retransmission times to detect trajectory offset, and abnormal clustering labels are constructed. Time sequence association storage and elastic architecture dynamic screening path segments reduce invalid data processing amount, compression mean reduces calculation complexity, multi-dimensional clustering improves detection precision, a time window optimizes a storage structure, a sliding difference enhances fluctuation capture capability, elastic storage improves mass data access efficiency, and an end-to-end quality analysis closed loop is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of service signaling technology, and in particular to a voice call data analysis system based on cloud computing. Background Art

[0002] The field of service signaling technology encompasses various signaling technologies used to control the communication process in mobile communication networks. The core of this field is the establishment, maintenance, and release of communication services such as voice, data, and SMS through network control channels, involving multiple aspects such as signaling interaction, call control, session management, and mobility management. This technical field is based on standardized communication protocols and ensures the normal operation and service quality of communication networks by defining different message types and process mechanisms. In the system architecture, service signaling plays a role in coordinating the behavior between user equipment and the network. It is widely used in communication systems such as cellular networks, LTE, and 5G, and cooperates with core network functions to achieve user status management, service access, and resource allocation.

[0003] Among them, the voice call data analysis system refers to a system that collects and analyzes service signaling data related to voice calls in a mobile communication network. The technical matters targeted by this patent subject cover the extraction of signaling processes during call establishment, analysis of call delays and call paths, and location of call failure causes. It extracts key control data of calls by classifying and processing the call establishment request signaling, call connection signaling, and release signaling generated in the network, and combining timestamps and signaling parameters for orderly sorting and correlation mapping. Subsequently, the system classifies and counts different types of call behaviors according to predefined analysis logic, and uses state transition rules to analyze and judge call state changes, construct call behavior sequences and link structures, and achieve structured expression of call business processes.

[0004] Traditional technologies rely on centralized signaling collection and static storage, which is prone to data loss and processing delays in high-concurrency scenarios. Path stability assessment uses fixed thresholds, which cannot distinguish between dynamic network fluctuations and persistent failures, resulting in a high rate of misjudgment. Full signaling storage generates redundant data, interfering with feature extraction efficiency. Anomaly detection relies solely on single-dimensional statistics, failing to integrate multi-metric correlations, resulting in weak recognition of complex failure patterns. Static storage architectures lack elastic scalability, limiting the depth of historical data backtracking and trend analysis. For example, centralized collection leads to synchronization deviations between node hop counts and latency data, full storage increases query latency, single-dimensional rule matching makes it difficult to identify anomalies associated with retransmissions and timeouts, and fixed storage capacity limits long-term data accumulation. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and propose a cloud computing-based voice call data analysis system.

[0006] To achieve the above objectives, the present invention adopts the following technical solutions: A cloud computing-based voice call data analysis system includes: The signaling collection module is used to obtain the node hop count, response delay, state maintenance time, timeout ratio, and retransmission number of the call link through distributed collection nodes deployed on the cloud platform. After performing time window slicing, it is stored in the cloud time series database to generate a cloud signaling time series dataset and pass it to the stability assessment module. A stability evaluation module is configured to call a sliding window difference algorithm to calculate a three-hop transfer rate difference between the node hop count and the response delay in the cloud signaling timing data set, mark path segments that meet a threshold, generate a path stability segment label, and pass it to the dynamic compression module; a dynamic compression module, configured to filter path segments according to the path stability segmentation labels, perform distributed mean calculation on the response delays of all nodes in the segment, generate a cloud compression feature vector set, and pass the cloud compression feature vector set and the cloud signaling time series dataset to an abnormal trajectory clustering module; The abnormal trajectory clustering module is used to extract the state maintenance time, the timeout ratio, and the number of retransmissions, perform trajectory deviation detection through a distributed K-means algorithm based on the cloud platform, generate abnormal trajectory clustering labels, and pass them to the cloud feature library module.

[0007] As a further solution of the present invention, the path stability segmentation label includes the difference in continuous three-hop transfer rates, the segment where the rate difference threshold is satisfied, and the path stability level identifier. The cloud compression feature vector set specifically includes the screening path segment identifier, the distributed mean response delay, and the compression window delay distribution. The abnormal trajectory clustering label includes the offset trajectory cluster identifier, the state maintenance time offset, the timeout ratio cluster center, and the abnormal retransmission number threshold. The distributed collection nodes deployed on the cloud platform work together with the sliding window difference algorithm to reduce network jitter interference and improve path screening accuracy; The initial clustering centers of the distributed K-means algorithm based on the cloud platform are 5 randomly selected sample points, and the convergence condition is that the center offset is less than 0.5%.

[0008] As a further solution of the present invention, the signaling acquisition module includes: The signaling collection submodule uses distributed collection nodes deployed on the cloud platform to obtain the node hop count, response delay, state maintenance time, timeout ratio, and retransmission count of the call link. It also performs cross-regional clock synchronization calibration on the timestamps in the original data. It calculates the mean of multiple indicators based on a sliding window and removes data points that deviate from the mean by more than three standard deviations. It calculates the correlation coefficient between the node hop count and response delay to generate a node performance indicator group. The data points that deviate from the mean by more than 3 standard deviations are determined based on statistical analysis of historical network jitter data; The timing slicing submodule counts the number of node hops in segments according to the fixed time window length based on the node performance indicator group, uses the sliding average method to eliminate instantaneous fluctuations in the response delay, converts the timeout ratio and the number of retransmissions into percentage values ​​based on the proportion of event triggering times within the window, and generates window timing parameters; The sliding average method uses a window length of 5 seconds and weights are distributed according to exponential decay; The cloud storage submodule calls the window timing parameters, converts the number of node hops into an integer sequence code, maps the response delay and state maintenance time into double-precision floating-point values, writes the timeout ratio and the number of retransmissions into the statistical field according to the preset label classification rules, and generates a cloud signaling timing data set.

[0009] As a further solution of the present invention, the stability assessment module includes: The transfer rate calculation submodule calls the node hop count and the response delay field of the cloud signaling timing data set, intercepts the hop count difference and delay difference of three consecutive hops in chronological order, divides the hop count difference by the response delay difference, calculates the hop count change per unit delay in each window, and generates a transfer rate difference sequence; The delay difference is uniformly converted into milliseconds; Based on the transfer rate difference sequence, the path marking submodule sets the path stability threshold as the absolute fluctuation upper limit of the difference between the two windows. It traverses all the differences in the sequence, records the start and end timestamps of the windows that continuously exceed the threshold, and merges the adjacent window intervals with timestamp intervals less than the preset value to generate an abnormal path segment identifier. The absolute fluctuation limit is determined by calculating the moving average of the first 10% window differences; The path label generation submodule calls the abnormal path segment identifier, extracts the starting hop count, ending hop count and difference mean of each path segment, divides the stability level according to the preset interval of the difference mean, concatenates the level number and the path segment position code into a character string, and generates a path stability segment label.

[0010] As a further solution of the present invention, the dynamic compression module includes: The path screening submodule calls the path stability segmentation label, traverses the stability level numbers corresponding to the path segment codes in the label, compares the numbers with the preset path stability threshold value item by item, screens the path segment codes whose numbers are greater than the threshold, and extracts the node hop count and delay field storage address of the corresponding path segment from the cloud signaling timing data set based on the codes to generate a screening path segment identifier; The mean calculation submodule divides the nodes in the path segment into continuous intervals based on the screening path segment identifier and the hop value from low to high, accumulates the response delay values ​​of the nodes in each interval and counts the number of nodes, divides the total delay by the number of nodes to obtain the single interval mean, and combines the mean with the interval start hop count and end hop count to form a three-dimensional parameter group to generate the node delay mean; The vector generation submodule calls the node delay average, converts the starting hop count into an integer index value, maps the delay average value to a floating-point value with four decimal places, adds the ending hop count to the index value to generate a composite index key, and integrates it with the delay average value according to the field type to generate a cloud compression feature vector.

[0011] As a further solution of the present invention, the abnormal trajectory clustering module includes: The feature extraction submodule calls the state maintenance time, the timeout ratio, and the retransmission number fields, normalizes the values ​​of the multiple fields according to the path segment number, and splices the normalized field values ​​into a multidimensional vector in the order of the path segments to generate trajectory feature parameters; The normalization process uses the Min-Max formula to map the value to the interval [0,1]; The clustering calculation submodule uses a cloud-based distributed K-means algorithm based on the trajectory feature parameters to randomly select initial cluster centers, calculate the sum of the squared Euclidean distances between all vectors and the cluster center, assign the vector to the nearest cluster center, recalculate the mean of all vectors in each cluster as the new center, and iterate until the center position change amplitude is lower than the preset convergence condition to generate the trajectory cluster center; The cluster label generation submodule calls the trajectory cluster center, counts the proportion of path segments in each cluster, selects cluster numbers whose proportion is lower than the preset abnormal threshold, maps the path segments under the corresponding cluster into an integer label sequence, and generates abnormal trajectory cluster labels.

[0012] As a further embodiment of the present invention, the system further comprises: A cloud feature library module is used to associate the cloud compressed feature vector set and the abnormal trajectory clustering label in a time series and store them in an elastic storage service, generate a cloud call behavior feature matrix, and provide query services to the outside world through a cloud API interface; The cloud call behavior feature matrix specifically includes a time series correlation feature vector, a cluster label index, an elastic storage service address identifier, and a cloud API interface access key.

[0013] As a further solution of the present invention, the cloud feature library module includes: The associated storage submodule calls the cloud compressed feature vector and the abnormal trajectory cluster label, extracts the timestamp field of both, arranges the timestamps in ascending order and matches the corresponding feature vector storage address with the label storage address, binds the floating-point value of the feature vector and the integer code of the label in timestamp order and stores them in the elastic storage service to generate an associated feature dataset; The matrix generation submodule extracts all floating-point values ​​in timestamp order based on the associated feature dataset to construct a row vector, appends integer codes as independent columns to the end of the row vector, detects and deletes fields with missing values ​​or data type conflicts in the row vector, and arranges the verified row vectors into a two-dimensional structure in time series to generate a cloud call behavior feature matrix; The interface configuration submodule calls the cloud call behavior feature matrix, defines the matrix row index as the timestamp field, the column index as the combination of the feature number field and the label encoding field, configures the parameter input format of the query interface as the time range starting value and the feature number list, and generates a cloud feature query interface.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, multi-dimensional parameters of the call link are obtained through distributed collection nodes and time window slicing storage is implemented. The sliding window difference algorithm is combined to dynamically calculate the difference in the transfer rate of three consecutive hops to accurately identify the path stability segment. Distributed mean calculation performs compression processing on the response delay in the screened path segment, generates feature vectors to reduce data redundancy. The distributed K-means algorithm based on the cloud platform integrates the state maintenance time, timeout ratio and number of retransmissions to perform trajectory offset detection and construct abnormal clustering labels. Supported by time-series associated storage and elastic architecture, dynamic screening of path segments reduces the amount of invalid data processing, mean compression reduces computational complexity, and multi-dimensional clustering improves the accuracy of anomaly detection. Time window slicing optimizes the storage structure of time-series data, sliding difference enhances the ability to capture path fluctuations, and elastic storage improves the efficiency of accessing massive data, forming an end-to-end quality analysis closed loop. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a system flow chart of the present invention; Figure 2 This is a flow chart of the signaling acquisition module of the present invention; Figure 3 This is a flow chart of the stability assessment module of the present invention; Figure 4 This is a flow chart of the dynamic compression module of the present invention; Figure 5 This is a flow chart of the abnormal trajectory clustering module of the present invention; Figure 6 This is the flow chart of the cloud feature library module of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.

[0018] Example 1: Please refer to Figure 1 , the cloud computing-based voice call data analysis system includes: The signaling collection module is used to obtain the node hop count, response delay, state maintenance time, timeout ratio, and retransmission number of the call link through distributed collection nodes deployed on the cloud platform. After performing time window slicing, it is stored in the cloud time series database to generate a cloud signaling time series dataset and pass it to the stability assessment module. The stability assessment module is used to call the sliding window difference algorithm to calculate the difference in three-hop transfer rates based on the number of node hops and response delay in the cloud signaling time series dataset, mark the path segments that meet the threshold, generate path stability segment labels, and pass them to the dynamic compression module; The dynamic compression module is used to filter path segments based on path stability segmentation labels, perform distributed mean calculation on the response delays of all nodes in the segment, generate a cloud compression feature vector set, and pass the cloud compression feature vector set and the cloud signaling time series dataset to the abnormal trajectory clustering module; The abnormal trajectory clustering module is used to extract the state maintenance time, timeout ratio, and retransmission number, perform trajectory deviation detection through the distributed K-means algorithm based on the cloud platform, generate abnormal trajectory cluster labels, and pass them to the cloud feature library module; The cloud feature library module is used to associate the cloud compressed feature vector set and the abnormal trajectory clustering label in time series and store them in the elastic storage service, generate the cloud call behavior feature matrix, and provide query services to the outside world through the cloud API interface.

[0019] The path stability segmentation label includes the difference in transfer rates of three consecutive hops, the segment where the rate difference threshold is met, and the path stability level identifier. The cloud compression feature vector set specifically refers to the screening path segment identifier, distributed mean response delay, and compression window delay distribution. The abnormal trajectory clustering label includes the offset trajectory cluster identifier, state maintenance time offset, timeout ratio cluster center, and retransmission number abnormal threshold. The cloud call behavior feature matrix specifically includes the timing correlation feature vector, clustering label index, elastic storage service address identifier, and cloud API interface access key.

[0020] Distributed collection nodes deployed on the cloud platform work together with the sliding window difference algorithm to reduce network jitter interference and improve path screening accuracy; The initial clustering centers of the distributed K-means algorithm based on the cloud platform are 5 randomly selected sample points, and the convergence condition is that the center offset is less than 0.5%.

[0021] See also Figure 2 , the signaling acquisition module includes: The signaling collection submodule uses distributed collection nodes deployed on the cloud platform to obtain the node hop count, response delay, state maintenance time, timeout ratio, and retransmission count of the call link. It also performs cross-regional clock synchronization calibration on the timestamps in the original data. It calculates the mean of multiple indicators based on a sliding window and removes data points that deviate from the mean by more than three standard deviations. It calculates the correlation coefficient between the node hop count and response delay to generate a node performance indicator group. Data points that deviate from the mean by more than three standard deviations, determined based on statistical analysis of historical network jitter data; The signaling collection submodule proactively acquires key performance data related to a specific call link, such as a VoIP call from user A in Beijing to user B in Shanghai, through distributed collection nodes deployed on a cloud platform along network paths. Specific collection items include the number of node hops through which the data packet is transmitted, the response delay between nodes, the duration of the connection status, the timeout ratio during the communication process, and the number of data packet retransmissions. The four collection nodes (Node1, Node2, Node3, and Node4) deployed in Beijing, Jinan, Nanjing, and Shanghai each record the key time points and performance indicators of the data packet's journey.

[0022] The timestamps in the collected raw data may be inaccurate due to the different geographical locations of the collection nodes (involving different time zones) or slight deviations in the clocks of each node. Cross-regional clock synchronization calibration is required. This step selects a unified reference clock source and selects the Global Positioning System (GPS) time as a high-precision reference. The deviation is calculated by comparing the difference between the local clock of each collection node and the GPS time. The specific operation is: if the local clock of Node1 is measured to be 2 milliseconds faster than the GPS time and the local clock of Node4 is 1 millisecond slower than the GPS time, then when processing the data, subtract 2 milliseconds from all timestamps recorded by Node1 and add 1 millisecond to all timestamps recorded by Node4. Similarly, the timestamps of all relevant nodes are adjusted to ensure that subsequent analysis is based on a unified and accurate time reference.

[0023] After calibration, the sliding window method is used to perform statistical analysis on multiple indicators. The sliding window size is set to 60 seconds and the sliding step is 10 seconds. In each window, the collected indicator data is processed. Taking response delay as an example, the response delay data sequence collected in a specific time window W1 (for example, 14:30:00 to 14:31:00) is milliseconds, calculate the arithmetic mean of the delay within the window milliseconds, while calculating the standard deviation of these data points , the calculation process is to first find the variance , then the standard deviation milliseconds. Next, data points that deviate from the mean by more than 3 standard deviations are removed. The basis for setting this removal rule (3 standard deviations) is based on statistical analysis of historical network jitter data. The analysis results show that under similar network environments, regular network delay fluctuations rarely exceed The data points outside this range are considered to be significantly abnormal and need to be eliminated. The upper limit of the elimination threshold is calculated as milliseconds, with a lower limit of milliseconds. Since the latency cannot be negative, the actual lower limit is 0 milliseconds. The data point of 150 milliseconds is in the interval [0, 161.115] milliseconds, so it is not removed in this example. If there is a data point of 170 milliseconds in the window, it is removed because it exceeds the upper limit of 161.115 milliseconds. After removing the outlier, the mean and standard deviation of the window need to be recalculated. The same sliding window statistics and outlier data removal process is also performed for other collection indicators such as node hop count and state maintenance time.

[0024] Finally, the correlation between the number of node hops and the response delay is calculated, and the Pearson Correlation Coefficient is used to quantify the strength of the relationship between the two. The node hop number sequence within the same sliding window after synchronization calibration and elimination of outliers is extracted. (Assuming no abnormal hops are removed, the average is 4.375) and the response delay sequence (The mean is 70.875ms), calculate the covariance of the two sequences and their respective standard deviations and ms, correlation coefficient , obtained by calculation The value range of this coefficient is [-1,1]. 0.75 indicates that within this time window, the increase in the number of node hops is strongly positively correlated with the increase in response delay. The calculated statistical indicators, including the average number of node hops of 4.375, the average response delay of 70.875ms, the average state maintenance time, the average timeout ratio, the average number of retransmissions, and the correlation coefficient of 0.75, are combined into a structured data record, namely the node performance indicator group, and the timestamp of the corresponding window is attached and output to the subsequent module.

[0025] The timing slicing submodule counts the number of node hops in segments according to the fixed time window length based on the node performance indicator group. It uses the sliding average method to eliminate instantaneous fluctuations in response delay. It converts the timeout ratio and the number of retransmissions into percentage values ​​based on the proportion of event triggering times within the window to generate window timing parameters. The sliding average method uses a window length of 5 seconds, and the weights are distributed according to exponential decay; The time series slicing submodule receives the node performance indicator group sequence generated by the previous module, which contains the statistical results of each sliding window arranged in chronological order, such as timestamp The corresponding indicator group {hop count average , mean response delay ms, average timeout ratio , the average number of retransmissions This module first counts the node hop counts in segments according to a fixed time window length. The fixed window length is set to 1 minute. The number of times the average node hop count (rounded off or the main value determined according to the distribution) in the input indicator group is a specific value within 1 minute from 14:30:00 to 14:31:00 is counted. The statistical results are {hop count 4: 3 times, hop count 5: 2 times, hop count 6: 1 time}.

[0026] At the same time, the response delay is smoothed using the sliding average method to eliminate instantaneous sharp fluctuations in the data and highlight trend changes. The sliding average method specifies a window length of 5 seconds and uses exponential decay to assign weights. The weight calculation is based on the smoothing factor , which is set with the window length (or equivalent points) related to Rule: If the window length is 5 seconds and corresponds to 5 data points (assuming the data frequency is 1 Hz), then , current time point Smooth delay By formula Calculated, where is the average original response delay at the current time point, is the smoothed delay value calculated at the previous time point. The specific calculation is: if Average response delay at time ms, and the smooth delay of the previous moment ms, the current smoothing delay ms, this smoothing calculation is continuously applied to each delay value in the time series data stream.

[0027] In addition, the timeout ratio and retransmission count are converted into percentage values ​​according to the frequency of events occurring within the set 1-minute window. The specific method is to count the total number of packets processed within the window, as well as the number of timeout events and retransmission events detected. If a total of 1000 packets are processed within 1 minute, of which 20 timeout events and 50 retransmission events occur, the timeout ratio is calculated as , the ratio of retransmission times is The segmented hop count statistics obtained through the above processing {hop 4: 3 times, hop 5: 2 times, hop 6: 1 time}, the delay value corresponding to the end point of the 1-minute window in the calculated smoothed response delay sequence (for example, 62.10ms), and the converted timeout ratio are combined. and retransmission ratio Parameters such as these are integrated with the corresponding timestamps to generate window timing parameters, which are then output to the next module.

[0028] The cloud storage submodule calls the window timing parameters, converts the number of node hops into an integer sequence code, maps the response delay and state maintenance time into double-precision floating-point values, writes the timeout ratio and the number of retransmissions into the statistical field according to the preset label classification rules, and generates a cloud signaling timing data set.

[0029] The cloud storage submodule receives the window timing parameters generated by the timing slicing submodule. A typical parameter set is {timestamp: "2025-04-24 14:31:00", hop count: {4:3, 5:2, 6:1}, smoothed response delay: 62.10ms, mean state maintenance time: 120.5s, timeout ratio: 2%, retransmission ratio: 5%}. This module is responsible for converting these parameters into a format suitable for storage and subsequent analysis, and storing them in the cloud storage system to form a cloud signaling timing dataset.

[0030] First, the node hop count is converted into an integer sequence code. The selected encoding rule is to record the hop value that appears the most times within the window. In {4:3,5:2,6:1}, hop number 4 appears the most times (3 times), so it is encoded as the integer value 4. Then, the response delay and state maintenance time are mapped to double-precision floating-point values. The smoothed response delay of 62.10ms is stored as the floating-point number 62.10, and the average state maintenance time of 120.5s is stored as the floating-point number 120.5.

[0031] Next, the timeout ratio and retransmission count are written to the statistics field according to the preset label classification rules. The rule here is to store percentage values ​​directly as floating-point numbers. A timeout ratio of 2% is stored as 2.0, and a retransmission ratio of 5% is stored as 5.0. No interval division and labeling are performed to preserve more detailed original information for subsequent analysis. All processed and converted data items are integrated to form a structured storage record: {Timestamp: "2025-04-24 14:31:00", HopCountCode: 4, AvgSmoothedLatency: 62.10, AvgStatusTime: 120.5, TimeoutPercent: 2.0, RetransmissionPercent: 5.0}.

[0032] By continuously processing the parameters passed in each time window, a series of such records are generated and stored in a cloud storage service (such as a distributed database or a time series database) in chronological order, thereby constructing a cloud signaling time series dataset.

[0033] Table 1 Example of cloud signaling timing data Timestamp Hop Coding Smooth response delay (ms) State maintenance time (s) Timeout ratio (%) 2025-04-24 14:30:00 4 61.65 118.2 1.5 4.0 2025-04-24 14:31:00 4 62.10 120.5 2.0 5.0 2025-04-24 14:32:00 5 75.30 115.0 2.5 5.5 2025-04-24 14:33:00 5 78.00 110.8 3.2 7.0 2025-04-24 14:34:00 4 65.50 122.1 1.8 4.5 2025-04-24 14:35:00 4 64.00 125.3 1.6 4.2 2025-04-24 14:36:00 5 71.00 119.8 2.2 5.0 As shown in Table 1, this table lists some record samples of the cloud signaling time series dataset. Each row represents a snapshot of the call link status in a time window (1 minute interval), including the encoded main hop count, smoothed response delay (unit: milliseconds), average state maintenance time (unit: seconds), and directly recorded timeout and retransmission ratios (unit: %).

[0034] See also Figure 3 , the stability assessment module includes: The transfer rate calculation submodule uses the node hop count and response delay fields in the cloud signaling timing data set, extracts the hop count difference and delay difference of three consecutive hops in chronological order, divides the hop count difference by the corresponding delay difference, calculates the hop count change per unit delay in each window, and generates a transfer rate difference sequence. The delay difference is uniformly converted into milliseconds; The transfer rate calculation submodule uses the node hop count encoding and smoothed response delay fields in the cloud signaling timing dataset, as shown in the data records in Table 1. Its purpose is to calculate the change in the number of node hops per unit delay, thereby quantifying the change rate of the path structure.

[0035] The calculation process is carried out in chronological order. Each time, the records of three consecutive time points are intercepted to calculate the transfer rate between two adjacent time windows. Taking the records of the three time points of 14:30:00, 14:31:00, and 14:32:00 in Table 1 as an example, the corresponding hop count coding sequence is , the corresponding smooth response delay sequence is ms, first calculate the hop count difference of the first time interval (14:30 to 14:31) (unitless), delay difference ms, and then calculate the hop count difference for the second time interval (14:31 to 14:32) , delay difference ms.

[0036] Next, we divide the hop count difference by the corresponding delay difference to calculate the transfer rate, which is the change in hop count per unit delay. We use milliseconds for the delay difference because this is a fine time unit commonly used in the communications field and is suitable for capturing rapid changes. We calculate the transfer rate for the first window (corresponding to 14:31:00, reflecting changes in the previous minute). ,because , calculated Jump / millisecond, calculate the transfer rate of the second window (corresponding to the 14:32:00 time point) Jumps / millisecond.

[0037] Continue processing along the time series, considering the data at 14:31:00, 14:32:00, and 14:33:00, the hop count , delay ms, calculated , ms, transfer rate Hop / millisecond, considering the data at 14:32:00, 14:33:00, and 14:34:00, the hop count , delay ms, calculated , ms, transfer rate Jump / millisecond, and repeat the calculation to generate a time-ordered sequence of transfer rate differences. ,This sequence reflects the dynamics of the network path topology (measured in hop count) relative to ,delay changes.

[0038] The path marking submodule sets the path stability threshold as the absolute fluctuation limit of the difference between the two windows based on the transfer rate difference sequence. It traverses all the differences in the sequence, records the start and end timestamps of the windows that continuously exceed the threshold, and merges the adjacent window intervals with timestamp intervals less than the preset value to generate abnormal path segment identification. The absolute volatility cap is determined by calculating the moving average of the first 10% window differences; The path marking submodule receives the transfer rate difference sequence generated by the previous module ,The task is to identify the time period when the network path is significantly unstable based on this sequence.

[0039] First, a path stability threshold is set. This threshold is defined as the absolute fluctuation limit of the difference between the transfer rates of the two windows. The specific value of the threshold is determined by analyzing the first 10% time windows of the transfer rate sequence in historical data (for example, the past month). The specific calculation process is: select the first 10% of data points in the sequence R, calculate the absolute value of the difference between the two adjacent points, and obtain an absolute difference sequence. Then calculate the arithmetic mean of this absolute difference sequence and set this mean as the stability threshold. , suppose the sequence R has 100 data points, take the first 10 points to , calculate 9 absolute differences for to , and obtain the difference sequence , calculate the average of these values Hops / milliseconds / window, therefore, setting the path stability threshold .

[0040] Next, traverse the entire transfer rate difference sequence R and calculate the absolute value of the difference between all adjacent points , and compare this value with the threshold Compare and record the number of times the threshold is exceeded continuously The start and end timestamps of the time window, continue to use the sequence ,calculate , , , , , , , found from arrive The absolute value of the difference exceeds the threshold for 5 consecutive times, and the corresponding time period is calculated from A window after the time point (corresponding to the record at 14:31:00, the change occurred between 14:30-14:31, and the mark point is 14:31:00) (i.e. 14:32:00, corresponding to The time from the impact point to the calculation A window after the time point (corresponding to the record at 14:35:00, the change occurred between 14:34-14:35) (i.e. 14:36:00, corresponding to The time point at which the impact occurs) ends, and the starting timestamp is recorded and end timestamp .

[0041] Finally, the identified consecutive anomaly windows are merged. A time interval threshold, such as 2 minutes, is set to merge anomaly segments that are close in time. If a new anomaly segment starting at 14:37:00 and lasting to 14:39:00 is discovered immediately after the anomaly segment (ending at 14:36:00), the time interval between 14:37:00 and 14:36:00 is only 1 minute, less than the 2-minute merging threshold. These two anomaly segments are merged into a longer anomaly segment with a start time of 14:32:00 and an end time of 14:39:00. The final anomaly segment and its start and end timestamps are packaged to generate an anomaly path segment identifier, which is output as {ID: "PathSeg_001", StartTime: "2025-04-24 14:32:00", EndTime: "2025-04-24 14:39:00"}.

[0042] The path label generation submodule calls the abnormal path segment identifier, extracts the starting hop count, ending hop count and difference mean of each path segment, divides the stability level according to the preset interval of the difference mean, concatenates the level number and the path segment position code into a string, and generates a path stability segment label.

[0043] The path label generation submodule receives the abnormal path segment identifier, such as {ID: "PathSeg_001", StartTime: "2025-04-2414:32:00", EndTime: "2025-04-2414:39:00"}, and its task is to generate a quantitative stability level label for this identified abnormal path segment.

[0044] First, the node hop count codes of the starting time point (14:32:00) and the ending time point (assuming that there is a corresponding record at 14:39:00) of the path segment are extracted from the cloud signaling time series dataset (refer to Table 1 and its subsequent data) to obtain the starting hop count and termination hop count At the same time, all calculated transfer rate values ​​within this path segment (covering the time from 14:32:00 to 14:39:00) are extracted. ,Right now (Added here ), calculate the arithmetic mean of these transfer rate values Jumps / millisecond.

[0045] Then, according to the calculated mean transfer rate , refer to the preset stability level division interval rule, assign a level number to the path segment, the preset interval division rule is: Level 1 (stable) corresponds to , Level 2 (slight fluctuation) corresponds to , Level 3 (moderate volatility) corresponds to , Level 4 (violent fluctuations) corresponds to The setting of these intervals is based on the analysis of the impact of different fluctuations in historical data on call quality. The current calculated mean is , its absolute value , falls into the interval , so the stability level of this path segment is rated as level 3.

[0046] Finally, the obtained level number (3) is concatenated with the unique identifier of the path segment ("PathSeg_001") to form a structured string label with the format defined as "Level[level number]_[path segment ID]", i.e. "Level3_PathSeg_001". This process is repeated for all identified abnormal path segments to generate a list of path stability segment labels, such as ["Level3_PathSeg_001", "Level4_PathSeg_002", ...], which is passed to subsequent modules.

[0047] See also Figure 4 , the dynamic compression module includes: The path screening submodule calls the path stability segment label, traverses the stability level number corresponding to the path segment code in the label, compares the number with the preset path stability threshold, and screens the path segment codes with numbers greater than the threshold. At the same time, based on the code, it extracts the node hop count and delay field storage address of the corresponding path segment from the cloud signaling timing data set to generate the screening path segment identifier; The path filtering submodule receives a list of path stability segment labels, for example, L = [“Level3_PathSeg_001”, “Level4_PathSeg_002”, “Level2_PathSeg_003”], and its goal is to filter out path segments that need further analysis based on preset conditions.

[0048] Traverse each label string in the label list and extract the path segment stability level number contained therein. Extract level number 3 from "Level3_PathSeg_001", extract level number 4 from "Level4_PathSeg_002", and extract level number 2 from "Level2_PathSeg_003". Compare the extracted level numbers with the preset path stability screening threshold. To compare values, set ,This threshold is set based on the operation and maintenance requirements, ,and is intended to focus on those path segments that show moderate or ,higher instability, so as to conduct in-depth root cause analysis or ,optimization.,Comparison process: Level 3 vs. Compare, The result is true; level 4 and Compare, The result is true; level 2 and Compare, The result is false.

[0049] According to the comparison results, filter out the level numbers that are greater than or equal to the screening threshold The identifier of the path segment is obtained to obtain the filtered path segment ID list ["PathSeg_001", "PathSeg_002"]. Next, based on these filtered path segment IDs, the node hop count encoding sequence and the corresponding smoothed response delay sequence of each path segment in its corresponding time interval (PathSeg_001 corresponds to 14:32:00-14:39:00, PathSeg_002 is assumed to correspond to 14:50:00-14:55:00) are extracted from the storage address index or direct data cache of the cloud signaling timing data set (such as Table 1 and subsequent data). Specifically, the extracted content is the hop count sequence in the PathSeg_001 time period. and time-delayed sequences (The complete sequence is provided here to avoid ellipsis), as well as the data corresponding to PathSeg_002, the filtered path segment ID is combined with its extracted data (or the storage reference / address of the data) to generate the filtered path segment identifier, whose structure is [{ID:“PathSeg_001”,Data:{Hop:[5,5,4,4,4,5,5],Latency:[75.30,78.00,65.50,66.00,64.80,70.00,72.50]}},{ID:“PathSeg_002”,Data:{Hop:[…],Latency:[…]}}], which is passed to the next module.

[0050] The mean calculation submodule, based on the filtered path segment identifier, divides the nodes in the path segment into continuous intervals according to the hop count from low to high. It accumulates the response delay values ​​of the nodes in each interval and counts the number of nodes. The sum of the delays is divided by the number of nodes to obtain the mean of the single interval. The mean is then combined with the start and end hop counts of the interval to form a three-dimensional parameter group to generate the node delay mean. The mean calculation submodule receives and filters the path segment identifier and its associated data. For example, it processes the data of PathSeg_001 {ID: "PathSeg_001", Data: {Hop: [5, 5, 4, 4, 4, 5, 5], Latency: [75.30, 78.00, 65.50, 66.00, 64.80, 70.00, 72.50]}}. The purpose is to calculate the average response delay corresponding to each different hop value in the path segment.

[0051] First, analyze the hop count sequence [5, 5, 4, 4, 4, 5, 5] within the path segment and identify the different hop count values, including hop counts 4 and 5. Then group the data by hop count value. The delay data corresponding to hop count 4 is {65.50, 66.00, 64.80} ms, and the delay data corresponding to hop count 5 is {75.30, 78.00, 70.00, 72.50} ms.

[0052] Then, for each hop group, all response delay values ​​in the group are accumulated, and the number of data points in the group (that is, the number of times the hop appears) is counted. For a hop group of 4, the total delay is ms, number of data points , for a hop count of 5 packets, the total delay is ms, number of data points .

[0053] Next, divide the total delay of each packet by the number of corresponding data points to calculate the average response delay under the hop value. The average delay of hop number 4 is ms, average delay of hop 5 ms.

[0054] Finally, the calculated average delay and its corresponding hop count (which serves as the representative value for the interval, with the starting and ending hop counts being the same in this case) are combined into a three-dimensional parameter group in the format of [hop count, hop count, average delay]. The result for PathSeg_001 is two parameter groups: [4,4,65.43] and [5,5,73.95]. This calculation process is performed for all filtered path segments to generate a set of node average delay parameter groups, for example, [[4,4,65.43],[5,5,73.95]] (from PathSeg_001) and [[6,6,88.20],[7,7,95.10]] (assuming it is from PathSeg_002).

[0055] The vector generation submodule calls the node latency average, converts the starting hop count into an integer index value, maps the latency average to a floating-point value with four decimal places, adds the ending hop count to the index value to generate a composite index key, and integrates it with the latency average according to field type to generate a cloud compression feature vector.

[0056] The vector generation submodule receives a set of node delay mean parameter groups, such as [[4,4,65.43],[5,5,73.95]] (from PathSeg_001). The goal is to convert these calculation results into feature vectors in a fixed format to facilitate subsequent machine learning model processing.

[0057] The vectorization method used is to create a feature vector for each path segment. The dimension (or index) of this vector corresponds to a possible node hop value, and the value of the vector in this dimension is the average response delay of the corresponding hop number. If a hop number does not appear in the path segment, the value in this dimension is a preset fill value (0 is used here). The hop number range covered by the vector is set to 1 to 10, that is, the vector length is 10.

[0058] Process the parameter groups [4,4,65.43] and [5,5,73.95] from PathSeg_001, and round the average delay values ​​to four decimal places to get 65.4300 and 73.9500. Taking hop number 4 as the index, the corresponding value is 65.4300; taking hop number 5 as the index, the corresponding value is 73.9500. For other dimensions in the vector (index 1, 2, 3, 6, 7, 8, 9, 10), since these hop numbers do not appear in PathSeg_001, their values ​​are filled with 0. The final feature vector V1 generated for PathSeg_001 is: [0.0000, 0.0000, 0.0000, 65.4300, 73.9500, 0.0000, 0.0000, 0.0000, 0.0000, 0.0000].

[0059] The same conversion process is performed on all input path segment parameter group sets. For example, for the parameter group [[6,6,88.20],[7,7,95.10]] from PathSeg_002, the generated feature vector V2 is: [0.0000, 0.0000, 0.0000, 0.0000, 0.0000, 88.2000, 95.1000, 0.0000, 0.0000, 0.0000], and each generated vector is associated with its corresponding path segment ID to form a cloud compression feature vector set: [{ID:“PathSeg_001”, Vector:V1},{ID:“PathSeg_002”, Vector:V2},…].

[0060] See also Figure 5 ,The abnormal trajectory clustering module includes: The feature extraction submodule calls the state maintenance time, timeout ratio, and retransmission count fields, normalizes the values ​​of multiple fields according to the path segment number, and concatenates the normalized field values ​​into a multidimensional vector in the order of the path segments to generate trajectory feature parameters; Normalization uses the Min-Max formula to map the values ​​to the [0,1] interval; The purpose of the feature extraction submodule is to extract features of other dimensions besides delay and hop count for each abnormal path segment (such as PathSeg_001 and PathSeg_002 that were previously screened out) and integrate them.

[0061] Call the cloud signaling time series data set and extract the original data sequence of the state maintenance time, timeout ratio, and retransmission count fields corresponding to the abnormal path segment identifier (such as PathSeg_001 corresponding to the time period 14:32:00-14:39:00). For PathSeg_001, the extracted sequence is: state maintenance time s (assuming this segment contains 5 recording points), timeout ratio %, retransmission ratio %.

[0062] Calculate the average value of each field in the path segment, and the average state maintenance time of PathSeg_001 s, average timeout ratio , average retransmission ratio , perform the same calculation on PathSeg_002 and get .

[0063] The average values ​​calculated for all abnormal path segments are normalized using the Min-Max normalization method. The formula is: , linearly map the values ​​to the [0,1] interval. Before normalization, it is necessary to first determine the global minimum and maximum values ​​of each feature dimension (state maintenance time, timeout ratio, retransmission ratio) in all abnormal path segment data to be processed. By scanning the average value data of all abnormal path segments, determine: state maintenance time s, s; timeout ratio , ;Retransmission ratio , .

[0064] Apply the Min-Max formula to normalize the three average values ​​of PathSeg_001: Normalized state maintenance time , normalized timeout ratio , normalized retransmission ratio .

[0065] The normalized field values ​​of each path segment are concatenated in the order of (state maintenance time, timeout ratio, retransmission ratio) to form a multidimensional vector. For PathSeg_001, the generated trajectory feature vector is , the same process is performed on PathSeg_002 to obtain , and finally generate a set of trajectory feature parameters, including the feature vectors of all abnormal path segments .

[0066] The clustering calculation submodule uses a cloud-based distributed K-means algorithm to randomly select initial cluster centers based on trajectory feature parameters, calculates the sum of squared Euclidean distances between all vectors and the cluster center, assigns the vector to the nearest cluster center, and recalculates the mean of all vectors in each cluster as the new center. It iterates until the change in the center position is lower than the preset convergence condition to generate the trajectory cluster center. The clustering calculation submodule receives the trajectory feature parameter set, namely ,in ,The goal is to divide abnormal path segments with similar characteristics into the same group (cluster).

[0067] Use the K-means clustering algorithm implemented by cloud distribution to set the number of clusters , the value Usually based on business understanding (what kinds of abnormal patterns are expected) or determined through technical evaluation such as the elbow rule, the algorithm starts by randomly selecting Initial cluster centers are selected from the input data set. Randomly draw without replacement vector as the initial center , let the randomly selected initial center be , , .

[0068] Entering the iterative process, in each iteration, the first step is to perform the distribution step: calculate each vector in the data set To all current cluster centers The square of the Euclidean distance, the formula for the square of the Euclidean distance is ,in is the vector dimension (here 3), for vector , and calculate the square of its distance from the three initial centers: ; ; ; Comparing these three distance squared values, distance The closest (0.0856 minimum), so Assigned to the first cluster, for all This allocation is performed for each vector.

[0069] Then perform the update step: for each cluster , recalculate its cluster center , the new center is the arithmetic mean (centroid) of all vectors currently assigned to the cluster, i.e. ,in It is After rounds of iteration, the clusters are assigned Vector collection of is the number of vectors in the set.

[0070] Repeat the allocation and update steps until the preset convergence condition is met. The convergence condition is set as follows: the sum of the moving distances (Euclidean distances) of all cluster centers in two consecutive iterations is less than a minimum threshold. , or reaches the preset maximum number of iterations, the threshold Set according to the required accuracy, such as , when the final iteration stops, the cluster centers obtained This is the final trajectory cluster center.

[0071] The cluster label generation submodule calls the trajectory cluster center, counts the proportion of path segments in each cluster, filters the cluster numbers whose proportion is lower than the preset abnormal threshold, maps the path segments under the corresponding cluster into an integer label sequence, and generates abnormal trajectory cluster labels.

[0072] The cluster label generation submodule receives the final trajectory cluster center output by the cluster calculation submodule and the cluster number to which each data point (abnormal path segment vector) ultimately belongs.

[0073] First, count the proportion of the number of path segments contained in each cluster to the total number of abnormal path segments. Suppose there are Abnormal path segments are clustered, clustering Contains path segments, the clustering The proportion of , calculate all The proportion of clusters .

[0074] Then, we filter out clusters representing abnormal behavior patterns and set a preset abnormality ratio threshold. The threshold is set based on experience, and it is believed that clusters with too low a proportion may represent rare but critical abnormal patterns. , filter out all Cluster number ,The clusters corresponding to these numbers are regarded as abnormal trajectory clusters that need special attention.

[0075] Map all path segments belonging to these filtered abnormal clusters into an integer label sequence, and the value of the label is the number of the abnormal cluster to which the path segment belongs. For example, if the proportion of cluster 2 and cluster 3 is less than 5%, then all path segments originally assigned to cluster 2 will have their cluster label set to 2, and all path segments originally assigned to cluster 3 will have their cluster label set to 3. ), its cluster label is also set to the cluster number 1 to which it belongs.

[0076] Finally, a unique cluster label is generated for each abnormal path segment participating in the clustering (such as "PathSeg_001", "PathSeg_002", etc.), which directly corresponds to the number of the cluster to which it is assigned during the clustering process. For example, if PathSeg_001 ultimately belongs to cluster 1, PathSeg_002 belongs to cluster 3 (and the proportion of cluster 3 is 4%, which is lower than the 5% threshold), PathSeg_003 belongs to cluster 1, PathSeg_004 belongs to cluster 2 (accounting for 36%), and PathSeg_005 belongs to cluster 3, then the output cluster label result can be expressed as: {"PathSeg_001": 1, "PathSeg_002": 3, "PathSeg_003": 1, "PathSeg_004": 2, "PathSeg_005": 3}. This result set identifies the characteristic pattern categories to which different abnormal path segments belong. In particular, by combining the cluster proportion information, those that belong to low proportions (lower than ) Clustered path segments that may represent more special or critical abnormal behaviors.

[0077] Table 2 Examples of abnormal path segment clustering labels Segment ID Trajectory feature vector (normalized: hold time, timeout%, retransmission%) Cluster number Cluster proportion (%) Is it below the threshold (5%)? Final Label PathSeg_001 [0.4767, 0.1800, 0.2489] 1 60 no 1 PathSeg_002 [0.8500, 0.7500, 0.8800] 3 4 yes 3 PathSeg_003 [0.5100, 0.2200, 0.2800] 1 60 no 1 PathSeg_004 [0.1500, 0.0800, 0.1200] 2 36 no 2 PathSeg_005 [0.7800, 0.8100, 0.9100] 3 4 yes 3 Table 2 shows the label generation process for some anomalous path segments after clustering. The table lists each path segment's ID, its corresponding normalized trajectory feature vector (with dimensions representing state maintenance time, timeout ratio, and retransmission ratio, in that order), the cluster number assigned using the K-means algorithm, the proportion of path segments included in that cluster relative to the total number of path segments, whether that proportion falls below the preset 5% anomaly threshold, and the final cluster label assigned to the path segment (i.e., the cluster number to which it belongs). This table shows that PathSeg_002 and PathSeg_005 are classified into Cluster 3, which accounts for only 4% of the total, marking them as anomalous patterns requiring special attention.

[0078] See also Figure 6 , the cloud feature library module includes: The associated storage submodule calls the cloud compressed feature vector and the abnormal trajectory clustering label, extracts the timestamp field of both, sorts the timestamps in ascending order, matches the corresponding feature vector storage address with the label storage address, and binds the floating-point value of the feature vector and the integer code of the label in timestamp order to the elastic storage service to generate the associated feature dataset. The association storage submodule receives a set of cloud-compressed feature vectors, such as {ID: "PathSeg_001", Vector: V1}, {ID: "PathSeg_002", Vector: V2}, ..., {ID: "PathSeg_005", Vector: V5}, and a set of abnormal trajectory cluster labels, { "PathSeg_001": 1, "PathSeg_002": 3,"PathSeg_003": 1, "PathSeg_004": 2, "PathSeg_005": 3}. The core task of this module is to associate the vectors representing path segment delay features with their corresponding behavior pattern cluster labels using timestamps and store them for subsequent construction of the time series feature matrix.

[0079] First, extract the timestamp information corresponding to each path segment identifier (such as "PathSeg_001"). This timestamp uses the end time of the abnormal path segment in the original time series data as a unique identifier. Based on the context examples in items 5 and 12, set the end timestamps of each path segment as follows: PathSeg_001 corresponds to "2025-04-24 14:39:00", PathSeg_002 corresponds to "2025-04-24 14:55:00", PathSeg_003 corresponds to "2025-04-24 15:05:00", PathSeg_004 corresponds to "2025-04-24 15:12:00", and PathSeg_005 corresponds to "2025-04-24 15:20:00".

[0080] Next, store the feature vector (vector) and cluster label (label) for each path segment in a designated storage location and retrieve their storage addresses or references. Feature vectors (floating-point arrays) are best stored in object storage services (such as AWS S3, Google Cloud Storage) or specialized vector databases for efficient storage and retrieval. Cluster labels (integers) can be stored in key-value stores (such as Redis) or simple database tables. Assume the following addresses are obtained after storage: PathSeg_001: Vector address vec_loc_1 = " / feature_store / vectors / seg001.vec", Label address lbl_loc_1 = " / feature_store / labels / seg001.lbl" PathSeg_002: Vector address vec_loc_2 = " / feature_store / vectors / seg002.vec", Label address lbl_loc_2 = " / feature_store / labels / seg002.lbl" PathSeg_003: Vector address vec_loc_3 = " / feature_store / vectors / seg003.vec", Label address lbl_loc_3 = " / feature_store / labels / seg003.lbl" PathSeg_004: Vector address vec_loc_4 = " / feature_store / vectors / seg004.vec", Label address lbl_loc_4 = " / feature_store / labels / seg004.lbl" PathSeg_005: Vector address vec_loc_5 = " / feature_store / vectors / seg005.vec", Label address lbl_loc_5 = " / feature_store / labels / seg005.lbl" Then, the timestamp of each path segment is bound to its corresponding feature vector storage address and label storage address to form an associated record. For example, the record corresponding to PathSeg_001 is: {"timestamp": "2025-04-24 14:39:00", "vector_loc": vec_loc_1, "label_loc": lbl_loc_1}.

[0081] Finally, all these associated records are sorted in ascending order by the timestamp field and stored in an elastic storage service optimized for time series data, such as a time series database (such as InfluxDB or TimescaleDB) or a NoSQL database configured with a time-ordered index. This storage service is chosen because it efficiently supports queries and data aggregation by time range, which is crucial for subsequent feature matrix generation and query interface provision. Once stored, a dataset of associated features is formed, logically structured as a time-ordered list, with each element containing a timestamp and a pointer to the storage location of the feature vector and cluster label corresponding to that point in time.

[0082] The matrix generation submodule extracts all floating-point values ​​in timestamp order from the associated feature dataset to construct a row vector. It then appends integer codes as independent columns to the end of the row vector. It then detects and deletes fields with missing values ​​or conflicting data types in the row vector. The validated row vectors are then arranged into a two-dimensional structure in time series to generate a cloud call behavior feature matrix. The matrix generation submodule is based on the associated feature dataset created in the previous step. Its goal is to construct a structured two-dimensional matrix, namely the cloud call behavior feature matrix, which integrates the temporal feature information of all screened abnormal path segments.

[0083] First, we traverse each record in the associated feature dataset in timestamp order. For each record, such as {"timestamp": "2025-04-24 14:39:00", "vector_loc": vec_loc_1, "label_loc": lbl_loc_1}, we retrieve the actual feature vector (floating-point array) and cluster label (integer encoding) from the corresponding storage location (object storage, vector database, key-value storage, etc.) based on the storage addresses vector_loc and label_loc contained in it. Based on the previous example data: Timestamp 14:39:00: Vector V1 = [0.0, 0.0, 0.0, 65.43, 73.95, 0.0,0.0, 0.0, 0.0, 0.0], Label L1 = 1 Timestamp 14:55:00: Vector V2 = [0.0, 0.0, 0.0, 0.0, 0.0, 88.20,95.10, 0.0, 0.0, 0.0], Label L2 = 3 Timestamp 15:05:00: Vector V3 = [0.0, 0.0, 0.0, 68.10, 75.50, 0.0,0.0, 0.0, 0.0, 0.0], Label L3 = 1 Timestamp 15:12:00: Vector V4 = [0.0, 0.0, 55.00, 62.30, 0.0, 0.0,0.0, 0.0, 0.0, 0.0], Label L4 = 2 Timestamp 15:20:00: Vector V5 = [0.0, 0.0, 0.0, 0.0, 0.0, 90.10,98.50, 105.20, 0.0, 0.0], Label L5 = 3 Next, a row vector is constructed for each timestamp. This row vector consists of two parts: first, the value of the retrieved 10-dimensional floating-point feature vector, and then the retrieved integer cluster label encoding is appended to the end of the vector.

[0084] Row for 14:39:00: [0.0, 0.0, 0.0, 65.43, 73.95, 0.0, 0.0, 0.0, 0.0,0.0, 1] Row for 14:55:00: [0.0, 0.0, 0.0, 0.0, 0.0, 88.20, 95.10, 0.0, 0.0,0.0, 3] Row for 15:05:00: [0.0, 0.0, 0.0, 68.10, 75.50, 0.0, 0.0, 0.0, 0.0,0.0, 1] Row for 15:12:00: [0.0, 0.0, 55.00, 62.30, 0.0, 0.0, 0.0, 0.0, 0.0,0.0, 2] Row for 15:20:00: [0.0, 0.0, 0.0, 0.0, 0.0, 90.10, 98.50, 105.20,0.0, 0.0, 3] After constructing the row vectors, perform data validation and cleaning. Each row vector is checked for missing values ​​(e.g., due to storage or retrieval errors that prevented the complete vector or label from being retrieved) or fields with conflicting data types (e.g., non-integer label values). A processing rule is set: if any row vector containing missing values ​​or conflicting data types is detected, the entire row is removed from the matrix to be constructed. This rule ensures the data integrity and consistency of the final matrix. In this example, it is assumed that all data was successfully retrieved and of the correct type, and no rows are deleted.

[0085] Finally, all verified row vectors are arranged vertically in the order of their corresponding timestamps to form a two-dimensional structure, which is the cloud call behavior feature matrix.

[0086] Table 3 Example of cloud call behavior feature matrix Timestamp (RowIndex) F1 (Hop1Latency) F2 (Hop2Latency) F3 (Hop3Latency) F4 (Hop4Latency) F5 (Hop5Latency) F6 (Hop6Latency) F7 (Hop7Latency) F8 (Hop8Latency) F9 (Hop9Latency) F10 (Hop10Latency) Label(Cluster ID) 2025-04-2414:39:00 0.00 0.00 0.00 65.43 73.95 0.00 0.00 0.00 0.00 0.00 1 2025-04-2414:55:00 0.00 0.00 0.00 0.00 0.00 88.20 95.10 0.00 0.00 0.00 3 2025-04-2415:05:00 0.00 0.00 0.00 68.10 75.50 0.00 0.00 0.00 0.00 0.00 1 2025-04-2415:12:00 0.00 0.00 55.00 62.30 0.00 0.00 0.00 0.00 0.00 0.00 2 2025-04-2415:20:00 0.00 0.00 0.00 0.00 0.00 90.10 98.50 105.20 0.00 0.00 3 Table 3 shows a portion of the generated cloud call behavior feature matrix. Each row represents the end time of an abnormal path segment, with the row index being the timestamp. Each column represents a feature dimension or label. The column indices, from left to right, are the average latency features (F1 to F10, in milliseconds) corresponding to 10 hops, followed by the cluster label encoding. This matrix integrates key temporal feature information for subsequent analysis and application.

[0087] The interface configuration submodule calls the cloud call behavior feature matrix, defines the matrix row index as the timestamp field, the column index as the combination of the feature number field and the label encoding field, configures the query interface parameter input format as the time range start value and the feature number list, and generates a cloud feature query interface.

[0088] The interface configuration submodule utilizes the cloud call behavior feature matrix generated in the previous step (as shown in Table 3 ) and aims to provide a standardized query interface that allows external systems or users to retrieve specific call behavior data based on time and feature dimensions.

[0089] First, define the matrix's data access structure. Explicitly specify the matrix's row index as a timestamp field with a datetime data type. Define the matrix's column index as a composite structure: the first 10 columns (F1 to F10) correspond to feature numbers 1 to 10 (representing the average latency at different hop counts), and the last column corresponds to the label encoding field.

[0090] Next, configure the parameter input format of the query interface. The interface is designed to accept the following parameters: startTime: The starting value of the query time range, in the format of an ISO 8601 timestamp string (for example, "2025-04-24T14:00:00Z").

[0091] endTime: The end value of the query time range, in the same format as above (for example, "2025-04-24T15:00:00Z").

[0092] featureIDs: A list (integer array) containing the desired feature IDs, for example [4, 5] means you want to query feature 4 (Hop4 Latency) and feature 5 (Hop5 Latency).

[0093] The internal logic of the query interface performs operations based on the input parameters: Based on startTime and endTime, filter out all rows in the cloud call behavior feature matrix whose row index (timestamp) falls within this time range.

[0094] For each filtered row, extract the corresponding feature column data based on the featureIDs list. At the same time, extract the data of the timestamp column and label encoding column.

[0095] Organize the extracted data into a structured response format.

[0096] Set the response format to a JSON array. Each object in the array represents an original matrix row (or a data snapshot at a time point) that meets the time range and contains the following key-value pairs: timestamp: The timestamp string corresponding to this row.

[0097] features: An object whose keys are the requested feature numbers (strings, such as "4", "5") and whose values ​​are the corresponding feature values ​​(floating point numbers).

[0098] label: The cluster label encoding corresponding to this row (integer).

[0099] Query example: Assume that a user initiates a query request with the following parameters: startTime = "2025-04-24T14:30:00Z" endTime = "2025-04-24T15:10:00Z" featureIDs = [4, 6] Interface execution logic: Filter the rows in Table 3 whose timestamps are between 14:30:00Z and 15:10:00Z to obtain the two rows corresponding to 14:39:00 and 15:05:00.

[0100] For the 14:39:00 row, extract feature 4 (65.43) and feature 6 (0.00), along with the label (1).

[0101] For the 15:05:00 row, extract feature 4 (68.10) and feature 6 (0.00), along with the label (1).

[0102] Organized into JSON response.

[0103] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. The cloud computing-based voice call data analysis system is characterized by: The system comprises: The signaling collection module is used to obtain the node hop count, response delay, state maintenance time, timeout ratio, and retransmission number of the call link through distributed collection nodes deployed on the cloud platform. After performing time window slicing, it is stored in the cloud time series database to generate a cloud signaling time series dataset and pass it to the stability assessment module. A stability evaluation module is configured to call a sliding window difference algorithm to calculate a three-hop transfer rate difference between the node hop count and the response delay in the cloud signaling timing data set, mark path segments that meet a threshold, generate a path stability segment label, and pass it to the dynamic compression module; a dynamic compression module, configured to filter path segments according to the path stability segmentation labels, perform distributed mean calculation on the response delays of all nodes in the segment, generate a cloud compression feature vector set, and pass the cloud compression feature vector set and the cloud signaling time series dataset to an abnormal trajectory clustering module; The abnormal trajectory clustering module is used to extract the state maintenance time, the timeout ratio, and the number of retransmissions, perform trajectory deviation detection through a distributed K-means algorithm based on the cloud platform, generate abnormal trajectory clustering labels, and pass them to the cloud feature library module.

2. The cloud computing-based voice call data analysis system according to claim 1, characterized in that: The path stability segmentation label includes the difference in continuous three-hop transfer rates, the segment where the rate difference threshold is met, and the path stability level identifier. The cloud compression feature vector set specifically includes the screening path segment identifier, the distributed mean response delay, and the compression window delay distribution. The abnormal trajectory clustering label includes the offset trajectory cluster identifier, the state maintenance time offset, the timeout ratio cluster center, and the abnormal retransmission number threshold. The distributed collection nodes deployed on the cloud platform work together with the sliding window difference algorithm to reduce network jitter interference and improve path screening accuracy; The initial clustering centers of the distributed K-means algorithm based on the cloud platform are 5 randomly selected sample points, and the convergence condition is that the center offset is less than 0.5%.

3. The cloud computing-based voice call data analysis system according to claim 2, characterized in that: The signaling acquisition module includes: The signaling collection submodule uses distributed collection nodes deployed on the cloud platform to obtain the node hop count, response delay, state maintenance time, timeout ratio, and retransmission count of the call link. It also performs cross-regional clock synchronization calibration on the timestamps in the original data. It calculates the mean of multiple indicators based on a sliding window and removes data points that deviate from the mean by more than three standard deviations. It calculates the correlation coefficient between the node hop count and response delay to generate a node performance indicator group. The data points that deviate from the mean by more than 3 standard deviations are determined based on statistical analysis of historical network jitter data; The timing slicing submodule counts the number of node hops in segments according to the fixed time window length based on the node performance indicator group, uses the sliding average method to eliminate instantaneous fluctuations in the response delay, converts the timeout ratio and the number of retransmissions into percentage values ​​based on the proportion of event triggering times within the window, and generates window timing parameters; The sliding average method uses a window length of 5 seconds and weights are distributed according to exponential decay; The cloud storage submodule calls the window timing parameters, converts the number of node hops into an integer sequence code, maps the response delay and state maintenance time into double-precision floating-point values, writes the timeout ratio and the number of retransmissions into the statistical field according to the preset label classification rules, and generates a cloud signaling timing data set.

4. The cloud computing-based voice call data analysis system according to claim 3, characterized in that: The stability assessment module includes: The transfer rate calculation submodule calls the node hop count and the response delay field of the cloud signaling timing data set, intercepts the hop count difference and delay difference of three consecutive hops in chronological order, divides the hop count difference by the response delay difference, calculates the hop count change per unit delay in each window, and generates a transfer rate difference sequence; The delay difference is uniformly converted into milliseconds; Based on the transfer rate difference sequence, the path marking submodule sets the path stability threshold as the absolute fluctuation upper limit of the difference between the two windows. It traverses all the differences in the sequence, records the start and end timestamps of the windows that continuously exceed the threshold, and merges the adjacent window intervals with timestamp intervals less than the preset value to generate an abnormal path segment identifier. The absolute fluctuation limit is determined by calculating the moving average of the first 10% window differences; The path label generation submodule calls the abnormal path segment identifier, extracts the starting hop count, ending hop count and difference mean of each path segment, divides the stability level according to the preset interval of the difference mean, concatenates the level number and the path segment position code into a character string, and generates a path stability segment label.

5. The cloud computing-based voice call data analysis system according to claim 4, characterized in that: The dynamic compression module includes: The path screening submodule calls the path stability segmentation label, traverses the stability level numbers corresponding to the path segment codes in the label, compares the numbers with the preset path stability threshold value item by item, screens the path segment codes whose numbers are greater than the threshold, and extracts the node hop count and delay field storage address of the corresponding path segment from the cloud signaling timing data set based on the codes to generate a screening path segment identifier; The mean calculation submodule divides the nodes in the path segment into continuous intervals based on the screening path segment identifier and the hop value from low to high, accumulates the response delay values ​​of the nodes in each interval and counts the number of nodes, divides the total delay by the number of nodes to obtain the single interval mean, and combines the mean with the interval start hop count and end hop count to form a three-dimensional parameter group to generate the node delay mean; The vector generation submodule calls the node delay average, converts the starting hop count into an integer index value, maps the delay average value to a floating-point value with four decimal places, adds the ending hop count to the index value to generate a composite index key, and integrates it with the delay average value according to the field type to generate a cloud compression feature vector.

6. The cloud computing-based voice call data analysis system according to claim 5, characterized in that: The abnormal trajectory clustering module includes: The feature extraction submodule calls the state maintenance time, the timeout ratio, and the retransmission number fields, normalizes the values ​​of the multiple fields according to the path segment number, and splices the normalized field values ​​into a multidimensional vector in the order of the path segments to generate trajectory feature parameters; The normalization process uses the Min-Max formula to map the value to the interval [0,1]; The clustering calculation submodule uses a cloud-based distributed K-means algorithm based on the trajectory feature parameters to randomly select initial cluster centers, calculate the sum of the squared Euclidean distances between all vectors and the cluster center, assign the vector to the nearest cluster center, recalculate the mean of all vectors in each cluster as the new center, and iterate until the center position change amplitude is lower than the preset convergence condition to generate the trajectory cluster center; The cluster label generation submodule calls the trajectory cluster center, counts the proportion of path segments in each cluster, selects cluster numbers whose proportion is lower than the preset abnormal threshold, maps the path segments under the corresponding cluster into an integer label sequence, and generates abnormal trajectory cluster labels.

7. The cloud computing-based voice call data analysis system according to claim 6, characterized in that: The system further comprises: A cloud feature library module is used to associate the cloud compressed feature vector set and the abnormal trajectory clustering label in a time series and store them in an elastic storage service, generate a cloud call behavior feature matrix, and provide query services to the outside world through a cloud API interface; The cloud call behavior feature matrix specifically includes a time series correlation feature vector, a cluster label index, an elastic storage service address identifier, and a cloud API interface access key.

8. The cloud computing-based voice call data analysis system according to claim 7, characterized in that: The cloud feature library module includes: The associated storage submodule calls the cloud compressed feature vector and the abnormal trajectory cluster label, extracts the timestamp field of both, arranges the timestamps in ascending order and matches the corresponding feature vector storage address with the label storage address, binds the floating-point value of the feature vector and the integer code of the label in timestamp order and stores them in the elastic storage service to generate an associated feature dataset; The matrix generation submodule extracts all floating-point values ​​in timestamp order based on the associated feature dataset to construct a row vector, appends integer codes as independent columns to the end of the row vector, detects and deletes fields with missing values ​​or data type conflicts in the row vector, and arranges the verified row vectors into a two-dimensional structure in time series to generate a cloud call behavior feature matrix; The interface configuration submodule calls the cloud call behavior feature matrix, defines the matrix row index as the timestamp field, the column index as the combination of the feature number field and the label encoding field, configures the parameter input format of the query interface as the time range starting value and the feature number list, and generates a cloud feature query interface.

Citation Information

Patent Citations

  • Method, device and equipment for positioning voice quality problem and medium

    CN109994128A

  • Intelligent network connection automobile anomaly detection system and method based on big data analysis

    CN119946640A

  • Digital marketing management system and method for telecommunication service

    CN120013614A

  • Automatic conference recording and abstract generating method for intelligent conference

    CN120340497A

  • Computer control system and method based on cloud platform

    CN120371525A

Cited By

  • Intelligent call data analysis and prediction method

    CN121030231A

  • Hearing-aid hearing state detection method and system

    CN121509891A