Time series data preprocessing system and method based on a communication network
By preprocessing the time series of the communication network, identifying and removing noise, and stabilizing data processing between nodes, the problem of signal transmission accuracy and node blockage is solved, and data quality and network efficiency are improved.
Patent Information
- Application Number
- CN202510368692.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-27
AI Technical Summary
Noise is easily mixed in signal transmission in the communication network, resulting in code errors and affecting signal transmission accuracy. The hardware parameters of different nodes lead to node blockage, affecting network operation efficiency.
The time series is preprocessed by signal monitoring module, wave frequency analysis module, standard fitting module and node scheduling module. Through window transformation of variable windows, T-shaped correlation coefficient calculation and node task management, noise is identified and removed to stabilize data processing between nodes.
Improve the data quality and interpretability of the time series, identify long-term trends and periodic patterns, stabilize data processing between nodes, and improve the overall information processing efficiency of the communication network.
Smart Images

Figure CN119906533B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing, and specifically to a time series data preprocessing system and method based on a communication network. Background Art
[0002] A communication network is a system composed of multiple communication nodes and communication links connecting the nodes, and its purpose is to realize the transmission and exchange of data information between different nodes. In a communication network, binary time series are commonly used to transmit signals. A time series is a series of data points arranged in chronological order, which reflects the change of information over time. Processing the time series helps to improve data quality and optimize resource allocation between nodes.
[0003] Since signals are easily mixed with noise during transmission, when digitizing samples, there are error codes in the restored time series signal, which affects the signal transmission accuracy. Since status codes, instruction statements, etc. in the communication network have the same expression form, these error codes can be filtered out through signal analysis. However, due to the high instability of the time series, existing processing technologies are difficult to accurately distinguish error information.
[0004] In addition, each node in the communication network is composed of computers with different specifications, and essentially has different hardware parameters. There are differences in the processing ability and adaptability to time series signals. If the time series is not preprocessed, it is easy to occur node congestion, which affects the operation efficiency of the communication network. Summary of the Invention
[0005] The purpose of the present invention is to provide a time series data preprocessing system and method based on a communication network to solve the problems raised in the above background art.
[0006] To solve the above technical problems, the present invention provides the following technical solution: A time series data preprocessing system based on a communication network, including: a signal monitoring module, a wave frequency analysis module, a standard fitting module, a sequence processing module, and a node scheduling module;
[0007] The signal monitoring module is used to load a monitoring control in the node server of the communication network, obtain the input data and output data of each communication network node, expand the communication data in a time series, set an information processing device or a signal transmission path at the input and output ports of the node to store the expanded time series, and perform arithmetic analysis on the time series;
[0008] The wave frequency analysis module is used to perform windowed transformation with a variable window on the expanded time series, determine the time-frequency window of the time series according to the window interval when the correlation coefficient is the highest, cut the sequence according to the time-frequency window to obtain subsequences, perform dimensionless line transformation on each subsequence and the previous subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence according to the information distribution between the subsequences;
[0009] The standard fitting module is used to calculate the correlation degree of each subsequence with other subsequences according to the T-type correlation coefficient between the subsequences, retain the set of all sequences with correlation degrees less than the threshold to obtain the first reference set, and discard the set with the number of sequences less than the threshold in the first reference set to obtain the second reference set. For all subsequences in each second reference set, perform differential fitting and output the standardized sequence and the observation error interval corresponding to each second reference set;
[0010] The sequence processing module is used to obtain the subsequent input or output sequence, cut the sequence according to the time-frequency window, separate the subsequent subsequences, calculate the correlation degree of the subsequent subsequences with each standardized sequence, and when the correlation degree is within the error interval, perform standardized processing on the subsequent sequence to make the subsequent sequence consistent with the standardized fitting sequence and then output it. Otherwise, repeat the time window cutting and fuse the cutting results into the first reference set;
[0011] The node scheduling module is used to obtain the number of input and output subsequences of each node in the communication network, determine the operation speed of the node according to the ratio of the number of input and output subsequences, and perform task-based node management on the communication network according to the operation speed of the node to make the number of input and output sequences of each node maintain a unified ratio to stabilize the data processing efficiency of the communication network.
[0012] Further, the signal monitoring module includes: a port control unit and a central processing unit;
[0013] The port control unit is arranged in the input and output ports of the node server and is used to monitor the time series of the input and output signals;
[0014] The central processing unit is used to provide computing power support for time series analysis through the terminal data operation chip or the feedback path.
[0015] Further, the wave frequency analysis module includes: a windowing transformation unit, a time-frequency cutting unit, and a composition analysis unit;
[0016] The windowing transformation unit is used to select a window function according to the data information entropy and perform discrete wavelet transformation on the time series based on the window function;
[0017] The time-frequency cutting unit is used to calculate the time delay of the window function when the discrete wavelet transform function takes the maximum value, cut the sequence with a time-frequency window, and obtain subsequences;
[0018] The composition analysis unit is used to calculate the composition difference, composition ratio and T-type correlation coefficient function of each subsequence according to the information distribution among the subsequences.
[0019] Further, the standard fitting module includes: an association grouping unit and a difference fitting unit;
[0020] The association grouping unit is used to calculate the association degree of the T-type correlation coefficients between each subsequence, and group the subsequences according to the calculation results;
[0021] The difference fitting unit is used to screen the grouping results, remove invalid groups and non-significant groups, and obtain a second reference set.
[0022] Further, the sequence processing module includes: a signal filtering unit and a sequence fusion unit;
[0023] The signal filtering unit is used to cut the subsequent subsequences and filter the subsequent subsequences according to the corresponding standard sequence information;
[0024] The sequence fusion unit is used to fuse the filtered subsequent subsequences, restore them to a time series and send them to a communication node for processing.
[0025] Further, the node scheduling module includes: a node testing unit, a task management unit and an efficiency stabilizing unit;
[0026] The node testing unit is used to calculate the information processing speed of the communication node according to the ratio of the information amount of the input subsequence and the output subsequence in the communication node;
[0027] The task management unit is used to plan the best task allocation method according to the data volume of the signal to be transmitted and the information processing speed of each node;
[0028] The efficiency stabilizing unit is used to calculate the overall efficiency of the communication network in real time and upload the information amount of the input and output time series of the communication network in real time.
[0029] A time series data preprocessing method based on a communication network includes the following steps:
[0030] Step S1. Listen to input and output signals at the node server port, expand the signals in the form of a time series, calculate the information entropy of the time series, select a window function according to the information entropy, and perform discrete wavelet transform on the time series with the window function as the base frequency to obtain a discrete wavelet transform function;
[0031] Step S2. Calculate the time delay of the window function when the discrete wavelet transform function takes the maximum value, denoted as the time-frequency window. Use the time-frequency window to cut the time series to obtain subsequences with a fixed symbol length. Perform a dimensionless broken-line transformation on each subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence.
[0032] Step S3. Calculate the T-type correlation coefficients between each subsequence and other subsequences. Classify all subsequences with a correlation degree less than the threshold into one category to obtain the first reference set. Discard the sets with the number of elements less than the threshold to obtain the second reference set.
[0033] Step S4. Superimpose and sample the subsequences in each reference set to obtain a standardized sequence and an error interval. Cut the subsequent time series according to the time-frequency window, and calculate the T-type correlation coefficients between the subsequences of the subsequent time series and each standardized sequence. When the correlation degree is within the error interval, standardize each subsequent subsequence.
[0034] Step S5. Represent the information processing speed of the communication node by the information quantity ratio of the input subsequence and the output subsequence in the communication node. Perform task allocation according to the data volume of the signal to be transmitted and the information processing speed of each node to maximize the overall efficiency of the communication network.
[0035] Further, Step S1 includes:
[0036] Step S11. Load a listening control in the node server of the communication network to listen to the input and output signals, and transmit the input and output signals into the processing chip through the terminal data operation chip or the feedback path.
[0037] Step S12. Expand the listened input and output signals into a time series in the processing chip. The time series is a finite-length sequence, and all elements in the sequence are binary. Calculate the information entropy according to the length of the time series, and select a window function from the pre-loaded function library according to the information entropy.
[0038] Step S13. Perform a discrete wavelet transform on the time series with the window function as the base frequency. The transform function is:
[0039] ;
[0040] where f(a, b) is the discrete wavelet transform function, a is the scale parameter, b is the time delay parameter, t0 is the sampling interval of the time series, δ(k) is the window function, E(k) is the value of the k-th element in the time series, N is the number of elements in the time series, and k is the sequence label.
[0041] Further, Step S2 includes:
[0042] Step S21. Limit the scale parameter within the range of one period length of the window function, calculate the maximum value of the discrete wavelet transform function, and output the magnitude of the delay parameter b when the function takes the maximum value, denoted as the time-frequency window.
[0043] Step S22. Use the time-frequency window to cut the time series so that each subsequence contains m elements. If there is a sequence with a length less than m, it is filled with a low-level signal. The subsequences are numbered in chronological order, and the previous sequences of each subsequence are determined in descending order of the numbers. The initial subsequence takes the last subsequence as the previous sequence.
[0044] Perform a dimensionless line transformation on each subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence:
[0045] ;
[0046] Among them, z represents the composition difference between the subsequence and the previous sequence, z1(x) represents the difference between the x-th element and the (x + 1)-th element in the previous subsequence, z2(x) represents the difference between the x-th element and the (x + 1)-th element in the current subsequence, s represents the composition ratio between the subsequence and the previous sequence, min and max respectively represent the maximum value and minimum value functions, and r represents the T-type correlation coefficient.
[0047] Further, step S3 includes:
[0048] Step S31. Calculate the T-type correlation coefficients between all subsequences and other sequences, group the subsequences so that the correlation coefficients between every two sequences within the group are all less than the threshold, and store all the groups in a set to obtain the first reference set.
[0049] Step S32. Screen out all groups in the first reference set that contain a number of sequences less than the threshold to obtain the second reference set.
[0050] Further, step S4 includes:
[0051] Step S41. For each group in the second reference set, calculate the average value of each element in all the subsequences within the group. If the average value is greater than the sampling standard, it is compiled into a high-level signal. If the average value is less than the sampling standard, it is compiled into a low-level signal. Obtain all the compiled signals and combine them into a standardized sequence, and the error interval is [-Σ (x=1,m) |z(x) - z0(x)|, Σ (x=1,m) |z(x) - z0(x)|], where z(x) represents the x-th element in the standardized sequence, and z0(x) represents the average value of the x-th element of all the subsequences within the group.
[0052] Step S42. When transmitting the subsequent time series in the communication node, the subsequent time series is cut according to the time-frequency window to obtain subsequences of the subsequent time series. Calculate the T-type correlation coefficients between the subsequences of each subsequent time series and the standardized sequence. If the correlation degree is within the error interval, the standardized sequence is used to replace the subsequence of the subsequent time series, and then it is restored and output.
[0053] Further, step S5 includes:
[0054] Step S51. Determine the information processing speed of the communication node according to the ratio of the information amount of the input subsequence and the output subsequence in the communication node, and identify the speed of each node in the communication network.
[0055] Step S52. Calculate the information amount of the data to be processed, and select a scheduling algorithm to allocate traffic among the nodes. The scheduling algorithms include: dynamic weighted scheduling, hybrid scheduling, neural network scheduling, and batch processing scheduling.
[0056] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0057] The present invention unfolds communication data in time series, adopts windowing transformation with a variable window to determine the time-frequency window of the time series, cuts the sequence according to the time-frequency window, performs dimensionless broken-line transformation on each sequence and the previous sequence, and calculates the T-type correlation coefficient of the technical sequence, which can identify long-term trends, periodic changes, and seasonal patterns in the data, can more accurately predict the cyclic pattern of the time series, and improve the integrity and quality of the data.
[0058] The present invention calculates the correlation degree according to the T-type correlation coefficient, retains the set of all sequences with a correlation degree less than the threshold. When the subsequent input is within the error interval of the standardized sequence, the subsequent sequence is pre-processed by standardization and then output. The abnormal noise in the data is identified and removed through the smoothing technology, making the time series clearer and more interpretable.
[0059] The present invention determines the operation speed of each node according to the ratio of the input and output standardized sequence numbers of each node, and performs task-based node management on the communication network, so that the input and output sequence numbers of each node maintain a unified ratio, to stabilize the data processing efficiency of the communication network, better understand the dynamic information processing state of the communication network, and improve the overall information processing efficiency of the communication network. Description of the Drawings
[0060] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0061] Figure 1It is a schematic structural diagram of the time series data preprocessing system based on the communication network of the present invention;
[0062] Figure 2 It is a schematic step diagram of the time series data preprocessing method based on the communication network of the present invention. Detailed implementation manners
[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] Please refer to Figure 1 , the present invention provides a technical solution: a time series data preprocessing system based on a communication network, including: a signal monitoring module, a wave frequency analysis module, a standard fitting module, a sequence processing module, and a node scheduling module;
[0065] The signal monitoring module is used to load a monitoring control in the node server of the communication network, obtain the input data and output data of each communication network node, expand the communication data in a time series, set an information processing device or a signal transmission path at the input and output ports of the node to store the expanded time series, and perform arithmetic analysis on the time series;
[0066] The signal monitoring module includes: a port control unit and a central processing unit;
[0067] The port control unit is set in the input and output ports of the node server and is used to monitor the time series of the input and output signals;
[0068] The central processing unit is used to provide computing power support for time series analysis through a terminal data operation chip or a feedback path.
[0069] The wave frequency analysis module is used to perform windowing transformation on the expanded time series by using a variable window, determine the time-frequency window of the time series according to the window interval when the correlation coefficient is the highest, cut the sequence according to the time-frequency window to obtain subsequences, perform dimensionless broken line transformation on each subsequence and the previous subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence according to the information distribution between the subsequences;
[0070] The wave frequency analysis module includes: a windowing transformation unit, a time-frequency cutting unit, and a composition analysis unit;
[0071] The windowing transformation unit is used to select a window function according to the data information entropy and perform discrete wavelet transformation on the time series based on the window function;
[0072] The time-frequency cutting unit is used to calculate the time delay of the window function when the discrete wavelet transform function takes the maximum value, cut the sequence with a time-frequency window, and obtain subsequences;
[0073] The composition analysis unit is used to calculate the composition difference, composition ratio and T-type correlation coefficient function of each subsequence according to the information distribution among the subsequences.
[0074] The standard fitting module is used to calculate the correlation degree between each subsequence and other subsequences according to the T-type correlation coefficient between the subsequences, retain the set of all sequences with a correlation degree less than the threshold to obtain the first reference set, and in the first reference set, discard the set with the number of sequences less than the threshold to obtain the second reference set. For all subsequences in each second reference set, perform differential fitting, and output the standardized sequence and the observation error interval corresponding to each second reference set;
[0075] The standard fitting module includes: a correlation grouping unit and a differential fitting unit;
[0076] The correlation grouping unit is used to calculate the correlation degree of the T-type correlation coefficient between each subsequence, and group the subsequences according to the calculation results;
[0077] The differential fitting unit is used to screen the grouping results, remove invalid groups and non-significant groups, and obtain the second reference set.
[0078] The sequence processing module is used to obtain subsequent input or output sequences, cut the sequences with a time-frequency window, separate subsequent subsequences, calculate the correlation degree between the subsequent subsequences and each standardized sequence. When the correlation degree is within the error interval, perform standardization processing on the subsequent sequences, make the subsequent sequences consistent with the standardized fitting sequences and then output them. Otherwise, repeat the time window cutting and fuse the cutting results into the first reference set;
[0079] The sequence processing module includes: a signal filtering unit and a sequence fusion unit;
[0080] The signal filtering unit is used to cut subsequent subsequences and filter the subsequent subsequences according to the corresponding standard sequence information;
[0081] The sequence fusion unit is used to fuse the filtered subsequent subsequences, restore them to a time series and send them to a communication node for processing.
[0082] The node scheduling module is used to obtain the number of input and output subsequences of each node in the communication network, determine the operation speed of the node according to the ratio of the number of input and output subsequences, perform task-based node management on the communication network according to the operation speed of the node, so that the number of input and output sequences of each node maintains a unified ratio to stabilize the data processing efficiency of the communication network.
[0083] The node scheduling module includes: a node testing unit, a task management unit, and an efficiency stabilizing unit;
[0084] The node testing unit is used to calculate the information processing speed of a communication node according to the ratio of the information amounts of the input subsequence and the output subsequence in the communication node;
[0085] The task management unit is used to plan the optimal task allocation method according to the data volume of the signal to be transmitted and the information processing speeds of each node;
[0086] The efficiency stabilizing unit is used to calculate the overall efficiency of the communication network in real time and upload the information amounts of the input-output time series of the communication network in real time.
[0087] As Figure 2 shown, a time series data preprocessing method based on a communication network includes the following steps:
[0088] Step S1. Listen for input and output signals at the node server port, expand the signals in the form of a time series, calculate the information entropy of the time series, select a window function according to the information entropy, and perform a discrete wavelet transform on the time series with the window function as the base frequency to obtain a discrete wavelet transform function;
[0089] Step S1 includes:
[0090] Step S11. Load a listening control in the node server of the communication network, listen for input and output signals, and transmit the input and output signals into the processing chip through a terminal data operation chip or a feedback path;
[0091] Step S12. Expand the listened input and output signals into a time series in the processing chip. The time series is a finite-length sequence, and the elements in the sequence are all binary. Calculate the information entropy according to the length of the time series, and select a window function from a pre-loaded function library according to the information entropy;
[0092] Step S13. Perform a discrete wavelet transform on the time series with the window function as the base frequency. The transform function is:
[0093] ;
[0094] where f(a, b) is the discrete wavelet transform function, a is the scale parameter, b is the delay parameter, t0 is the sampling interval of the time series, δ(k) is the window function, E(k) is the value of the k-th element in the time series, N is the number of elements in the time series, and k is the sequence label.
[0095] Step S2. Calculate the time delay of the window function when the discrete wavelet transform function takes the maximum value, denoted as the time-frequency window. Use the time-frequency window to cut the time series to obtain subsequences with a fixed symbol length. Perform a dimensionless polyline transformation on each subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence;
[0096] Step S2 includes:
[0097] Step S21. Limit the scale parameter within one period length of the window function, calculate the maximum value of the discrete wavelet transform function, and output the magnitude of the time delay parameter b when the function takes the maximum value, denoted as the time-frequency window;
[0098] Step S22. Use the time-frequency window to cut the time series so that each subsequence contains m elements. If there is a sequence with a length less than m, it is filled with a low-level signal. Number the subsequences in chronological order, determine the previous sequences of each subsequence in descending order of the numbers, and use the last subsequence as the previous sequence for the initial subsequence;
[0099] Perform a dimensionless polyline transformation on each subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence:
[0100] ;
[0101] Among them, z represents the composition difference between the subsequence and the previous sequence, z1(x) represents the difference between the x-th element and the (x + 1)-th element in the previous subsequence, z2(x) represents the difference between the x-th element and the (x + 1)-th element in the current subsequence, s represents the composition ratio between the subsequence and the previous sequence, min and max respectively represent the maximum value and minimum value functions, and r represents the T-type correlation coefficient.
[0102] Step S3. Calculate the T-type correlation coefficients between each subsequence and other subsequences, classify all subsequences with a correlation degree less than the threshold into one category to obtain the first reference set, and discard the sets with the number of elements less than the threshold in it to obtain the second reference set;
[0103] Step S3 includes:
[0104] Step S31. Calculate the T-type correlation coefficients between all subsequences and other sequences, group the subsequences so that the correlation coefficients between every two sequences within the group are all less than the threshold, and store all the groups in a set to obtain the first reference set;
[0105] Step S32. Screen out all groups in the first reference set that contain a number of sequences less than the threshold to obtain the second reference set.
[0106] Step S4. Superpose and sample the subsequences in each reference set to obtain a standardized sequence and an error interval. Cut the subsequent time series according to the time-frequency window, and calculate the T-type correlation coefficient between the subsequences of the subsequent time series and each standardized sequence. When the correlation degree is within the error interval, standardize each subsequent subsequence;
[0107] Step S4 includes:
[0108] Step S41. For each group in the second reference set, calculate the average value of each element in all subsequences within the group. If the average value is greater than the sampling standard, it is compiled into a high-level signal; if the average value is less than the sampling standard, it is compiled into a low-level signal. Obtain all the compiled signals and combine them into a standardized sequence. The error interval is [-Σ (x=1,m) |z(x) - z0(x)|, Σ (x=1,m) |z(x) - z0(x)|], where z(x) represents the x-th element in the standardized sequence, and z0(x) represents the average value of the x-th element of all subsequences within the group;
[0109] Step S42. When transmitting the subsequent time series in the communication node, cut the subsequent time series according to the time-frequency window to obtain the subsequences of the subsequent time series. Calculate the T-type correlation coefficient between the subsequences of each subsequent time series and the standardized sequence. If the correlation degree is within the error interval, replace the subsequence of the subsequent time series with the standardized sequence and output it after restoration.
[0110] Step S5. Represent the information processing speed of the communication node by the information amount ratio of the input subsequence and the output subsequence in the communication node, and perform task allocation according to the data amount of the signal to be transmitted and the information processing speed of each node to maximize the overall efficiency of the communication network.
[0111] Step S5 includes:
[0112] Step S51. Determine the information processing speed of the communication node according to the ratio of the information amounts of the input subsequence and the output subsequence in the communication node, and identify the speeds of each node in the communication network;
[0113] Step S52. Calculate the information amount of the data to be processed, and select a scheduling algorithm to allocate traffic among the nodes. The scheduling algorithms include: dynamic weighted scheduling, hybrid scheduling, neural network scheduling, and batch processing scheduling.
[0114] Example: The time series input to the communication node is [10110110]. After windowing transformation, the time domain window is 3, so the sequence is cut into three subsequences:
[101] ,
[101] , and
[100] . Group
[101] and
[101] into one group, and
[100] into another group. Filter out the group of
[100] . When the subsequent signal input is
[111] , normalize the subsequent signal to
[101] and output it.
[0115] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0116] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for preprocessing time series data based on a communication network, characterized in that, The method includes the following steps: Step S1. Listen for input and output signals at the node server port, expand the signals in the form of a time series, calculate the information entropy of the time series, select a window function according to the information entropy, and perform a discrete wavelet transform on the time series with the window function as the base frequency to obtain a discrete wavelet transform function; Step S2. Calculate the time delay of the window function when the discrete wavelet transform function takes the maximum value, denoted as the time-frequency window, cut the time series with the time-frequency window to obtain subsequences with a fixed symbol length, perform a dimensionless polyline transform on each subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence; Step S3. Calculate the T-type correlation coefficients between each subsequence and other subsequences, classify all subsequences with a correlation degree less than the threshold into one category to obtain a first reference set, and discard the sets with an element number less than the threshold to obtain a second reference set; Step S4. Superimpose and sample the subsequences in the second reference set to obtain a standardized sequence and an error interval, cut the subsequent time series with the time-frequency window, calculate the T-type correlation coefficients between the subsequences of the subsequent time series and each standardized sequence, and when the correlation degree is within the error interval, perform a standardized process on each subsequent subsequence; Step S5. Represent the information processing speed of the communication node by the information amount ratio of the input subsequence and the output subsequence in the communication node, and perform task allocation according to the data volume of the signal to be transmitted and the information processing speed of each node.
2. The time series data preprocessing method based on a communication network according to claim 1, wherein: Step S1 includes: Step S11. Load a listening control in the node server of the communication network, listen for input and output signals, and transmit the input and output signals into the processing chip through the terminal data operation chip or the feedback path; Step S12. Expand the monitored input and output signals into a time series in the processing chip. The time series is a finite-length sequence, and the elements in the sequence are all binary. Calculate the information entropy according to the length of the time series, and select a window function from the pre-loaded function library according to the information entropy; Step S13. Perform a discrete wavelet transform on the time series with the window function as the base frequency. The discrete wavelet transform function is: ; where f(a, b) is the discrete wavelet transform function, a is the scale parameter, b is the time delay parameter, t0 is the sampling interval of the time series, δ(k) is the window function, E(k) is the kth element value in the time series, and N is the number of elements in the time series.
3. The method for preprocessing time series data based on a communication network according to claim 2, wherein: Step S2 includes: Step S21. Limit the scale parameter within one period length range of the window function, calculate the maximum value of the discrete wavelet transform function, and output the magnitude of the time delay parameter b when the function takes the maximum value, denoted as the time-frequency window; Step S22. Cut the time series with the time-frequency window so that each subsequence contains m elements. If there is a sequence with a length less than m, it is filled with a low-level signal. Number the subsequences in chronological order, determine the pre-sequence of each subsequence in descending order of the number, and the initial subsequence uses the last subsequence as the pre-sequence; Perform a dimensionless polyline transform on each subsequence, and calculate the composition difference, composition ratio, and T-type correlation coefficient of each subsequence: ; Among them, z represents the compositional difference between the subsequence and the pre-order sequence, z1(x) represents the difference between the x-th element and the (x + 1)-th element in the pre-order subsequence, z2(x) represents the difference between the x-th element and the (x + 1)-th element in the current subsequence, s represents the compositional ratio between the subsequence and the pre-order sequence, min and max respectively represent the functions of taking the maximum value and the minimum value, and r represents the T-type correlation coefficient.
4. The time series data preprocessing method based on a communication network according to claim 3, wherein: Step S3 includes: Step S31. Calculate the T-type correlation coefficients between all subsequences and other sequences, group the subsequences so that the correlation coefficients between every two sequences within the group are all less than the threshold, store all the groups in a set to obtain the first reference set; Step S32. Screen out the groups in the first reference set that contain a number of sequences less than the threshold to obtain the second reference set; Step S4 includes: Step S41. For each group in the second reference set, calculate the average value of each element in all subsequences within the group. If the average value is greater than the sampling standard, it is compiled into a high-level signal; if the average value is less than the sampling standard, it is compiled into a low-level signal. Obtain all the compiled signals and combine them into a standardized sequence, with an error range of [-Σ (x=1,m) |z(x) - z0(x)|, Σ (x=1,m) |z(x) - z0(x)|], where z(x) represents the x-th element in the standardized sequence, and z0(x) represents the average value of the x-th element in all subsequences within the group; Step S42. When transmitting the post-order time series in the communication node, cut the post-order time series according to the time-frequency window to obtain the subsequences of the post-order time series, calculate the T-type correlation coefficients between the subsequences of each post-order time series and the standardized sequence. If the degree of association is within the error interval, replace the subsequences of the post-order time series with the standardized sequence, restore and output it. Otherwise, repeat the time-frequency window cutting and fuse the cutting results into the first reference set.
5. The time series data preprocessing method based on a communication network according to claim 4, wherein: Step S5 includes: Step S51. Determine the information processing speed of the communication node according to the ratio of the information amount of the input subsequence and the output subsequence in the communication node, and identify the speeds of each node in the communication network; Step S52. Calculate the information amount of the data to be processed, and select a scheduling algorithm to allocate traffic among the nodes. The scheduling algorithms include: dynamic weighted scheduling, hybrid scheduling, neural network scheduling, and batch processing scheduling.
6. A time series data preprocessing system based on a communication network, the system performing the time series data preprocessing method based on a communication network according to claim 1, characterized in that, The system includes the following modules: a signal monitoring module, a wave frequency analysis module, a standard fitting module, a sequence processing module, and a node scheduling module; The signal monitoring module is used to load a monitoring control in the node server of the communication network, obtain the input data and output data of each communication network node, expand the communication data in a time series, set an information processing device or a signal transmission path at the input and output ports of the node to store the expanded time series, and perform arithmetic analysis on the time series; The wave frequency analysis module is used to perform windowing transformation on the expanded time series using a variable window, determine the time-frequency window of the time series, cut the sequence according to the time-frequency window to obtain subsequences, perform dimensionless polyline transformation on each subsequence and the pre-order subsequence, and calculate the compositional difference, compositional ratio, and T-type correlation coefficient of each subsequence according to the information distribution between the subsequences; The standard fitting module is used to calculate the degree of association between each subsequence and other subsequences according to the T-type correlation coefficients between the subsequences, retain the set of all sequences with degrees of association less than the threshold to obtain the first reference set. In the first reference set, discard the sets with a number of sequences less than the threshold to obtain the second reference set, and output the standardized sequence and the observed error interval corresponding to each second reference set; The sequence processing module is used to obtain the subsequent input or output sequence, cut the sequence according to the time-frequency window, separate the subsequent subsequences, calculate the correlation degree between the subsequent subsequences and each standardized sequence, and when the correlation degree is within the error range, perform standardized processing on the subsequent sequence to make the subsequent sequence consistent with the standardized fitting sequence and then output; The node scheduling module is used to obtain the number of input and output subsequences of each node in the communication network and determine the operation speed of the node according to the ratio of the number of input and output subsequences.
7. The time series data preprocessing system based on a communication network according to claim 6, wherein: The signal monitoring module includes: a port control unit and a central processing unit; The port control unit is arranged in the input and output ports of the node server and is used to monitor the time series of input and output signals; The central processing unit is used to provide computing power support for time series analysis through the terminal data operation chip or the feedback path.
8. The time series data preprocessing system based on a communication network according to claim 7, wherein: The wave frequency analysis module includes: a windowing transformation unit, a time-frequency cutting unit, and a composition analysis unit; The windowing transformation unit is used to select a window function according to the data information entropy and perform discrete wavelet transformation on the time series based on the window function; The time-frequency cutting unit is used to calculate the time delay of the window function when the discrete wavelet transformation function takes the maximum value and cut the sequence with the time-frequency window to obtain subsequences; The composition analysis unit is used to calculate the composition difference, composition ratio, and T-type correlation coefficient function of each subsequence according to the information distribution between the subsequences.
9. The time series data preprocessing system based on a communication network according to claim 8, wherein: The standard fitting module includes: a correlation grouping unit and a difference fitting unit; The correlation grouping unit is used to calculate the correlation degree of the T-type correlation coefficients between each subsequence and group the subsequences according to the calculation results; The difference fitting unit is used to screen the grouping results, remove invalid groups and non-significant groups, and obtain a second reference set; The sequence processing module includes: a signal filtering unit and a sequence fusion unit; The signal filtering unit is used to cut the subsequent subsequences and filter the subsequent subsequences according to the corresponding standard sequence information; The sequence fusion unit is used to fuse the filtered subsequent subsequences, restore them to a time series, and send them to the communication node for processing.
10. The time series data preprocessing system based on a communication network according to claim 9, wherein: The node scheduling module includes: a node testing unit, a task management unit, and an efficiency stabilization unit; The node testing unit is used to calculate the information processing speed of the communication node according to the ratio of the information amounts of the input subsequence and the output subsequence in the communication node; The task management unit is used to plan the optimal task allocation method according to the data volume of the signal to be transmitted and the information processing speed of each node; The efficiency stabilization unit is used to calculate the overall efficiency of the communication network in real time and upload the information amount of the input and output time series of the communication network in real time.
Citation Information
Patent Citations
Wavelet transform and Markov model combined time series anomaly detection method
CN116975760A
Pediatric abnormal breathing detection method and system based on physiological signals
CN119538119A