Data processing method, data screening method and data compression method

By dynamically generating the window data information range, only representative data in massive sensing data is stored, which solves the problem of low storage and query efficiency of massive data, and realizes the streamlining of data volume and retention of features.

CN119669537BActive Publication Date: 2025-05-16ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510202323.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-16
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In the scenario of massive sensing data, direct centralized storage and analysis of original collected data leads to excessive bandwidth pressure, huge storage costs and low data query efficiency.

Method used

By obtaining the reference time sequence data information and its corresponding time points, the data time difference is calculated, and the window data information range is generated based on the time difference and the starting data information, the data processing window is dynamically adjusted, and only the representative data information is stored in the window.

Benefits of technology

It effectively reduces the amount of data, reduces storage and bandwidth requirements, while retaining the key features of the data, improving data query efficiency, and avoiding the risk of data forgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669537B_ABST
    Figure CN119669537B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a data processing method, a data screening method, and a data compression method, wherein the data processing method includes: obtaining reference time series data information and a corresponding reference time point, and obtaining a reference data processing window corresponding to the reference starting time series data information and the reference starting time point. The reference data time difference is calculated based on the reference time point and the reference starting time point, and a window data information range is generated according to the time difference and the reference starting time series data information. When the reference time series data information falls within this range and the time difference does not exceed the preset threshold, the reference time series data information continues to be obtained; otherwise, the target time series data information is obtained and stored based on the reference time point and the reference time series data information is used as the reference starting time series data information of the next reference data processing window. The window data information range is generated by the time difference, and the key information of the data is retained during data processing, thereby ensuring the accuracy of the data while reducing the amount of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present specification relate to the field of computer technology, and more particularly to a data processing method, a data screening method, and a data compression method. Background Art

[0002] With the continuous development of the Internet of Things and industrial intelligent technology, the access and analysis of massive sensor data has gradually become possible, bringing significant results such as real-time monitoring and accurate decision-making. However, in ultra-large-scale data scenarios, the high-frequency collection and storage of large amounts of sensor information will lead to a rapid increase in bandwidth and storage costs, bringing challenges to subsequent data query and analysis.

[0003] Currently, the method of directly storing and analyzing all the original collected data in a centralized manner can meet the needs of device access, data collection and some intelligent analysis. However, when the amount of data reaches the tens of billions, it is easy to have problems such as excessive bandwidth pressure, huge storage costs and low data query efficiency. Therefore, in order to solve the above problems, a data processing method is needed. Summary of the invention

[0004] In view of this, the embodiments of this specification provide a data processing method, a data screening method, and a data compression method. One or more embodiments of this specification also relate to a data processing device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a data processing method is provided, including:

[0006] Acquire reference timing data information and a reference time point corresponding to the reference timing data information, and acquire a reference data processing window corresponding to the reference timing data information, and reference starting timing data information and a reference starting time point corresponding to the reference data processing window;

[0007] Determine a reference data time difference based on the reference time point and the reference starting time point;

[0008] Generate a window data information range according to the reference data time difference and the reference starting timing data information;

[0009] When the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to perform the step of obtaining the reference time series data information and the reference time point corresponding to the reference time series data information;

[0010] In the case that the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target timing data information is acquired based on the reference time point, and the target timing data information and the reference starting timing data information are stored; the reference timing data information is determined to be the reference starting timing data information of the next reference data processing window, and the steps of acquiring the reference timing data information and the reference time point corresponding to the reference timing data information are continued, wherein the target timing data information is representative data information in the reference data processing window.

[0011] According to a second aspect of an embodiment of this specification, a data screening method is provided, comprising:

[0012] Receive reference timing data information and a reference time point corresponding to the reference timing data information, and obtain a reference data processing window corresponding to the reference timing data information, and reference starting timing data information and a reference starting time point corresponding to the reference data processing window;

[0013] Determine a reference data time difference based on the reference time point and the reference starting time point;

[0014] Generate a window data information range according to the reference data time difference and the reference starting timing data information;

[0015] When the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to perform the step of obtaining the reference time series data information and the reference time point corresponding to the reference time series data information;

[0016] In the case that the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target timing data information is acquired based on the reference time point, and the target timing data information and the reference starting timing data information are stored; the reference timing data information is determined to be the reference starting timing data information of the next reference data processing window, and the steps of receiving the reference timing data information and the reference time point corresponding to the reference timing data information are continued, wherein the target timing data information is representative data information in the reference data processing window.

[0017] According to a third aspect of an embodiment of this specification, a data compression method is provided, including:

[0018] Acquire reference time series data information and a reference time point corresponding to the reference time series data information in the time series data set to be compressed, and acquire a reference data processing window corresponding to the reference time series data information, and reference starting time series data information and a reference starting time point corresponding to the reference data processing window;

[0019] Determine a reference data time difference based on the reference time point and the reference starting time point;

[0020] Generate a window data information range according to the reference data time difference and the reference starting timing data information;

[0021] When the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to perform the step of obtaining the reference time series data information and the reference time point corresponding to the reference time series data information;

[0022] In the case that the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target timing data information is obtained based on the reference time point, and the target timing data information and the reference starting timing data information are stored in a target timing data set; the reference timing data information is determined to be the reference starting timing data information of the next reference data processing window, and the steps of receiving the reference timing data information and the reference time point corresponding to the reference timing data information are continued until all the reference timing data information in the timing data set to be compressed are processed, and a data compression result is generated based on the target timing data set, wherein the target timing data information is representative data information in the reference data processing window.

[0023] According to a fourth aspect of an embodiment of this specification, a computing device is provided, including:

[0024] Memory and processor;

[0025] The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the above-mentioned data processing method, data screening method and data compression method are implemented.

[0026] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data processing method, data screening method and data compression method are implemented.

[0027] According to the sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which implements the steps of the above-mentioned data processing method, data screening method and data compression method when executed by a processor.

[0028] By applying the scheme of the embodiment of this specification, by obtaining the reference time series data information and the reference time point, and calculating the reference data time difference based on the reference time point and the reference starting time point, it can be judged whether the current data falls within the generated window data information range. When the data is within the range and the time difference does not exceed the preset threshold, the new reference time series data information can continue to be included in the same window to make full use of the stable time period, thereby reducing redundant sampling points. When the data is not within this range or the time difference is too large, it is necessary to end the current window and open the next window so that the subsequent time series data can be managed and recorded separately. With the help of this window data information range dynamically generated based on the time difference and reference information, the method can flexibly respond to different types of local changes: aggregate more data in the stable stage to reduce storage and bandwidth requirements, and open a new window in time when fluctuations or mutations occur to retain important features. In addition, while greatly reducing the amount of data, it can still effectively retain core time series features such as mutations and trends. All output sampling points come from the original data itself, avoiding the potential risk of data forgery, thereby achieving a higher query efficiency in subsequent data queries through the streamlined data. Data simplification through real-time adjustment of windows is also highly engineering-friendly and can adapt to common IoT system architectures to achieve efficient management and analysis of multi-source, multi-level time series data. This not only reduces the bandwidth and storage pressure brought by massive data, but also provides a reliable foundation for subsequent real-time monitoring, intelligent diagnosis, and deeper data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flow chart of a data processing method provided by an embodiment of this specification;

[0030] Figure 2 is a schematic diagram of initial data information and reference data information provided by an embodiment of this specification;

[0031] Figure 3 is a schematic diagram of a data processing window provided by an embodiment of this specification;

[0032] Figure 4 It is a schematic diagram of reference data information and stored data information provided by an embodiment of this specification;

[0033] Figure 5 is a flow chart of a data screening method provided by an embodiment of this specification;

[0034] Figure 6 is a flow chart of a data compression method provided by an embodiment of this specification;

[0035] Figure 7 is an architecture diagram of a data screening system provided by an embodiment of this specification;

[0036] Figure 8 is an architecture diagram of a data compression system provided by an embodiment of this specification;

[0037] Fig. 9 is a process flow chart of a network performance data screening method provided by an embodiment of this specification;

[0038] Fig.10 is a structural schematic diagram of a data processing device provided by an embodiment of this specification;

[0039] Fig.11 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0040] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0041] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0042] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0043] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0044] In this specification, a data processing method, a data screening method, and a data compression method are provided. One or more embodiments of this specification also relate to a data processing device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0045] With the widespread application of the Internet of Things and the Industrial Internet of Things, many platforms face technical challenges in processing massive amounts of data, especially in terms of device access, data collection, and intelligent analysis. For medium and large customers, the huge amount of sensor data usually leads to problems such as excessive data bandwidth and excessive storage costs, which in turn limits the effective use of data and subsequent analysis applications, making it difficult to effectively query and use sensor data. In order to solve these problems, reducing the time series density of data to control storage and query costs has become a necessary measure, and downsampling processing has become one of the solutions.

[0046] In order to ensure that downsampling can meet the needs of actual applications, three core requirements need to be followed: first, effectively reduce data density and the total amount of data; second, the downsampling process should retain important data variation characteristics, such as mutations and trend changes, to ensure that subsequent applications are not affected; then, the authenticity of the data must be ensured to avoid generating false data points during the downsampling process.

[0047] See also Figure 1 , Figure 1 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0048] Step 102: Obtain reference timing data information and a reference time point corresponding to the reference timing data information, and obtain a reference data processing window corresponding to the reference timing data information, as well as reference starting timing data information and a reference starting time point corresponding to the reference data processing window.

[0049] In practical applications, the reference time series data information is the time series data point; the reference time point is the time tag corresponding to the time series data point; the reference data processing window is the time window defined in the data processing process, which is used to define the processing range; the reference starting time series data information is the time series data at the starting moment in the reference data processing window; the reference starting time point is the starting time of the reference data processing window.

[0050] Reference time series data information can be understood as time series data that has been collected and prepared for processing during the data processing process. For example, in a performance monitoring system of a network cluster, the reference time series data information may be the collected network bandwidth usage data, which indicates the specific usage of network bandwidth at a certain point in time. These data can be denoised, compressed or other pre-processing steps to ensure that they meet the requirements of subsequent processing and are used as input for subsequent window judgment, data collection and analysis tasks.

[0051] The reference time point can be understood as a time tag associated with the reference time series data information, which marks the specific moment of data collection. In the scenario of network cluster monitoring, the reference time point can be an accurate timestamp of a certain moment. For example, the time point for recording network bandwidth usage may be "2025-01-18 10:30:00". This time point is not only used to calculate the time difference, but also provides a time reference for subsequent window judgment and data collection, ensuring that data is processed within the correct time range.

[0052] The reference data processing window can be understood as a time range window set during the data processing process for efficient management and analysis of time series data. For example, for the performance data of a network cluster, the reference data processing window can be the network bandwidth data collected every hour. The data collected during this time period will be processed together for statistical analysis. The size and range of the reference data processing window directly affect the accuracy and efficiency of data processing. For bandwidth utilization analysis, hours or minutes are usually selected as the window length.

[0053] The reference start time series data information can be understood as the data point that serves as the start of each data processing window. In the example of network cluster performance monitoring, the reference start time series data information can be the network bandwidth usage data at the beginning of each hour, usually the network bandwidth usage value at that moment. This information is used to determine whether the data meets the processing range of the current window and is the basis for subsequent data collection and processing.

[0054] The reference starting time point can be understood as the starting moment of the reference data processing window, which is used to mark the time limit of the window. In the example of network cluster performance monitoring, the reference starting time point can be the start time of each hour, such as "2025-01-18 10:00:00", which marks the start of the current data window. All data collected after this time point will be included in this window for processing. It should be noted that if the reference time series data information is not obtained before the reference time series data information, the reference time series data information is automatically set to the reference starting time series data information of the first reference data processing window, and its corresponding reference time point is set to the reference starting time point of the first reference data processing window.

[0055] It should be noted that obtaining reference timing data information and a reference time point corresponding to the reference timing data information can be understood as obtaining the data to be processed. The specific method can be to obtain the reference timing data information obtained by the data acquisition device in real time, and use the time point of acquiring the reference timing data information as the reference time point; or it can be to determine the reference timing data information in a reference timing data information set including multiple data information and obtain the reference time point corresponding to the reference timing data information, etc. This specification does not impose any restrictions on this.

[0056] The collection of reference time series data information and the acquisition of the corresponding reference data processing window provide the basis for subsequent data processing and are important steps in the entire data processing process.

[0057] Further, obtaining reference time series data information and a reference time point corresponding to the reference time series data information includes:

[0058] Acquire initial time series data information and an initial time point corresponding to the initial time series data information;

[0059] Generate a reference time series data information set based on the initial time series data information and an initial time point corresponding to the initial time series data information;

[0060] Reference time series data information and a reference time point corresponding to the reference time series data information are determined in the reference time series data information set.

[0061] In practical applications, the initial time series data information is the original data obtained by the data acquisition module from the source device; the initial time point is the acquisition time tag corresponding to the initial time series data information, which is used to mark the specific moment of data acquisition; the reference time series data information set is the time series data set generated by processing the initial time series data information and the initial time point for subsequent screening and analysis. The close combination of these information supports the complete process of data screening, downsampling and feature retention.

[0062] Initial time series data information can be understood as the unprocessed raw data points directly obtained from the data acquisition system. For example, in industrial equipment monitoring, initial time series data information may include the temperature, pressure or vibration data of the equipment. As the starting point of the entire data processing process, the integrity and accuracy of these data directly affect the subsequent processing effect. Since they are unprocessed, the initial time series data information may contain noise or outliers, so they need to be optimized through subsequent denoising and screening steps.

[0063] The initial time point can be understood as the time mark of the initial time series data information, which is used to mark the specific collection time of each data point. For example, in network performance monitoring, the initial time point may be "2025-01-18 10:30:00", which represents the network bandwidth value recorded at that time point. The role of the initial time point is not only to provide a reference on the timeline, but also to provide a basic basis for subsequent time difference calculation and window division.

[0064] The reference time series data information set can be understood as a more refined and reliable time series data set generated by the initial time series data information and the initial time point. For example, by denoising and filtering the initial time series data information, a set containing only valid data is obtained. In industrial scenarios, the reference time series data information set may include key data points of the equipment operating status (such as fluctuation data within the normal operating temperature range) for subsequent feature extraction and data compression.

[0065] It should be noted that determining the reference timing data information and the reference time point corresponding to the reference timing data information in the reference timing data information set can be understood as obtaining the data information in the data information set. Considering that the data information stored in the reference timing data information set has timing characteristics, the method for determining the reference timing data information and the reference time point corresponding to the reference timing data information in the reference timing data information set can be to obtain the reference timing data information in sequence based on the order of each reference time point and obtain the reference time point corresponding to the reference timing data information.

[0066] refer to Figure 2 , Figure 2A schematic diagram of initial data information and reference data information provided for an embodiment of the present specification, wherein green dots represent initial data information, which are raw data directly obtained from the acquisition device and contain real-time records of the device's operating status, but may contain certain fluctuations or noise. The blue broken line represents reference data information, which is generated by performing the above-mentioned denoising processing on the initial data information, retaining the core change trend of the time series data, and removing abnormal fluctuations to ensure the accuracy and representativeness of the data. On the timeline, the initial data information recorded at different time points is gradually processed to form reference data information, which is specifically reflected in the blue broken line being smooth and fitting the main trend of the initial data. Through this processing method, the figure shows the conversion process from initial data information to reference data information, and at the same time emphasizes the importance of data processing, that is, while reducing data redundancy, retaining the key features of the data, thereby providing high-quality input for subsequent analysis and application.

[0067] The generation of a reference time series data information set is a key step to ensure the accuracy of data analysis, especially when there are many data fluctuations or anomalies, providing high-quality input for subsequent processing.

[0068] Further, obtaining the initial time series data information and the initial time point corresponding to the initial time series data information includes:

[0069] Receiving initial time series data information and initial time point acquired by a data acquisition device; or,

[0070] The initial time series data information and the initial time point are obtained from the initial time series database.

[0071] In practical applications, the data acquisition device is a hardware or software module used to obtain real-time data from a device or environment; the initial time series database is a database that stores the collected raw data and its timestamps to support subsequent data processing and analysis.

[0072] A data acquisition device can be understood as a tool for real-time monitoring and recording data. It is directly connected to a sensor or device and obtains raw data by reading the output of the sensor (online mode). For example, in industrial equipment monitoring, a data acquisition device may include an embedded sensor module or an independent data recording device to collect data such as temperature, vibration or pressure during equipment operation. These devices are generally real-time and continuous, and can collect data at fixed time intervals to provide dynamic time series data input for the system.

[0073] The initial time series database can be understood as a storage system (offline mode) used to save the original data acquired by the data acquisition device and its corresponding time points. For example, in the Internet of Things platform, the initial time series database may include a series of equipment operation logs with timestamps. This database not only records the original value of the data, but also retains the timeline information of data acquisition, providing a reliable basis for subsequent time series analysis and data screening. In addition, in online mode, data can be stored directly from the data acquisition device in real time; in offline mode, historical data is loaded into the database by batch import to meet the needs of different scenarios.

[0074] Further, generating a reference time series data information set based on the initial time series data information and an initial time point corresponding to the initial time series data information includes:

[0075] Acquire a data denoising window corresponding to the initial time series data information, and acquire initial starting time series data information and an initial starting time point corresponding to the data denoising window;

[0076] Acquire an initial data time difference based on the initial time point and the initial starting time point, and acquire a denoising window data information range corresponding to the initial time series data information based on the initial starting time series data information and the initial time series data information;

[0077] When the initial time series data information belongs to the window data information range and the initial data time difference is less than or equal to the preset data denoising window time threshold, continue to perform the step of acquiring the initial time series data information and the initial time point corresponding to the initial time series data information;

[0078] In the case that the initial timing data information does not belong to the window data information range or the initial data time difference is greater than the data denoising window time threshold, the initial predecessor timing data information is obtained based on the initial time point, and the initial predecessor timing data information and the initial starting timing data information are stored in a reference timing data information set; the initial timing data information is determined to be the initial starting timing data information of the next data denoising window, and the steps of obtaining the initial timing data information and the initial time point corresponding to the initial timing data information are continued.

[0079] In practical applications, the data denoising window is a logical unit used to segment time series data; the initial starting time series data information is the value of the first sampling point in the denoising window; the denoising window data information range is the range information that describes the upper and lower limits of the current denoising window for the current initial time series information; the initial data time difference is used to measure the time interval between sampling points; the initial predecessor time series data information is the previous valid sampling point of the initial time series data information; the reference time series data information set is a set of sampling points after denoising, which contains multiple groups of valid data.

[0080] The data denoising window can be understood as a strategy for segmenting the original time series data. Its function is to divide the original data into multiple independent segments for processing according to the time and data distribution characteristics. For example, in industrial equipment monitoring, sensors generate a large amount of data per second. The data denoising window can divide this data into several small segments based on a certain time or data characteristic, so as to facilitate subsequent denoising and feature retention operations. The length and width of each data denoising window are determined by the preset threshold and data fluctuation characteristics, ensuring the aggregation of stable segments and timely response to fluctuating segments.

[0081] The initial starting time series data information can be understood as the first data point in the denoising window, which is the reference basis for subsequent data processing in the entire window. For example, in a 10-second sensor monitoring window, the sensor value recorded in the 1st second can be used as the initial starting time series data information, which is used to define the upper and lower limits of the window and as the basis for subsequent data time difference calculation. The denoising window data information range can be understood as the upper and lower boundary values ​​used to determine whether the current data information belongs to the current window. For example, if a window has an upper and lower bound range of 3 and 7 for the current data information, the current data information must fall within this range. If the data information exceeds the range, the end of the window will be triggered and a new window will be opened. By dynamically adjusting the boundary range for each data information, the denoising window can flexibly adapt to changes in different data characteristics.

[0082] The time difference of obtaining the initial data can be understood as a calculation method used to measure the time interval between the current sampling point and the initial sampling point in the data denoising window. For example, in a real-time monitoring system, the time difference between sampling points can reflect the frequency of data updates. When the time difference exceeds the preset threshold, it means that the data may have entered a new stage, and a new denoising window needs to be opened.

[0083] The initial predecessor time series data information can be understood as the previous valid sampling point of the current initial time series data information, which is used to be stored together with the initial starting time series data information at the end of the denoising window, so as to generate a reference time series data information set later. For example, in a scenario with large data fluctuations, when a sampling point does not meet the conditions of the current window, the system will record the previous valid sampling point of the point to ensure the integrity and logic of the transition to the next window. The reference time series data information set can be understood as a valid data set retained after denoising, which is used for subsequent data analysis and processing.

[0084] In the data denoising process, dynamic screening and window division of initial time series data information are key steps to achieve noise reduction effects. For the initial time series data information, when it belongs to the data information range of the current window and the time difference between it and the initial starting time series data information does not exceed the preset time threshold, the significance of continuing to perform data collection and analysis operations is to efficiently utilize the current window and avoid adding additional processing overhead due to premature window termination. Under this condition, the system can fully aggregate valid data in a stable time period, thereby reducing storage pressure while retaining the integrity of the data in the window. Especially in scenarios such as industrial monitoring that require a large amount of time series data, this logic ensures that as much data as possible is aggregated into one window during the stable phase to improve data management efficiency.

[0085] On the other hand, when the initial time series data information no longer meets the window conditions, that is, it is not within the upper and lower bounds of the current window for the initial time series data information, it means that the data may have entered a new stage or fluctuation area. At this time, the current window is closed and a new denoising window is generated. The necessity of this division logic lies in that it can quickly capture the key features of data changes, such as sudden changes in trends or large fluctuations, thereby preventing important information from being diluted or missed. For example, in environmental monitoring, when the sensor reads abnormal temperature fluctuations, the system can use this logic to close the current window in time and start a new window, focusing on recording the detailed data of the abnormal change.

[0086] In addition, storing the initial starting time series data information of the current window and the last valid sampling point (initial predecessor time series data information) in the reference time series data information set can not only provide complete time series features for subsequent analysis, but also reduce the amount of redundant data storage through simplified data point representation. The rationality of this storage mechanism lies in that the starting and predecessor points at the end of the window can fully represent the data characteristics of the window, such as fluctuation range and time span, providing concise and effective input for subsequent modeling.

[0087] By dynamically adjusting the window boundaries and time thresholds, the system can adapt to the diversity and complexity of the data, ensuring the aggregation effect in the stable stage and responding to the data characteristics of the mutation area in a timely manner. This adaptive windowing strategy greatly improves the flexibility of data processing, making data denoising not only reduce noise, but also a time series data management method that takes into account efficiency and feature integrity.

[0088] Further, based on the initial starting time series data information and the initial time series data information, obtaining a denoising window data information range corresponding to the initial time series data information includes:

[0089] Obtaining initial time series data upper limit information and initial time series data lower limit information;

[0090] Based on the initial time series data upper limit information and the initial time series data lower limit information, obtaining the initial time series data range information;

[0091] A denoising window data information range of the data denoising window is generated according to the initial starting time series data information and the range information of the initial time series data.

[0092] In practical applications, the upper limit information of initial time series data is the restriction information generated based on the higher data value in the current window, which is used to indicate the upper limit of the allowable range of data in the window; the lower limit information of initial time series data is the restriction information generated based on the lower data value in the current window, which is used to define the lower limit of the allowable range of data; the range information of initial time series data is the difference between the upper limit information of initial time series data and the lower limit information of initial time series data, which is used to describe the fluctuation range of data distribution in the current window.

[0093] Exemplarily, the upper limit information of the initial time series data can be understood as being used to determine the upper limit of the data in the current window, and its value is affected by the local fluctuations of the data in the window. For example, when monitoring the operation of industrial equipment, the temperature data collected by the sensor may rise slowly within a window, and the upper limit information is dynamically adjusted to reflect this changing trend, thereby avoiding misjudging normal data as noise. The lower limit information of the initial time series data can be understood as the lower allowable value of the data distribution within the window, which is used to filter abnormal data below the normal range. For example, in environmental monitoring, the humidity sensor may occasionally collect abnormally low values ​​due to interference, and the lower limit information can help eliminate these noise points while retaining the valid range of the actual data.

[0094] It should be noted that the method of obtaining the upper limit information and the lower limit information of the initial time series data can be understood as determining the upper and lower limits by obtaining the distribution characteristics of the data in the window. The specific method can be to sort the data collected in the window from large to small, and use the data information ranked first as the upper limit information of the initial time series data, and use the data information ranked last as the lower limit information of the initial time series data.

[0095] Correspondingly, the range information of the initial time series data can be understood as the range between the upper and lower limit information, which is used to measure the fluctuation amplitude and trend change of the data in the current window. Its calculation results can be further used to define the window width. For example, in the acceleration monitoring of vehicle driving, the range information can accurately capture the dynamic changes in the acceleration and deceleration process, providing a basis for subsequent data denoising. At the same time, the range information can also be used to generate the denoising range of the window to ensure that important features such as sudden acceleration or deceleration points are retained.

[0096] It should be noted that the denoising window data information range of the data denoising window is generated according to the initial starting time series data information and the initial time series data extreme difference information. It can be understood that the upper and lower bounds of the window are dynamically adjusted through this information to adapt to the local fluctuation characteristics of the data and ensure effective denoising. The specific method can be to perform weighted summation of the initial starting time series data information and the initial time series data extreme difference information to generate the upper and lower bounds of the denoising window data information range respectively, thereby realizing adaptive adjustment of the dynamic window range; it can also be generated by analyzing the extreme value gap of the data distribution in the current window and combining the time information to reflect the trend change of the data; it can also be combined with contextual data to determine the denoising range of the current window by calculating the superposition effect of multiple windows, so as to more accurately capture abnormal fluctuations or key trends. This manual does not impose any restrictions on this.

[0097] In one embodiment provided in this specification, the method for obtaining the denoising window data information range is as shown in Formula 1:

[0098] ...Formula 1

[0099] in, is the lower bound of the denoising window data information range; is the upper bound of the denoising window data range, i is the i-th initial time series data; The initial starting time series data information corresponding to the current denoising window; To calculate the weight of the information range of the denoising window data; The larger value from the initial starting time series data information to the i-th initial time series data information represents the initial time series data upper limit information; The smaller value from the initial starting time series data information to the i-th initial time series data information represents the initial time series data lower limit information.

[0100] Combining the upper limit information, lower limit information and range information of the initial time series data, the system can adaptively adjust the range of each data denoising window to adapt to the data distribution and volatility characteristics in different scenarios. This method not only ensures the accuracy of denoising, but also effectively reduces errors and improves the reliability of data analysis and decision-making.

[0101] Step 104: Determine a reference data time difference based on the reference time point and the reference start time point.

[0102] In practical applications, the reference data time difference is the time interval between the time point corresponding to the reference time series data information and the reference start time point. The reference data time difference can be understood as the difference between the time of the current data point and the time of the window start point during data processing. For example, in the performance monitoring of a network cluster, the reference data time difference may be the time difference between the acquisition time of the network bandwidth at a certain moment and the start time of the window. By calculating the reference data time difference, the system can determine whether the current data should continue to be included in the existing processing window or a new window should be allocated for the new data. This time difference provides a time constraint for subsequent data acquisition and processing. In specific implementations, the calculation of the reference data time difference is usually based on the reference time point and the reference start time point. For example, in a network bandwidth monitoring system, if the acquisition time of the reference time series data information is "2025-01-18 10:30:00", and the reference start time point is "2025-01-18 10:00:00", the reference data time difference is 30 minutes. Based on this time difference, the system can decide whether to continue collecting data in the current data window, or whether to end the current window and open a new data collection window. This mechanism ensures efficient data processing and storage and avoids the accumulation of invalid data.

[0103] The calculation of the reference data time difference not only helps determine the continuity of the data window, but also can dynamically optimize the window length and width in complex data scenarios, especially when the data fluctuates significantly or mutations occur, the upper and lower bounds can be adjusted through the time difference to capture the mutation characteristics. For example, when the reference data time difference is large and accompanied by significant data fluctuations, the system can expand the window range to ensure that important data is not missed; when the time difference is small, the window range can be narrowed to improve processing efficiency. This dynamic adjustment mechanism based on time difference not only improves the accuracy of data processing, but also can adapt to the time series data analysis needs of different scenarios.

[0104] By calculating these time differences, the system determines whether the data meets the current processing window in subsequent data processing, and decides the data collection and processing flow accordingly.

[0105] Step 106: Generate a window data information range according to the reference data time difference and the reference starting timing data information.

[0106] In practical applications, the window data information range is a numerical range defined by upper and lower bounds during data processing; the upper and lower bounds are calculated by referring to the starting time series data information and the time difference of the reference data, and are used to determine whether the current data belongs to the processing range of the window. The dynamic adjustment of the window data information range can adapt to local changes in the data, thereby effectively controlling the accuracy and efficiency of data processing.

[0107] The window data information range can be understood as the upper and lower limits of the data values ​​defined for each window in the time series data processing. In the specific processing process, the upper and lower limits of the window are calculated by the time difference between the reference starting time series data information and the reference data, and are dynamically corrected in combination with the extreme value of the data in the window. For example, in the network cluster performance monitoring scenario, the upper and lower limits of the window may correspond to the upper and lower limits of the network bandwidth usage in a certain period of time. This range can be dynamically generated by combining the initial value of the bandwidth with the real-time fluctuation, thereby ensuring that data can still be accurately filtered and processed in a changing network environment.

[0108] It should be noted that the window data information range generated according to the reference data time difference and the reference starting time series data information can be understood as dynamically calculating the upper and lower bounds of the window through the time difference and the initial data value, which is used to clarify the scope of the current data processing to ensure the flexibility and accuracy of the data processing. This specific method can dynamically adjust the upper and lower bounds according to the changes in the time interval by combining the reference data time difference and the initial data value, and adaptively expand or reduce the window width according to the local fluctuations of the data to capture the trend changes of the data; it can also generate a constant range directly based on the starting data value by presetting fixed upper and lower bound widths, without considering the changes in the time difference, which is suitable for scenes with relatively stable data; it can also use a prediction model trained based on historical data, combine time difference and historical features to generate a window range, adapt to the data processing needs of multi-variable complex scenes, etc. This manual does not impose any restrictions on this.

[0109] This approach ensures the flexibility and accuracy of data processing, while reducing redundant sampling and improving the processing efficiency of the system. By dynamically adjusting the window range, the system can reduce storage and computing costs in the stable data stage, while accurately recording key features in the mutation stage, providing a reliable basis for subsequent analysis and prediction.

[0110] Further, according to the reference data time difference and the reference starting timing data information, a window data information range is generated, including:

[0111] Acquire predecessor reference timing data information and a predecessor reference time point and a predecessor window data information range corresponding to the predecessor reference timing data information;

[0112] Generate data change characteristic information based on the predecessor reference time point, the reference starting time point, the predecessor window data information range and the reference starting time series data information;

[0113] Based on the data change characteristic information and the reference data time difference, a window data information range corresponding to the reference time series data information is generated.

[0114] In practical applications, the predecessor reference timing data information is the previous valid sampling point in the processed data sequence in the reference data processing window; the predecessor reference time point is the time point corresponding to the predecessor reference timing data information, which is usually used to calculate the time difference; the predecessor window data information range is the upper and lower bounds generated based on the predecessor reference timing data information and the reference starting timing data information, reflecting the distribution characteristics of the data in the predecessor window; the data change characteristic information is quantitative information generated based on the predecessor window data information range, time difference and reference starting timing data information, which is used to describe the trend and amplitude of data changes over time.

[0115] Exemplarily, the predecessor reference time series data information can be understood as a key point used to define the window boundary in time series data processing, and its role is to serve as a reference point for the current window range, thereby ensuring the continuity and accuracy of the data. For example, when monitoring the operation of industrial equipment, the predecessor reference time series data information can be the equipment temperature recorded at the previous time point, and this information is used to determine whether the current temperature exceeds the allowable fluctuation range. The predecessor reference time point can be understood as a time mark paired with the predecessor reference time series data information, which is used to calculate the time interval to analyze the dynamic change characteristics of the data. For example, in traffic flow monitoring, the predecessor reference time point may record the time point when the previous vehicle passed the sensing device, and combined with its time series data information, the speed and interval of the vehicle can be calculated. The predecessor window data information range can be understood as the upper and lower bounds generated by the predecessor reference time series data information, reflecting the statistical characteristics of the data in the window and the allowable variation range. For example, in environmental monitoring, the predecessor window data information range may define the upper and lower limits of the air quality index within a certain period of time, which is used to determine whether the current data is abnormal.

[0116] Data change feature information can be understood as an indicator describing the dynamic change of data based on the previous window data information range, time point and reference start time series data information, aiming to capture the trend of data change over time. For example, in financial market analysis, data change feature information can be used to quantify the rate of change of stock prices and provide key feature information for trend prediction models.

[0117] It should be noted that the generation of data change feature information based on the predecessor reference time point, the reference start time point, the predecessor window data information range and the reference start time series data information can be understood as extracting the dynamic change trend of the time series data by combining the multiple associations of the time dimension and the value range, thereby providing key feature support for subsequent data analysis or processing. The specific method may include: based on the predecessor reference time point and the reference start time point, calculating the time point change information, quantifying the time difference between the current data point and the reference point, and judging whether the data change occurs within a reasonable time range; then, data change information can also be generated based on the predecessor window data information range and the reference start time series data information, and the fluctuation amplitude and trend of the data can be evaluated by comparing the degree of deviation between the current data point and the upper and lower bounds of the window; then, data change feature information can also be generated based on the comprehensive calculation of the time point change information and the data change information, and the description of the dynamic change of the data can be further refined by integrating the characteristics of time and value changes, such as extracting indicators such as acceleration or trend change rate, etc., and this specification does not impose any restrictions on this.

[0118] It should be noted that the window data information range corresponding to the reference time series data information generated based on the data change characteristic information and the reference data time difference can be understood as combining the dynamic change trend of the data with the difference in the time dimension to define the upper and lower bounds that adapt to the current data state, thereby more effectively capturing the key features of the data. The specific method may include: generating an intermediate window data information range based on the data change characteristic information and the reference data time difference, for example, by calculating the dynamic changes of the upper and lower bounds of the current data relative to the reference starting point to obtain the intermediate range; then the reference time series data upper limit information and the reference time series data lower limit information can also be obtained, and the reference time series data range information can be calculated using these two parameters to define the larger amplitude of data fluctuations; then the window data information range corresponding to the reference time series data information can also be further generated based on the intermediate window data information range and the reference time series data range information, and the window boundary can be more accurately defined by combining the dynamic range and range information. This specification does not impose any restrictions on this.

[0119] By obtaining data change feature information and generating a window data information range based on the data change feature information, with the help of this window data information range dynamically generated based on data change feature information, it is possible to flexibly respond to different types of local changes, so that more data can be aggregated in the stable stage to reduce storage and bandwidth requirements, and new windows can be opened in time when fluctuations or mutations occur to retain important features. This significantly improves the accuracy and efficiency of data processing and provides strong support for time series data management in complex systems.

[0120] Further, based on the predecessor reference time point, the reference start time point, the predecessor window data information range and the reference start time series data information, data change characteristic information is generated, including:

[0121] Generate time point change information based on the previous reference time point and the reference start time point;

[0122] Generate data change information according to the preceding window data information range and the reference starting time sequence data information;

[0123] Data change characteristic information is generated based on the time point change information and the data change information.

[0124] In practical applications, the time point change information is a parameter describing the time difference between the previous reference time point and the reference starting time point; the data change information is a parameter describing the change range between the previous window data information range and the reference starting timing data information.

[0125] Exemplarily, the time point change information can be understood as being generated by calculating the time difference between the previous reference time point and the reference start time point. The significance of this information is to reflect the dynamic changes of the data on the time axis and provide the necessary time reference for subsequent data feature extraction. For example, in temperature monitoring, the previous time point may be the recording time of the previous hour, and the reference start time point is the time when a certain day starts. This difference in time points can help analyze the fluctuation trend in a short period of time.

[0126] The data change information can be understood as being generated by the difference between the preceding window data information range and the reference starting time series data information. This information describes the changes in the upper and lower bounds of the data within the window, and is used to reflect local data characteristics, such as the difference between the upper and lower bounds. In practical applications, for example, when analyzing the air pollution index, the data change information can represent the deviation of the upper and lower bounds of the pollution index within a certain window, which helps to understand the amplitude of pollution fluctuations. The method of generating data change information by the difference between the preceding window data information range and the reference starting time series data information can be to calculate the difference between the upper and lower bounds in the preceding window data information range and the reference starting time series data information as the data change information; it can also be to calculate the average of the upper and lower bounds in the preceding window data information range, and then calculate the difference between the average and the reference starting time series data information as the data change information, and so on. This specification does not impose any restrictions on this.

[0127] It should be noted that generating data change feature information based on time point change information and data change information can be understood as extracting important parameters that can reflect data characteristics by combining the dynamic changes of time dimension and data amplitude to support subsequent window processing and feature extraction. The specific method can generate feature information based on the proportional relationship between time point change information and data change information, for example, by calculating the change rate by the ratio of the time point difference to the data upper and lower boundary difference, forming feature information to describe the trend of data change; it can also be obtained by using the time point change information as a weight, and then weighted summing the data change information and the reference starting time point, so as to obtain feature information. This method can better reflect the contribution of changes at different time points to data characteristics; it can also be combined with the time point change information and the data change information by introducing restrictive conditions, such as setting a threshold range for data fluctuations, to generate feature information that meets preset conditions, such as screening out a set of data points with significant change characteristics. This specification does not impose any restrictions on this.

[0128] refer to Figure 3 , Figure 3A schematic diagram of a data processing window provided for an embodiment of the present specification shows the information generation and processing process of the window data range, wherein the horizontal axis represents the reference starting time point and the reference data time point, and the vertical axis represents the amplitude variation range of the reference data information. The upper and lower bounds of the window are jointly determined by the reference starting time point and the reference starting data information. By adjusting the upper and lower bounds of the data in the window, the data fluctuation can be dynamically responded to. The window range is composed of dynamic boundaries corresponding to each reference data information, which are comprehensively generated by the reference starting time series data information and the data change feature information, ensuring that the window can flexibly adapt to the trend changes of the data while avoiding unnecessary data redundancy. Through the dynamic boundary adjustment mechanism shown in the figure, the effective range of reference data processing can be accurately limited in the time dimension, and the window range can be adaptively expanded or contracted according to the changes in the reference data, providing an accurate data basis for subsequent data screening, downsampling, and anomaly detection applications. The data processing window in the figure formally defines the limiting conditions for data changes, so that the processing of each data point has a clear basis, further improving the accuracy and efficiency of the overall data analysis.

[0129] The combination of time point change information and data change information can provide multi-dimensional support for the generation of data change feature information. Through the fusion of time and data information, the generated data change feature information can more comprehensively reflect the dynamic characteristics of the data and provide a reliable basis for application scenarios such as data downsampling, trend analysis and anomaly detection.

[0130] Further, based on the data change characteristic information and the reference data time difference, generating a window data information range corresponding to the reference time series data information includes:

[0131] Generate an intermediate window data information range according to the data change characteristic information and the reference data time difference;

[0132] Obtaining reference time series data upper limit information and reference time series data lower limit information, and obtaining reference time series data range information based on the reference time series data upper limit information and the reference time series data lower limit information;

[0133] A window data information range corresponding to the reference time series data information is generated according to the intermediate window data information range and the reference time series data range information.

[0134] In practical applications, the data information range of the intermediate window is the preliminary window range generated based on the data change characteristic information and the reference data time difference; the reference time series data upper limit information is the upper limit numerical information allowed in the reference time series data set, the reference time series data lower limit information is the corresponding lower limit numerical information, and the reference time series data range information is the difference between the reference time series data upper limit information and the reference time series data lower limit information, which is used to describe the range of data fluctuations.

[0135] The data information range of the intermediate window can be understood as a range generated by linear mapping or nonlinear rules based on the time difference between the data change feature information and the reference data, which is used as the basis for preliminary screening and adjustment. By setting this range, the initial boundaries of the window can be quickly defined, reducing unnecessary data processing workload. For example, in a stage where data is stable, the data information range of the intermediate window can be wider to cover more data points; in a stage where data fluctuates violently, the range can be adaptively reduced to capture data changes more accurately.

[0136] The upper limit information of the reference time series data can be understood as the upper limit data value allowed in a window selected from the reference data set. This information is usually determined in combination with statistical characteristics or real-time data characteristics and is used to identify the upper limit of the data. For example, in a real-time monitoring system, the higher temperature value in a window can be used as the upper limit information of the reference time series data. The lower limit information of the reference time series data can be understood as the smaller data value in the window corresponding to the upper limit information, which is used to identify the lower limit of the data. Similarly, this information can be generated through statistical analysis or adaptively adjusted according to specific task scenarios. For example, in environmental data collection, the lower humidity value of a certain area can be used as the lower limit information of the reference time series data. The range information of the reference time series data can be understood as the difference between the upper limit information and the lower limit information, which is used to indicate the data fluctuation range in a certain data window. Through this range information, the discrete degree or trend change of the data can be effectively measured.

[0137] It should be noted that the window data information range corresponding to the reference time series data information is generated according to the intermediate window data information range and the reference time series data range information. It can be understood as a process of dynamically adjusting the window boundary by combining the preliminary range of the intermediate window with the data fluctuation characteristics to adapt to the characteristics of different data distributions and improve the accuracy of data processing. The specific method is to multiply the upper and lower bounds of the intermediate window data information range by the ratio of the reference time series data range information, respectively, to obtain the upper and lower bounds of the window data information range, and ensure that the window range contains trend-changing data features; it is also possible to directly multiply the range information by the preset weight, and then directly add the upper and lower bounds of the intermediate window data information range to generate a dynamically adjusted window data information range to more sensitively capture local changes in the data; it is also possible to calculate the average value of the data in the window, and then add and subtract the range information to generate an adjusted window data information range to adapt to the data processing requirements of different scenarios. This manual does not impose any restrictions on this.

[0138] In one embodiment provided in this specification, the method of generating data with reference to the window data information range corresponding to the time series data information is as shown in Formula 2:

[0139] ...Formula 2

[0140] Wherein, i represents the i-th reference time series data information; i-1 represents the i-1-th reference time series data information (that is, the predecessor reference data information corresponding to the i-th reference time series data information); is the reference starting time series data information in the reference data processing window (current reference data processing window) corresponding to the i-th reference time series data information; is the reference starting time point in the current reference data processing window; is the reference time point corresponding to the i-th reference time series data information; is the previous reference time point corresponding to the i-th reference time series data information; is the lower bound of the range of the intermediate window data information corresponding to the i-th reference time series data information; is the upper bound of the range of the intermediate window data information corresponding to the i-th reference time series data information; It is the lower bound of the previous window data information range; It is the upper bound of the previous window data information range; and They are respectively the lower bound change information and the upper bound change information in the data change information; It is the time point change information; is the reference time difference.

[0141] is the lower bound of the window data information range corresponding to the i-th reference time series data information; is the upper bound of the range of the intermediate window data information corresponding to the i-th reference time series data information; To calculate the weight of the information range of the denoising window data; The larger value from the reference starting time series data information to the i-th reference time series data information represents the reference time series data upper limit information; The smaller value from the reference starting timing data information to the i-th reference timing data information represents the reference timing data lower limit information.

[0142] By combining the data information range of the intermediate window, the upper limit information of the reference time series data, the lower limit information of the reference time series data, and the range information of the reference time series data, accurate dynamic adjustment of the data window can be achieved, providing accurate range limitation for subsequent noise reduction, sampling or analysis.

[0143] Step 108: When the reference timing data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to execute the step of obtaining the reference timing data information and the reference time point corresponding to the reference timing data information, wherein the target timing data information is representative data information in the reference data processing window.

[0144] When the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, the step of obtaining the reference time series data information and the reference time point is continued to ensure the continuity and context relevance of data processing. The execution of this step can optimize the data grouping logic, improve processing efficiency, and provide a consistent data basis for subsequent analysis and feature extraction.

[0145] It should be noted that this judgment logic can be understood as a dynamic screening process for data, and its main purpose is to retain time series data points that meet the upper and lower bounds and have reasonable time differences in the current data window, thereby reducing data redundancy. For example, when the numerical fluctuation of the reference time series data information meets the preset window upper and lower bounds (such as the network bandwidth usage rate is within a certain stable range), and the time interval of data collection is within the allowable range, it means that these data points are statistically consistent and can be classified as the same processing window.

[0146] Specifically, this step ensures that the current data point maintains continuity with the existing window by comparing it with the upper and lower bounds of the window and the time threshold. Once the conditions are met, the system will continue to collect subsequent data points, thereby extending the window usage time. For example, in industrial monitoring scenarios, when the data collected by the sensor (such as equipment temperature) is within the current window range and the time interval is reasonable, these data points can be regarded as a continuation of the stable state, thereby reducing the frequency of opening new windows and reducing computing and storage costs.

[0147] In addition, this step is closely related to the extraction of subsequent target time series data information. By continuously processing qualified data points, complete data support can be provided for the extraction of target data at the end of the window. For example, in a fluctuating scenario, data points with trend changes can be captured through continued collection, providing more reliable basic data for subsequent downsampling and feature extraction. This approach not only improves the efficiency of data processing, but also accurately retains key features, providing strong support for a variety of downstream applications.

[0148] Step 110: When the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, obtain the target timing data information based on the reference time point, and store the target timing data information and the reference starting timing data information; determine that the reference timing data information is the reference starting timing data information of the next reference data processing window, and continue to execute the steps of obtaining the reference timing data information and the reference time point corresponding to the reference timing data information, wherein the target timing data information is representative data information in the reference data processing window.

[0149] In practical applications, the target time series data information is the key data point in the current window for representative storage and processing; the reference time series data information is the time series data point collected or processed, and the reference start time series data information is the first data point in the window that is included in the processing range. The combination of this information is used to select and store key data in the window, and supports subsequent data analysis and downsampling.

[0150] Acquiring the target time series data information based on the reference time point can be understood as using the reference time point to select representative data points in the processing window when the window end condition is met. Since the reference time series information corresponding to the reference time point does not belong to the data processing window that has ended, the method of acquiring the target time series data information based on the reference time point can be to determine the reference time series data information corresponding to the previous time point of the reference time point as the target time series information.

[0151] The target time series data information can be understood as a representative data point selected through a series of judgments and processing in the current window. It is usually a key point in the window, such as the end sampling point of the window, which is used to simplify data storage while retaining the core characteristics of the data. For example, in a network performance monitoring scenario, if multiple bandwidth usage change points are recorded in the window, the target time series data information is the data point at the end of the window.

[0152] The selection of target time series data information is usually determined by comparing it with the reference time series data information. For example, when the reference time series data information exceeds the upper and lower bounds of the window or its time difference exceeds the preset threshold, the system will mark the previous sampling point of the current data as the target time series data information and store it in the data set. This method ensures that the end point of the window can accurately represent the data characteristics in the current window, while supporting subsequent data downsampling and feature extraction tasks.

[0153] The determination of the target time series data information can also adapt to scenarios with data variation. For example, when the trend of change in the time series data is more obvious, the system can select certain feature points within the window as the target time series data information, such as peaks or troughs, to better reflect the fluctuation characteristics of the data. In this case, the target time series data information is not only the end point of the window, but may also be a data point with significant features within the window, thereby providing richer basic data for subsequent anomaly detection and trend analysis. Through this flexible selection mechanism, the target time series data information can effectively improve the system's data processing efficiency and analysis capabilities in a variety of scenarios.

[0154] refer to Figure 4 , Figure 4 A schematic diagram of a reference data information and stored data information provided for an embodiment of the present specification, wherein the blue broken line represents the reference data information, and the red broken line represents the data information for storage after processing. The reference data information is generated from the initial data through steps such as denoising and feature retention, reflecting the changing trend of the data at a point in time. The processed stored data information further optimizes the structure of the data, and only retains key change points, such as peaks and turning points, through downsampling, thereby ensuring the retention of key features while reducing data storage space. In the figure, the horizontal axis of the time point represents the acquisition time, and the vertical axis represents the data information. The blue reference data information shows the original changing trend, and the red stored data information is stored at important change points of the reference data information through window processing.

[0155] When the reference time series data information does not belong to the window data information range or the reference data time difference is greater than the preset data processing window time threshold, the main purpose of executing this step is to draw a boundary for the current window processing process and divide an independent window for the new data characteristics or time period. This logic not only ensures the effective organization and storage of data, but also significantly improves the flexibility and adaptability of the system in processing diverse data scenarios.

[0156] It should be noted that this step effectively distinguishes the continuity and independence of the data by judging whether the current data exceeds the upper and lower bounds of the window or the time difference limit. When the reference time series data information exceeds the window range, it may indicate that the data characteristics have changed significantly, such as a sudden change or trend turning point in sensor data; and when the time difference exceeds the threshold, it indicates that the time interval is long enough and the current window can no longer represent the next data point. Through such judgment logic, the system can ensure that the data contained in each window is consistent in features, thereby providing high-quality input for subsequent feature extraction, analysis, and downsampling.

[0157] On this basis, storing the target time series data information and the reference starting time series data information is an important part of the data processing process. The target time series data information is used as the representative data point of the current window. The key data points at the end of the window, such as the larger value, the smaller value or the trend turning point, are usually selected; and the reference starting time series data information is used as the starting point of the next window to ensure that subsequent processing has a clear starting basis. For example, in an industrial equipment monitoring scenario, when the operating status of the equipment is recorded in a window, when the equipment status suddenly changes or the sampling time interval is too long, the system will store the end data point of the current window and the start data point of the next window, thereby providing a complete time series basis for subsequent anomaly detection.

[0158] The context-associated logic also includes continuous processing of data streams. After completing the processing of the current window, the reference time series data information is used as the starting point of the new window, which helps to quickly connect the collection and analysis of new data and avoid unnecessary processing interruptions. This design method not only reduces the redundancy of data processing, but also provides technical support for real-time analysis and efficient decision-making. By dividing the window and storing the target data points, this step lays a solid foundation for the system to achieve downsampling, feature retention and data optimization in large-scale time series data processing.

[0159] By applying the solution of the embodiment of this specification, the method can effectively manage and optimize the data processing process by dynamically generating a window data information range based on the reference time series data information. In this process, the relationship between the reference time point and the reference start time point helps define the judgment condition of whether the current data belongs to the same window. When the data meets the condition and the time difference does not exceed the preset threshold, the new data can be smoothly included in the same processing window, thereby effectively reducing redundant sampling and reducing storage and transmission requirements. In addition, when the data no longer meets the range, the window automatically ends and a new window is allocated for the new data to ensure independent storage and processing of important data features.

[0160] This window processing method based on dynamic adjustment of time difference and reference information can flexibly adapt to different types of local data changes. In the data stability stage, the system can merge more time series information to further reduce data redundancy; when data fluctuations or mutations occur, the system can close the current window and open a new window in time to ensure that the fluctuation characteristics are fully preserved. In this way, the system can not only significantly reduce the pressure of storage and bandwidth, but also provide a more accurate data foundation for subsequent real-time analysis and monitoring, avoiding the risk of data falsification. In addition, the real-time window adjustment mechanism of this method also has high flexibility and adaptability, which can meet the processing requirements of multi-source and large-scale data environments such as the Internet of Things.

[0161] Corresponding to the above method embodiment, this specification also provides a data screening method embodiment, see Figure 5 , Figure 5 A flow chart of a data screening method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0162] Step 502: Receive reference timing data information and a reference time point corresponding to the reference timing data information, and obtain a reference data processing window corresponding to the reference timing data information, as well as reference starting timing data information and a reference starting time point corresponding to the reference data processing window.

[0163] Step 504: Determine a reference data time difference based on the reference time point and the reference start time point.

[0164] Step 506: Generate a window data information range according to the reference data time difference and the reference starting timing data information.

[0165] Step 508: When the reference timing data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to execute the step of obtaining the reference timing data information and the reference time point corresponding to the reference timing data information.

[0166] Step 510: When the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, obtain the target timing data information based on the reference time point, and store the target timing data information and the reference starting timing data information in a data information screening result set; determine that the reference timing data information is the reference starting timing data information of the next reference data processing window, and continue to execute the steps of receiving the reference timing data information and the reference time point corresponding to the reference timing data information, wherein the target timing data information is representative data information in the reference data processing window.

[0167] In practical applications, the data information screening result set is the key data set extracted from the time series data through the screening method; the reference time series data information is the input data in the screening process, and the target time series data information is the representative data point selected in each data processing window. These information combined together form an important support for time series data downsampling and feature extraction, and provide a basis for subsequent analysis and storage.

[0168] The data information screening result set can be understood as a subset of time series data obtained through a series of data screening steps. These data points are the simplification and optimization of the original data while retaining important feature information. For example, in the scenario of network performance monitoring, the data information screening result set may contain the starting bandwidth, ending bandwidth, and larger or smaller bandwidth values ​​in each time window. These data points are selected through precise screening rules and represent the changing trend and core features of the data in the entire window.

[0169] During the screening process, the generation of the data information screening result set is usually combined with the upper and lower bounds of the window, the reference data time difference and other conditions. For example, when the reference time series data information exceeds the window data information range or the time difference exceeds the threshold, the system will store the start and end sampling points of the current window into the screening result set. This method ensures that each window is reasonably divided and the storage of each data point has a clear meaning and purpose.

[0170] The generation of the data information screening result set can also be optimized based on the data variation characteristics. In some scenarios, such as industrial monitoring or meteorological analysis, the screening result set may include peak or trough data points within the window to capture sudden changes or extreme situations. Through this feature-driven screening method, the data information screening result set not only reflects the representativeness of the data, but also provides accurate data support for subsequent trend prediction, anomaly detection and other applications.

[0171] The above is a schematic scheme of a data screening method of this embodiment. It should be noted that the technical scheme of the data screening method and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the data screening method that are not described in detail can be found in the description of the technical scheme of the above data processing method.

[0172] By applying the scheme of the embodiments of this specification, different types of local changes can be flexibly handled by screening and dynamically managing the reference information of the time series data. In the data processing process, a window data information range is first generated based on the reference time series data information and the corresponding reference time point, combined with the calculated time difference. This range allows data to continuously enter the same window during the stable phase, thereby effectively utilizing the data consistency within the time period, reducing the generation of redundant sampling points, and reducing storage and bandwidth consumption. When the data changes significantly or fluctuates, the system can automatically determine and close the current window, and start a new window to process subsequent data, thereby effectively retaining the key change characteristics of the data.

[0173] The flexibility of this method not only ensures the efficiency of data storage, but also retains core features such as trend changes and mutations, ensuring that key time series information is not lost. Compared with traditional fixed windows or methods that are not adapted to data fluctuations, the strategy of adaptively adjusting the window range and length can greatly reduce the amount of data without sacrificing data accuracy. This feature enables efficient data stream processing and analysis in high-demand applications such as real-time monitoring and intelligent diagnosis, while ensuring the integrity of data authenticity, because all output data comes from the original time series data, avoiding the risk of forged data. This method can not only significantly reduce the storage cost of data, but also improve the system's processing power and engineering friendliness, so that it can better adapt to the needs of the Internet of Things system architecture, thereby laying a solid foundation for more complex data analysis tasks.

[0174] Corresponding to the above method embodiment, this specification also provides a data compression method embodiment, see Figure 6 , Figure 6 A flow chart of a data compression method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0175] Step 602: Obtain reference time series data information and a reference time point corresponding to the reference time series data information in the time series data set to be compressed, and obtain a reference data processing window corresponding to the reference time series data information, as well as reference starting time series data information and a reference starting time point corresponding to the reference data processing window.

[0176] Step 604: Determine a reference data time difference based on the reference time point and the reference start time point.

[0177] Step 606: Generate a window data information range according to the reference data time difference and the reference starting timing data information.

[0178] Step 608: When the reference timing data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to execute the step of obtaining the reference timing data information and the reference time point corresponding to the reference timing data information.

[0179] Step 610: When the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target timing data information is obtained based on the reference time point, and the target timing data information and the reference starting timing data information are stored in a target timing data set; the reference timing data information is determined to be the reference starting timing data information of the next reference data processing window, and the steps of receiving the reference timing data information and the reference time point corresponding to the reference timing data information are continued until all the reference timing data information in the timing data set to be compressed are processed, and a data compression result is generated based on the target timing data set, wherein the target timing data information is representative data information in the reference data processing window.

[0180] In practical applications, the time series data set to be compressed is the original time series data before data compression processing; the data compression result is the key data set output after the compression algorithm is processed, which is used to replace the original data for storage or analysis. The two are closely related in the data compression process. The former is the input of the compression operation, and the latter is the output of the compression operation. The two work together to reduce the cost of data storage and processing while retaining the important characteristics of the data.

[0181] The time series data set to be compressed can be understood as the original time series data set used as input in the data compression process. These data are usually continuous data points collected from sensors. For example, in a network monitoring scenario, the time series data set to be compressed may be a series of bandwidth usage data collected. The characteristics of these data are that they cover a wide time range and have dense data points. Therefore, they need to be compressed to reduce redundant data points to meet storage and transmission requirements.

[0182] The data compression result can be understood as a streamlined data set obtained after processing by a compression algorithm, which is used to replace the original time series data for storage and analysis. The generation of data compression results needs to retain the key features of the time series data while reducing the amount of data. For example, in an industrial monitoring scenario, the data compression result may only contain the start point, end point, and certain feature points (such as larger or smaller values) of each time window. These results can significantly reduce the data volume while still retaining enough information to support downstream analysis tasks.

[0183] In specific operations, the generation of data compression results is often combined with a data window mechanism. By calculating the time difference of the reference data, the data is divided into multiple windows, and each window only retains the starting point and the end point as the compression result. This method is suitable for relatively stable data scenarios; in scenarios with large data fluctuations, the compression results can also include key feature points in the fluctuations (such as trend change points) to ensure the integrity and representativeness of the data. In this way, the data compression results not only effectively reduce storage costs, but also provide an efficient input data foundation for subsequent trend prediction and anomaly detection.

[0184] The above is a schematic scheme of a data compression method of this embodiment. It should be noted that the technical scheme of the data compression method and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the data compression method that are not described in detail can all be referred to the description of the technical scheme of the above data processing method.

[0185] By applying the scheme of the embodiment of this specification, the data processing window is dynamically generated based on the time difference and the reference time series information, and the scheme realizes the flexible adjustment of the window during the data compression process. When the time difference of the time series data is within the preset threshold range and still belongs to the generated window range, the new data will continue to be included in the current window for processing, thereby effectively reducing the collection of redundant data. Through this adaptive adjustment method of the dynamic window, more data can be aggregated in a stable time period to reduce storage and bandwidth requirements, while avoiding the loss of key information when the data fluctuates. In the mutation or fluctuation stage of the data, the rapid start of the new window can capture important features in time to ensure that core data such as trend changes and mutations are effectively retained. Through intelligent control and real-time adjustment of each window, this scheme not only significantly reduces the redundant sampling caused by the data compression process, but also ensures the authenticity of the data and avoids the risk of forged data. This method can reduce the data volume in large-scale data processing without losing important information, greatly improving the efficiency and reliability of data analysis and storage systems.

[0186] See also Figure 7 , Figure 7The structure diagram of a data screening system provided by an embodiment of the present specification is shown. The data screening system may include a data collection terminal 100, a data screening terminal 200 and a data storage terminal 300;

[0187] The data collection terminal 100 is used to send reference time series data information and reference time points to the data screening terminal 200;

[0188] The data screening terminal 200 is used to receive reference timing data information and a reference time point corresponding to the reference timing data information, and obtain a reference data processing window corresponding to the reference timing data information, as well as reference starting timing data information and a reference starting time point corresponding to the reference data processing window; determine a reference data time difference based on the reference time point and the reference starting time point; generate a window data information range according to the reference data time difference and the reference starting timing data information; and continue to obtain the reference timing data information and the reference timing data information when the reference timing data information belongs to the window data information range and the reference data time difference is less than or equal to a preset data processing window time threshold. The step of determining a reference time point corresponding to the reference timing data information; when the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, obtaining the target timing data information based on the reference time point, and storing the target timing data information and the reference starting timing data information; determining that the reference timing data information is the reference starting timing data information of the next reference data processing window, and continuing to perform the step of receiving the reference timing data information and the reference time point corresponding to the reference timing data information, wherein the target timing data information is representative data information in the reference data processing window; and sending a data information screening result set to the data storage terminal 300;

[0189] The data storage end 300 is used to receive and store the data information screening result set sent by the data screening end 200.

[0190] By applying the solution of the embodiments of this specification, through the architectural design of the data screening system, it is possible to effectively manage and process data streams from different sources. The system can generate a suitable data processing window range based on the time difference of the reference data by dynamically managing the reference time series data information and the reference time point. This method ensures that within a stable data segment, the collected time series data can continuously enter the same window, avoiding the generation of redundant data, thereby reducing the pressure on storage and bandwidth. When the data changes significantly or fluctuates, the system can quickly identify and terminate the current window and start a new window to process subsequent data, which can not only effectively retain important trend and mutation information, but also maintain the integrity and authenticity of the data.

[0191] The data screening system may include multiple data collection terminals 100 and data screening terminals 200, wherein the data collection terminals 100 may be referred to as terminal-side devices, and the data screening terminals 200 may be referred to as cloud-side devices. Multiple data collection terminals 100 may establish communication connections through the data screening terminals 200. In the data screening scenario, the data screening terminals 200 are used to provide data screening services between multiple data collection terminals 100. Multiple data collection terminals 100 may serve as sending terminals or receiving terminals, respectively, and realize communication through the data screening terminals 200.

[0192] The user can interact with the data screening terminal 200 through the data collection terminal 100 to receive data sent by other data collection terminals 100, or send data to other data collection terminals 100, etc. In the data screening scenario, the user can publish a data stream to the data screening terminal 200 through the data collection terminal 100, and the data screening terminal 200 generates a data information screening result set according to the data stream, and pushes the data information screening result set to other data storage terminals with which communication is established.

[0193] The data collection terminal 100 and the data screening terminal 200 are connected to each other through a network, and the data screening terminal 200 and the data storage terminal 300 are connected to each other through a network. The network provides a medium for the communication link between the data collection terminal 100 and the data screening terminal 200, and between the data screening terminal 200 and the data storage terminal 300. The network may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The data transmitted by the data collection terminal 100 and the data screening terminal 200 may need to be encoded, transcoded, compressed, etc. before being released to the data screening terminal 200 and the data storage terminal 300.

[0194] Data acquisition terminal 100 is a front-end device in the system, responsible for regularly collecting sensor data and reporting it to other modules in the system. The terminal device may include multiple sensors and a periodic acquisition module, wherein the periodic acquisition module periodically acquires data from the sensor at a preset time interval and transmits it to the subsequent data processing system. Each sensor can monitor a specific physical quantity and transmit the measured data to the data acquisition module through a communication module. The data acquisition terminal 100 can be flexibly deployed on various terminal devices, such as mobile terminals, computers, embedded devices, etc.

[0195] The data screening end 200 may include servers that provide various services, such as servers that provide communication services for multiple clients, servers for background training that support models used on clients, and servers that process data sent by clients. It should be noted that the data screening end 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0196] The data storage end 300 is an important component of the system, which is used to store and manage the time series data processed by the data screening end 200. The data storage end 300 usually includes a time series database, which is specially optimized to process time series data. The time factor is crucial in query and storage operations of such data. The data storage end 300 is usually connected to the downstream application platform, and the stored time series data is used by multiple modules, such as time series forecasting, industrial composition, data reporting and anomaly detection modules, which all rely on the stored data to perform their respective functions. The data storage end ensures that the collected time series data can be effectively accessed for subsequent processing, analysis, trend identification and decision-making, and therefore plays a vital role in ensuring data integrity and accessibility.

[0197] It is worth noting that the data screening method provided in the embodiments of this specification is generally executed by the server, but in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the data screening method provided in the embodiments of this specification. In other embodiments, the data screening method provided in the embodiments of this specification may also be jointly executed by the client and the server.

[0198] See also Figure 8 , Figure 8 An architecture diagram of a data compression system provided by an embodiment of the present specification is shown, and the data compression system may include a client 400 and a server 500;

[0199] The client 400 is used to send a data compression request to the server 500, wherein the data compression request includes a set of time series data to be compressed;

[0200] The server 500 is used to obtain reference time series data information and a reference time point corresponding to the reference time series data information in the time series data set to be compressed, and obtain a reference data processing window corresponding to the reference time series data information, as well as reference starting time series data information and a reference starting time point corresponding to the reference data processing window; determine a reference data time difference based on the reference time point and the reference starting time point; generate a window data information range according to the reference data time difference and the reference starting time series data information; if the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to a preset data processing window time threshold, continue to execute the steps of obtaining reference time series data information and a reference time point corresponding to the reference time series data information; in the reference time series data When the information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target time series data information is obtained based on the reference time point, and the target time series data information and the reference starting time series data information are stored in the target time series data set; the reference time series data information is determined to be the reference starting time series data information of the next reference data processing window, and the steps of receiving the reference time series data information and the reference time point corresponding to the reference time series data information are continued until all the reference time series data information in the time series data set to be compressed are processed, and a data compression result is generated based on the target time series data set, wherein the target time series data information is representative data information in the reference data processing window; and the data compression result is sent to the client 400;

[0201] The client 400 is also used to receive the data compression result sent by the server 500.

[0202] The data compression system may include multiple clients 400 and a server 500, wherein the client 400 may be referred to as a terminal device and the server 500 may be referred to as a cloud device. Multiple clients 400 may establish a communication connection through the server 500. In the data compression scenario, the server 500 is used to provide data compression services between multiple clients 400. Multiple clients 400 may serve as a sender or a receiver respectively and realize communication through the server 500.

[0203] The user can interact with the server 500 through the client 400 to receive data sent by other clients 400, or send data to other clients 400, etc. In the data compression scenario, the user can publish a data stream to the server 500 through the client 400, and the server 500 generates a data compression result according to the data stream and pushes the data compression result to other clients that establish communication.

[0204] The client 400 and the server 500 are connected via a network. The network provides a medium for a communication link between the client 400 and the server 500. The network may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The data transmitted by the client 400 may need to be encoded, transcoded, compressed, etc. before being released to the server 500.

[0205] The client 400 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language Version 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 400 can be based on the software development kit (SDK, Software Development Kit) of the corresponding service provided by the server 500, such as based on the real-time communication (RTC, Real Time Communication) SDK development and acquisition. The client 400 can be deployed in an electronic device and needs to rely on the device to run or some APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0206] The server 500 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers for background training that support models used on clients, and servers that process data sent by clients. It should be noted that the server 500 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0207] It is worth noting that the data compression method provided in the embodiments of this specification is generally executed by the server, but in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the data compression method provided in the embodiments of this specification. In other embodiments, the data compression method provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0208] The above is a schematic scheme of a data compression system of this embodiment. It should be noted that the technical scheme of the data compression system and the technical scheme of the data compression method and the data screening system described above are of the same concept, and the details of the technical scheme of the data compression system not described in detail can be found in the description of the technical scheme of the data compression method and the data screening system described above.

[0209] By applying the scheme of the embodiment of this specification, the compression and transmission of time series data can be efficiently realized through the data compression system architecture constructed between the client and the server. In this system, the server is responsible for dynamically generating a data processing window and controlling the data compression process based on the reference time series data and the corresponding time point information. By accurately controlling the data time difference and reference information, the processing range of the data can be flexibly adjusted to ensure that the compression of the data does not lose important time series features. When the data is within the appropriate window range and the time difference is within the preset threshold, the system will continue to collect data to reduce redundant sampling, reduce bandwidth and storage costs. At the same time, when the data fluctuates greatly or exceeds the window range, the system will end the current window in time and create a new data processing window to ensure that important features can be effectively captured when the data changes suddenly and the trend changes. The implementation of this system not only improves the flexibility and efficiency of the data compression process, but also significantly reduces the network burden and storage pressure on the basis of ensuring data integrity and authenticity. In addition, the system supports data interaction between multiple clients, has good scalability, can adapt to the compression requirements in different scenarios, and effectively improves the overall performance and response speed of the system.

[0210] The following combination Fig. 9 , taking the application of the data processing method provided in this specification in network performance data screening as an example, the data processing method is further described. Fig. 9 A processing flow chart of a network performance data screening method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0211] Step 902: Acquire initial time series data information and initial time point collected by the network cluster performance data collection module.

[0212] Step 904: Determine a data denoising window corresponding to the initial time series data information.

[0213] Step 906: Calculate the initial data time difference between the initial time point and the initial starting time point in the data denoising window.

[0214] Step 908: Determine whether the above initial data time difference is less than or equal to the preset data denoising window time threshold. If so, execute step 910; if not, execute step 918.

[0215] Step 910: Calculate the range information of the initial time series data based on the upper limit information of the initial time series data and the lower limit information of the initial time series data.

[0216] Step 912: perform weighted summation of the initial starting data information and the initial time series data range information in the above data denoising window to obtain the denoising window data upper limit information and the denoising window data lower limit information respectively.

[0217] Step 914: Determine the denoising window data information range based on the denoising window data upper limit information and the denoising window data lower limit information.

[0218] Step 916: Determine whether the above-mentioned initial time series data information belongs to the above-mentioned denoising window data information range. If so, return to execute step 902, if not, execute step 918.

[0219] Step 918: Obtain the initial predecessor time series data information at a time point before the above-mentioned initial time series data information, and store the above-mentioned initial predecessor time series data information and the initial starting data information in the data denoising window into a reference data information set.

[0220] Step 920: Use the above initial time series data information as the initial starting time series data information of the next data denoising window.

[0221] Step 922: Determine whether the network cluster performance data collection module continues to collect initial time series data information. If so, execute step 902; if not, execute step 924.

[0222] Step 924: Obtain reference time series data information and a reference time point corresponding to the reference time series data information in the reference data information set.

[0223] Step 926: Determine the reference data processing window corresponding to the initial time series data information, and calculate the reference data time difference between the above reference time point and the reference starting time point in the reference data processing window.

[0224] Step 928: Determine whether the reference data time difference is less than or equal to a preset data processing window time threshold. If so, execute step 930; if not, execute step 942.

[0225] Step 930: Obtain the predecessor window data information range and the predecessor reference time point corresponding to the reference time series data information at the previous time point of the reference time series data information.

[0226] Step 932: Generate time point change information based on the above-mentioned predecessor reference time point and the reference starting time point in the reference data processing window, and generate data change information based on the above-mentioned predecessor window data information range and the reference starting timing data information in the reference data processing window.

[0227] Step 934: Calculate data change characteristic information based on the above time point change information and data change information, and multiply the above data change characteristic information and the above reference data time difference to generate an intermediate window data information range.

[0228] Step 936: Obtain the range information of the reference time series data, and perform weighted summation of the reference time series data range information and the reference starting data information in the reference data processing window to obtain the window data upper limit information and the window data lower limit information.

[0229] Step 938: Compare the intermediate window data information range with the window data upper limit information and the window data lower limit information to obtain the window data information range corresponding to the reference time series data information.

[0230] Step 940: Determine whether the reference timing data information belongs to the window data information range. If yes, execute step 924; if no, execute step 942.

[0231] Step 942: Obtain the reference predecessor time series data information at a time point before the reference time series data information, and store the reference predecessor time series data information and the reference start data information in the data processing window into a data acquisition result set.

[0232] Step 944: Use the above-mentioned reference timing data information as the reference starting timing data information of the next data processing window.

[0233] Step 946: Determine whether the reference timing data information has subsequent reference timing data information. If so, execute step 924; if not, execute step 948.

[0234] Step 948: Write the above data collection result set into the network cluster database.

[0235] By applying the scheme of the embodiment of this specification, the system can effectively manage and optimize the network data screening process by combining the network performance data collection module with the denoising processing window. After data collection, the denoising window is first generated based on the initial time series data and time point information, and the current data is determined by calculating the data time difference whether it should be included in the same window. When the data meets the preset conditions and the time difference does not exceed the threshold, the system continues to process the data in the window, effectively reducing redundant sampling and reducing bandwidth and storage pressure. For cases where the data fluctuates greatly or exceeds the window range, the system will automatically end the current window and allocate a new processing window for the new data, thereby ensuring the independence of the data and the complete retention of important features. Through this dynamic window management based on time difference and data change characteristics, the system can flexibly respond to different types of local changes, merge more time series data during the data stability period, and further reduce redundancy; and in the case of fluctuations or mutations, the window is switched in time to ensure feature integrity. This method significantly improves the processing efficiency of network performance data, not only reduces storage and transmission requirements, but also provides a reliable data basis for subsequent data analysis and real-time monitoring, avoids the risk of forged data, and has strong adaptability, which can meet the processing requirements of multi-source, large-scale network data environments.

[0236] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Fig.10 FIG. 1 is a schematic diagram showing the structure of a data processing device provided by an embodiment of the present specification. Fig.10 As shown, the device comprises:

[0237] The acquisition module 1002 is configured to acquire reference time series data information and a reference time point corresponding to the reference time series data information, and acquire a reference data processing window corresponding to the reference time series data information, and reference starting time series data information and a reference starting time point corresponding to the reference data processing window;

[0238] A time difference calculation module 1004 is configured to determine a reference data time difference based on the reference time point and the reference start time point;

[0239] A range calculation module 1006 is configured to generate a window data information range according to the reference data time difference and the reference starting timing data information;

[0240] A loop module 1008 is configured to continue to execute the step of obtaining reference time series data information and a reference time point corresponding to the reference time series data information when the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to a preset data processing window time threshold;

[0241] The storage module 1010 is configured to obtain the target timing data information based on the reference time point and store the target timing data information and the reference starting timing data information when the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold; determine that the reference timing data information is the reference starting timing data information of the next reference data processing window, and continue to execute the steps of obtaining the reference timing data information and the reference time point corresponding to the reference timing data information, wherein the target timing data information is representative data information in the reference data processing window.

[0242] Optionally, the acquisition module 1002 is further configured to:

[0243] Acquire initial time series data information and an initial time point corresponding to the initial time series data information;

[0244] Generate a reference time series data information set based on the initial time series data information and an initial time point corresponding to the initial time series data information;

[0245] Reference time series data information and a reference time point corresponding to the reference time series data information are determined in the reference time series data information set.

[0246] Optionally, the acquisition module 1002 is further configured to:

[0247] Receiving initial time series data information and initial time point acquired by a data acquisition device; or,

[0248] The initial time series data information and the initial time point are obtained from the initial time series database.

[0249] Optionally, the acquisition module 1002 is further configured to:

[0250] Acquire a data denoising window corresponding to the initial time series data information, and acquire initial starting time series data information and an initial starting time point corresponding to the data denoising window;

[0251] Acquire an initial data time difference based on the initial time point and the initial starting time point, and acquire a denoising window data information range corresponding to the initial time series data information based on the initial starting time series data information and the initial time series data information;

[0252] When the initial time series data information belongs to the window data information range and the initial data time difference is less than or equal to the preset data denoising window time threshold, continue to perform the step of acquiring the initial time series data information and the initial time point corresponding to the initial time series data information;

[0253] In the case that the initial timing data information does not belong to the window data information range or the initial data time difference is greater than the data denoising window time threshold, the initial predecessor timing data information is obtained based on the initial time point, and the initial predecessor timing data information and the initial starting timing data information are stored in a reference timing data information set; the initial timing data information is determined to be the initial starting timing data information of the next data denoising window, and the steps of obtaining the initial timing data information and the initial time point corresponding to the initial timing data information are continued.

[0254] Optionally, the acquisition module 1002 is further configured to:

[0255] Obtaining initial time series data upper limit information and initial time series data lower limit information;

[0256] Based on the initial time series data upper limit information and the initial time series data lower limit information, obtaining the initial time series data range information;

[0257] A denoising window data information range of the data denoising window is generated according to the initial starting time series data information and the range information of the initial time series data.

[0258] Optionally, the range calculation module 1006 is further configured to:

[0259] Acquire predecessor reference timing data information and a predecessor reference time point and a predecessor window data information range corresponding to the predecessor reference timing data information;

[0260] Generate data change characteristic information based on the predecessor reference time point, the reference starting time point, the predecessor window data information range and the reference starting time series data information;

[0261] Based on the data change characteristic information and the reference data time difference, a window data information range corresponding to the reference time series data information is generated.

[0262] Optionally, the range calculation module 1006 is further configured to:

[0263] Generate time point change information based on the previous reference time point and the reference start time point;

[0264] Generate data change information according to the preceding window data information range and the reference starting time sequence data information;

[0265] Data change characteristic information is generated based on the time point change information and the data change information.

[0266] Optionally, the range calculation module 1006 is further configured to:

[0267] Generate an intermediate window data information range according to the data change characteristic information and the reference data time difference;

[0268] Obtaining reference time series data upper limit information and reference time series data lower limit information, and obtaining reference time series data range information based on the reference time series data upper limit information and the reference time series data lower limit information;

[0269] A window data information range corresponding to the reference time series data information is generated according to the intermediate window data information range and the reference time series data range information.

[0270] The above is a schematic scheme of a data processing device of this embodiment. It should be noted that the technical scheme of the data processing device and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the data processing device that are not described in detail can be referred to the description of the technical scheme of the above data processing method.

[0271] By applying the scheme of the embodiment of this specification, the system can effectively manage and optimize the processing flow of time series data by constructing a flexible data processing device. Each module in the device ensures that each data point can be accurately assigned to a suitable processing window by dynamically calculating the time difference and generating the window data information range. Specifically, by calculating the relationship between the reference time series data information and the corresponding time point, the system can determine whether the current data should be included in the same window. If the data time difference does not exceed the preset threshold, the data will continue to be processed in the same window, thereby reducing redundant sampling and reducing storage and bandwidth consumption. When the data fluctuates or mutates, the system can end the current window in time and start a new window. This adaptive window switching not only ensures the complete retention of the mutation characteristics, but also ensures the flexibility and accuracy of data processing. In addition, by dynamically adjusting each window, the device can effectively adapt to different types of local changes, integrating more information to reduce redundancy during the data stable period, and adjusting in time during the fluctuation period to capture key change characteristics. This solution significantly improves the system's processing efficiency of time series data and the scalability of data storage, while avoiding the risk of forged data, and provides a reliable foundation for subsequent data analysis and real-time monitoring.

[0272] Fig.11 The block diagram of a computing device 1100 according to one embodiment of the present specification is shown. The components of the computing device 1100 include but are not limited to a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0273] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0274] In one embodiment of the present specification, the above components of the computing device 1100 and Fig.11 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.11 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0275] The computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1100 may also be a mobile or stationary server.

[0276] The processor 1120 is used to execute the following computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned data processing method, data screening method and data compression method.

[0277] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical schemes of the above-mentioned data processing method, data screening method and data compression method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical schemes of the above-mentioned data processing method, data screening method and data compression method.

[0278] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned data processing method, data screening method and data compression method.

[0279] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the data processing method, data screening method and data compression method described above belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the data processing method, data screening method and data compression method described above.

[0280] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned data processing method, data screening method and data compression method when executed by a processor.

[0281] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of the computer program and the technical schemes of the above-mentioned data processing method, data screening method and data compression method belong to the same concept, and the details not described in detail in the technical scheme of the computer program can be referred to the description of the technical schemes of the above-mentioned data processing method, data screening method and data compression method.

[0282] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0283] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0284] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0285] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0286] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: Acquire reference time series data information and a reference time point corresponding to the reference time series data information, and acquire a reference data processing window corresponding to the reference time series data information, and reference starting time series data information and a reference starting time point corresponding to the reference data processing window, wherein the time series data information is monitoring data in a monitoring system; Determine a reference data time difference based on the reference time point and the reference starting time point; Obtain the predecessor reference timing data information and the predecessor reference time point and the predecessor window data information range corresponding to the predecessor reference timing data information, wherein the predecessor reference timing data information is the data corresponding to the last valid sampling point in the reference data processing window, and the predecessor window data information range is the upper and lower bounds generated according to the predecessor reference timing data information and the reference starting timing data information; based on the predecessor reference time point, the reference starting time point, the predecessor window data information range and the reference starting timing data information, generate data change characteristic information; based on the data change characteristic information and the reference data time difference, generate the window data information range corresponding to the reference timing data information, wherein the window data information range is the upper and lower bounds of the data value defined for each data processing window in the timing data information processing; When the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to perform the step of obtaining the reference time series data information and the reference time point corresponding to the reference time series data information; In the case that the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target timing data information is acquired based on the reference time point, and the target timing data information and the reference starting timing data information are stored; the reference timing data information is determined to be the reference starting timing data information of the next reference data processing window, and the steps of acquiring the reference timing data information and the reference time point corresponding to the reference timing data information are continued, wherein the target timing data information is representative data information in the reference data processing window.

2. The method according to claim 1, obtaining reference time series data information and a reference time point corresponding to the reference time series data information, comprising: Acquire initial time series data information and an initial time point corresponding to the initial time series data information; Generate a reference time series data information set based on the initial time series data information and an initial time point corresponding to the initial time series data information; Reference time series data information and a reference time point corresponding to the reference time series data information are determined in the reference time series data information set.

3. The method according to claim 2, wherein obtaining the initial time series data information and the initial time point corresponding to the initial time series data information comprises: Receiving initial time series data information and initial time point acquired by a data acquisition device; or, The initial time series data information and the initial time point are obtained from the initial time series database.

4. The method according to claim 2, generating a reference time series data information set based on the initial time series data information and the initial time point corresponding to the initial time series data information, comprising: Acquire a data denoising window corresponding to the initial time series data information, and acquire initial starting time series data information and an initial starting time point corresponding to the data denoising window; Acquire an initial data time difference based on the initial time point and the initial starting time point, and acquire a denoising window data information range corresponding to the initial time series data information based on the initial starting time series data information and the initial time series data information; When the initial time series data information belongs to the window data information range and the initial data time difference is less than or equal to the preset data denoising window time threshold, continue to perform the step of acquiring the initial time series data information and the initial time point corresponding to the initial time series data information; In the case that the initial timing data information does not belong to the window data information range or the initial data time difference is greater than the data denoising window time threshold, the initial predecessor timing data information is obtained based on the initial time point, and the initial predecessor timing data information and the initial starting timing data information are stored in a reference timing data information set; the initial timing data information is determined to be the initial starting timing data information of the next data denoising window, and the steps of obtaining the initial timing data information and the initial time point corresponding to the initial timing data information are continued.

5. The method according to claim 4, obtaining the denoising window data information range corresponding to the initial time series data information based on the initial starting time series data information and the initial time series data information, comprising: Obtaining initial time series data upper limit information and initial time series data lower limit information; Based on the initial time series data upper limit information and the initial time series data lower limit information, obtaining the initial time series data range information; A denoising window data information range of the data denoising window is generated according to the initial starting time series data information and the range information of the initial time series data.

6. The method according to claim 1, generating data change characteristic information based on the predecessor reference time point, the reference start time point, the predecessor window data information range and the reference start time series data information, comprising: Generate time point change information based on the previous reference time point and the reference start time point; Generate data change information according to the preceding window data information range and the reference starting time sequence data information; Data change characteristic information is generated based on the time point change information and the data change information.

7. The method according to claim 1, generating a window data information range corresponding to the reference time series data information based on the data change characteristic information and the reference data time difference, comprising: Generate an intermediate window data information range according to the data change characteristic information and the reference data time difference, wherein the intermediate window data information range is generated according to a linear mapping or a nonlinear rule between the data change characteristic information and the reference data time difference; Obtaining reference time series data upper limit information and reference time series data lower limit information, and obtaining reference time series data range information based on the reference time series data upper limit information and the reference time series data lower limit information; A window data information range corresponding to the reference time series data information is generated according to the intermediate window data information range and the reference time series data range information.

8. A data screening method comprising: Receive reference time series data information and a reference time point corresponding to the reference time series data information, and obtain a reference data processing window corresponding to the reference time series data information, as well as reference starting time series data information and a reference starting time point corresponding to the reference data processing window, wherein the time series data information is monitoring data in a network cluster monitoring system; Determine a reference data time difference based on the reference time point and the reference starting time point; Obtain the predecessor reference timing data information and the predecessor reference time point and the predecessor window data information range corresponding to the predecessor reference timing data information, wherein the predecessor reference timing data information is the data corresponding to the last valid sampling point in the reference data processing window, and the predecessor window data information range is the upper and lower bounds generated according to the predecessor reference timing data information and the reference starting timing data information; based on the predecessor reference time point, the reference starting time point, the predecessor window data information range and the reference starting timing data information, generate data change characteristic information; based on the data change characteristic information and the reference data time difference, generate the window data information range corresponding to the reference timing data information, wherein the window data information range is the upper and lower bounds of the data value defined for each data processing window in the timing data information processing; When the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to perform the step of obtaining the reference time series data information and the reference time point corresponding to the reference time series data information; In the case that the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target timing data information is acquired based on the reference time point, and the target timing data information and the reference starting timing data information are stored in a data information screening result set; the reference timing data information is determined to be the reference starting timing data information of the next reference data processing window, and the steps of receiving the reference timing data information and the reference time point corresponding to the reference timing data information are continued, wherein the target timing data information is representative data information in the reference data processing window.

9. A data compression method, comprising: Obtain reference time series data information and a reference time point corresponding to the reference time series data information in the time series data set to be compressed, and obtain a reference data processing window corresponding to the reference time series data information, and reference starting time series data information and a reference starting time point corresponding to the reference data processing window, wherein the time series data information is monitoring data in a network cluster monitoring system; Determine a reference data time difference based on the reference time point and the reference starting time point; Obtain the predecessor reference timing data information and the predecessor reference time point and the predecessor window data information range corresponding to the predecessor reference timing data information, wherein the predecessor reference timing data information is the data corresponding to the last valid sampling point in the reference data processing window, and the predecessor window data information range is the upper and lower bounds generated according to the predecessor reference timing data information and the reference starting timing data information; based on the predecessor reference time point, the reference starting time point, the predecessor window data information range and the reference starting timing data information, generate data change characteristic information; based on the data change characteristic information and the reference data time difference, generate the window data information range corresponding to the reference timing data information, wherein the window data information range is the upper and lower bounds of the data value defined for each data processing window in the timing data information processing; When the reference time series data information belongs to the window data information range and the reference data time difference is less than or equal to the preset data processing window time threshold, continue to perform the step of obtaining the reference time series data information and the reference time point corresponding to the reference time series data information; In the case that the reference timing data information does not belong to the window data information range or the reference data time difference is greater than the data processing window time threshold, the target timing data information is obtained based on the reference time point, and the target timing data information and the reference starting timing data information are stored in a target timing data set; the reference timing data information is determined to be the reference starting timing data information of the next reference data processing window, and the steps of receiving the reference timing data information and the reference time point corresponding to the reference timing data information are continued until all the reference timing data information in the timing data set to be compressed are processed, and a data compression result is generated based on the target timing data set, wherein the target timing data information is representative data information in the reference data processing window.

10. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

11. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.

12. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Electric energy meter safe electricity utilization monitoring method based on Internet of Things

    CN119179042A

  • Index anomaly analysis method and apparatus, and electronic device and storage medium

    WO2021212756A1