Adaptive sampling and approximate query method for streaming data
By employing adaptive sampling and approximate query methods, the problems of low efficiency, poor real-time performance, and large sampling errors in streaming data processing are solved, achieving efficient, real-time, and accurate data processing and querying to meet diverse business needs.
Patent Information
- Application Number
- CN202510802141.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Traditional methods struggle to effectively address the high efficiency, real-time performance, and accuracy requirements of streaming data, particularly in terms of large data volumes, real-time query response latency, and sampling error control.
An adaptive sampling and approximate query method is adopted, which uses techniques such as weight allocation based on data features, machine learning anomaly detection, dynamic adjustment of sampling rate, bucket sampling, construction of data summary structure and deep learning result correction to achieve efficient processing and accurate querying of streaming data.
It improves the efficiency of streaming data processing, reduces resource consumption, ensures real-time query response, reduces sampling errors, and achieves a balance between the accuracy and real-time performance of query results, adapting to the complex needs of different application scenarios.
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of streaming data, in particular to an adaptive sampling and approximate query method for streaming data. BACKGROUND
[0002] With the rapid development of information technology, the amount of data generated by various application systems is growing exponentially, and the speed of data generation is also getting faster, and streaming data has emerged as the times require. Streaming data has the characteristics of continuous data arrival, large data volume, and real-time data processing, etc. For example, in the Internet field, the real-time access log of a website may generate millions of records per second; in the industrial Internet of Things, sensors collect real-time operation data of equipment, and data sources continuously flow into the system.
[0003] Traditional data processing methods cannot meet the processing needs of streaming data, and the following problems exist:
[0004] 1. Large data volume leads to low processing efficiency: Due to the continuous and high-speed generation of streaming data, if all data is processed, the system needs to have extremely high computing and storage capabilities, which is often too costly and difficult to achieve in practical applications. For example, a large e-commerce platform may generate tens of thousands of transaction data per second during a promotion event. If all data is analyzed in detail, the existing server resources will soon be exhausted, causing the system to crash and unable to respond to user queries and business decision-making needs in a timely manner.
[0005] 2. Real-time query response delay: In real-time application scenarios such as financial transaction monitoring and network traffic real-time analysis, users expect to quickly obtain query results in order to make timely decisions. However, processing complex queries on large-scale streaming data requires a lot of time, and traditional exact query methods cannot meet the real-time requirements. For example, in the financial market, stock transaction data changes in real time, and traders need to understand the transaction trends and abnormal situations of specific stocks in real time. If the query response is delayed for several seconds, the trader may miss the best trading opportunity, causing significant economic losses.
[0006] 3. Sampling error is difficult to control: When sampling streaming data, if simple random sampling or fixed interval sampling methods are used, it may not accurately reflect the overall characteristics of the data, resulting in large sampling errors. For example, when monitoring urban traffic flow, if fixed interval sampling is performed on a specific time period or a specific road segment, some sudden traffic congestion or traffic flow changes on special road segments may be ignored, making it difficult to accurately assess the overall traffic conditions of the city and make reasonable traffic management strategies for traffic management departments.
[0007] 4. The accuracy and real-time performance of query results are difficult to balance: In order to improve the accuracy of query results, more data often needs to be processed, which increases the processing time and reduces the real-time performance; while too much pursuit of real-time performance, the use of simple approximation method, may lead to large deviation of query results. For example, in the hot topic monitoring of social media platform, if you want to accurately analyze the heat trend and spread range of the topic, you need to deeply analyze the data of a large number of users, such as articles, comments and other data, but this may lead to the lag of analysis results, and cannot timely capture the instantaneous change of topic heat; if too simple algorithm is used to quickly get the result, some important information may be missed, and the real heat and influence of the topic cannot be accurately reflected.
[0008] Based on the above, therefore, the application provides a method for adaptive sampling and approximate query of streaming data. SUMMARY
[0009] To solve the above technical problems, according to one aspect of the application, the application provides the following technical scheme:
[0010] The method for adaptive sampling and approximate query of streaming data includes the following specific steps:
[0011] S1, weight distribution based on data characteristics: analyze the characteristics of streaming data, and assign corresponding weights to each data point; give higher weight to data points with rapid changes and large fluctuations; give lower weight to data points with gentle changes;
[0012] S2, anomaly detection mechanism based on machine learning: train an anomaly detection model using historical data, and perform real-time anomaly detection on streaming data;
[0013] S3, dynamically adjust the sampling rate: dynamically adjust the sampling rate according to the weight distribution of the data and the system resource status; when there are more high-weight data points in the data, appropriately increase the sampling rate to ensure that key information can be accurately captured; when the data is relatively stable and the proportion of low-weight data points is large, reduce the sampling rate to reduce the data processing amount; at the same time, monitor the CPU and memory resource usage of the system in real time to avoid system resource exhaustion due to too high sampling rate;
[0014] S4. Bucket sampling: divide the streaming data into buckets according to time window or data characteristics, and sample the data in each bucket independently; in each bucket, according to the characteristics of the data in the bucket and the preset sampling strategy, select the appropriate sampling method to ensure the sampling accuracy, improve the sampling efficiency, and facilitate the separate analysis of data in different time periods or with different characteristics;
[0015] S5, constructing a data summary structure: on the basis of the sampled data, a data summary structure is constructed, the summary structure can store main features of the data with a small space overhead, and is used for quickly answering approximate queries;
[0016] S6, query processing based on the summary structure: when a query request is received, query processing is first performed on the data summary structure, and approximate results are quickly calculated by using characteristics of the summary structure according to the type and conditions of the query;
[0017] S7, result correction based on deep learning: a deep learning model is constructed, the query conditions, the approximate results obtained by the original summary structure query, and context information of the data are taken as model inputs, and the corrected query results are output;
[0018] S8, error estimation and feedback adjustment: while returning the approximate query results, the error of the results is estimated; the error range is given by comparing with part of the original data or by a statistical model-based method; if the user is not satisfied with the error, the sampling strategy or the query algorithm can be automatically adjusted according to the error feedback, and the query is re-performed to improve the accuracy of the results.
[0019] As a preferred scheme of the adaptive sampling and approximate query method of the stream data, the specific steps of S2 are as follows:
[0020] S21, data preparation stage: first, the data is collected, then the data is cleaned, then the effective features related to anomaly detection are extracted, and then the data is labeled;
[0021] S22, model selection and training stage: first, the model is selected, then the model is trained, and after the training, the model is evaluated;
[0022] S23, real-time detection stage: first, the data is accessed in real time, and then the real-time prediction is performed;
[0023] S24, result processing stage: first, the anomaly is labeled and recorded, and then the anomaly alarm is performed.
[0024] As a preferred scheme of the adaptive sampling and approximate query method of the stream data, the specific steps of S21 are as follows:
[0025] S211, data collection: historical data is collected from a source of stream data;
[0026] S212, data cleaning: the collected historical data is cleaned to remove noise data, duplicate data and incomplete data;
[0027] S213, feature engineering: analyze the raw data and extract effective features related to anomaly detection;
[0028] S214, data labeling: for supervised anomaly detection algorithms, label the data according to specific rules to determine which data belongs to normal data and which data belongs to abnormal data; for unsupervised anomaly detection algorithms, record the real abnormal situation in the data preparation stage.
[0029] As a preferred scheme of the adaptive sampling and approximate query method of the stream data of the application, the specific steps of S22 are as follows:
[0030] S221, model selection: select a suitable machine learning anomaly detection model according to data characteristics and business requirements;
[0031] S222, model training: divide the prepared data into training set and test set, then use the training set to train the selected model, and adjust the parameters of the model during the training process to optimize the performance of the model, and through multiple iterations of training, the model can learn the distribution mode and feature rule of normal data;
[0032] S223, model evaluation: use the test set to evaluate the trained model, and use accuracy, recall rate, F1 value and AUC index to measure the performance of the model; if the performance of the model does not meet the requirements, return to adjust the parameters of the model or reselect the model until the satisfactory effect is achieved.
[0033] As a preferred scheme of the adaptive sampling and approximate query method of the stream data of the application, the specific steps of S23 are as follows:
[0034] S231, real-time data access: access the real-time generated stream data to the trained anomaly detection model to ensure that the format and features of the data are consistent with the training data;
[0035] S232, real-time prediction: the model predicts the accessed real-time data to determine whether the data is abnormal data; and according to the output result of the model, determine whether each data point is an abnormal point.
[0036] As a preferred scheme of the adaptive sampling and approximate query method of the stream data of the application, the specific steps of S24 are as follows:
[0037] S241, anomaly labeling and recording: label and record the relevant information of the detected abnormal data points, and then store the abnormal information in the database for subsequent analysis and tracing;
[0038] S242, Abnormal alarm: according to the preset alarm rule, when the abnormal data is detected, the alarm mechanism is triggered.
[0039] As a preferred scheme of the adaptive sampling and approximate query method of the stream data, wherein the specific steps of S7 are as follows:
[0040] S71, data preparation stage: first, data integration is performed, then data preprocessing is performed, and then data annotation and label generation are performed;
[0041] S72, model construction and training stage: first, model selection and architecture design are performed, then model training is performed, and after training, model optimization and evaluation are performed;
[0042] S73, real-time data input: when receiving new approximate query results, the query conditions, approximate results and related context data are arranged according to the data format of the training stage, and are input into the trained deep learning model in real time;
[0043] S74, correction result generation: the model performs forward propagation calculation according to the input data, and outputs the correction value or the corrected result of the approximate query result;
[0044] S75, correction effect evaluation: the corrected result is compared with the actual accurate result, the same performance index as the training stage is used to evaluate the correction effect; at the same time, the application effect of the correction result in the actual business scenario is analyzed;
[0045] S76, continuous optimization: according to the evaluation result, if the correction effect is not ideal, more data can be collected to retrain the model, or the model architecture can be fine-tuned, so as to continuously optimize the result correction model based on deep learning, and improve the accuracy of the approximate query result.
[0046] As a preferred scheme of the adaptive sampling and approximate query method of the stream data, wherein the specific steps of S71 are as follows:
[0047] S711, data integration: collecting multi-source data related to approximate query, including original stream data, approximate query results generated based on summary structure, query condition information and context information of data;
[0048] S712, data preprocessing: cleaning the integrated data, removing error values and missing values, and processing missing data by interpolation or filling method; at the same time, normalizing or standardizing the data, converting different scales and different distributions of data into a unified format for deep learning model learning;
[0049] S713, data labeling and label generation: for data samples with clear and accurate results, they are used as labeled data, and the difference between the approximate query result and the accurate result is used as the label for model training.
[0050] As a preferred scheme of the adaptive sampling and approximate query method of the stream data of the application, wherein: the specific steps of S72 are as follows:
[0051] S721, model selection and architecture design: according to the data characteristics and query task requirements, select a suitable deep learning model architecture;
[0052] S722, model training: divide the preprocessed data into training set, validation set and test set, then use the training set to train the model, and in the training process, define a suitable loss function, and update the parameters of the model through the back propagation algorithm;
[0053] S723, model tuning and evaluation: use the validation set to evaluate the model during training to monitor the performance indicators of the model, and optimize the performance of the model by adjusting the hyperparameters of the model.
[0054] Compared with the prior art:
[0055] 1. For the problem of low processing efficiency caused by too large data volume:
[0056] 1.1 Reduce resource consumption: by setting weight distribution based on data characteristics, dynamically adjusting sampling rate and bucket sampling, it can reduce data processing volume according to data weight and system resource status, without processing all stream data, greatly reducing the demand for computing and storage capacity, improving processing efficiency, avoiding server resource exhaustion, saving cost, and ensuring system stable operation.
[0057] 2. For the problem of real-time query response delay:
[0058] 2.1 Fast result generation: by constructing data summary structure, approximate results can be quickly calculated based on these structures when receiving query requests. Compared with processing original large-scale data, the query processing time is greatly shortened, which can realize real-time or quasi-real-time response, meet the requirements of financial transaction monitoring, network traffic real-time analysis and other scenarios for fast query results, and help users make timely decisions.
[0059] 2.2 Dynamic optimization mechanism: through error estimation and feedback adjustment, the sampling strategy or query algorithm can be automatically optimized according to the user's feedback on the query result error, continuously improving the query efficiency under the premise of ensuring a certain accuracy, and ensuring that the query response speed always meets the real-time application requirements.
[0060] 3. For the problem of difficult to control sampling error:
[0061] 3.1 Accurate reflection of data characteristics: Based on the weight distribution of data characteristics, the data points with sharp changes and large fluctuations can be highlighted, so that the sampling pays more attention to the key information; the bucket sampling selects appropriate sampling methods according to the data characteristics in different buckets, and the combination of the two can more accurately reflect the overall characteristics of the data, effectively reduce the sampling error, and improve the accuracy and reliability of data sampling.
[0062] 3.2 Error quantification and adjustment: Through error estimation and feedback adjustment, the error range can be given by comparing with the original data or statistical model, and the sampling and query process can be automatically optimized according to user feedback, so that the error is within a controllable range, and the sampled data can provide effective support for subsequent analysis and decision-making.
[0063] 4. For the problem of balancing the accuracy and real-time performance of query results:
[0064] 4.1 Balancing both needs: The present application can reduce the amount of data processing by reasonable sampling, quickly generate approximate results using data summary structure, and ensure real-time performance; at the same time, through error estimation and feedback adjustment mechanism, when the user is not satisfied with the error, the processing process is optimized to improve the accuracy, and a good balance between the accuracy and real-time performance of query results is achieved.
[0065] 4.2 Flexible adaptation to different scenarios: The present application can flexibly adjust the sampling strategy and parameters of the query algorithm according to the characteristics and needs of different application scenarios, such as appropriately increasing the sampling rate in scenarios with high accuracy requirements, and prioritizing query speed in scenarios with strict real-time requirements, to meet the complex needs of diversified business scenarios. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the drawings.
[0067] The present application provides an adaptive sampling and approximate query method for streaming data, including the following specific steps:
[0068] S1, weight distribution based on data characteristics: analyze the characteristics of streaming data, and assign appropriate weights to each data point; give higher weight to data points with sharp changes and large fluctuations; give lower weight to data points with smooth changes;
[0069] S2, anomaly detection mechanism based on machine learning: train an anomaly detection model using historical data, and perform real-time anomaly detection on streaming data;
[0070] The specific steps of S2 are as follows:
[0071] S21, data preparation stage: first, collect data, then clean the data, then extract effective features related to anomaly detection, and then label the data;
[0072] The specific steps of S21 are as follows:
[0073] S211, data collection: collect historical data from the source of streaming data;
[0074] S212, data cleaning: clean the collected historical data to remove noise data, duplicate data and incomplete data;
[0075] S213, feature engineering: analyze the original data and extract effective features related to anomaly detection; for numerical data, the statistical quantities such as mean, variance and standard deviation can be calculated; for time series data, trend features and periodic features can be extracted; for classification data, one-hot encoding and other conversion operations can be performed;
[0076] S214, data labeling: for supervised anomaly detection algorithms, label the data according to specific rules to determine which data belongs to normal data and which data belongs to abnormal data; for unsupervised anomaly detection algorithms, the real abnormal situation needs to be recorded in the data preparation stage;
[0077] S22, model selection and training stage: first, select the model, then train the model, and after training, evaluate the model;
[0078] The specific steps of S22 are as follows:
[0079] S221, model selection: according to the data characteristics and business requirements, select the appropriate machine learning anomaly detection model; common unsupervised models include Isolation Forest, which is suitable for high-dimensional data and can quickly identify abnormal points; One-Class SVM can distinguish normal data from abnormal data by constructing a hyperplane; DBSCAN algorithm based on density judges abnormality by the density of data points; supervised models such as logistic regression and random forest can also achieve accurate anomaly detection with sufficient labeled data; for example, for high-dimensional and complex distributed streaming data, the Isolation Forest model can be considered first;
[0080] S222, model training: divide the prepared data into training set and test set, then use the training set to train the selected model, and adjust the parameters of the model during training to optimize the performance of the model, and through multiple iterations of training, the model can learn the distribution pattern and feature rule of normal data;
[0081] S223, model evaluation: using the test set to evaluate the trained model, and using accuracy, recall, F1 value, AUC index to measure the performance of the model; if the model performance does not meet the requirements, return to adjust the model parameters or reselect the model until the satisfactory effect is achieved;
[0082] S23, real-time detection stage: first, real-time data access, then real-time prediction;
[0083] The specific steps of S23 are as follows:
[0084] S231, real-time data access: real-time streaming data is accessed to the trained anomaly detection model, ensuring that the format and features of the data are consistent with the training data;
[0085] S232, real-time prediction: the model predicts the real-time data accessed, judges whether the data is abnormal data, and determines whether each data point is an abnormal point according to the output result of the model;
[0086] S24, result processing stage: first, abnormal marking and recording, then abnormal alarm;
[0087] The specific steps of S24 are as follows:
[0088] S241, abnormal marking and recording: for the detected abnormal data points, marking and recording relevant information, then storing the abnormal information into the database for subsequent analysis and tracing;
[0089] S242, abnormal alarm: according to the preset alarm rule, when detecting abnormal data, triggering the alarm mechanism; through email, SMS, system pop-up window and other ways, timely inform the relevant personnel, so that they can take measures to handle in time. For example, in network security monitoring, once detecting abnormal network access behavior, immediately send alarm information to the security operation and maintenance personnel, so that they can handle the potential security threat in time;
[0090] By setting the anomaly detection mechanism based on machine learning, it can timely capture the abnormal situation in the data, avoid missing key information due to ignoring abnormal data by conventional sampling strategy; for example, in industrial production, it can quickly find abnormal data of equipment operation, provide more timely and accurate data support for fault early warning, and improve the response ability of the system to sudden situations;
[0091] S3, dynamically adjusting the sampling rate: dynamically adjusting the sampling rate according to the weight distribution of the data and the system resource status; when there are more high-weight data points in the data, appropriately increasing the sampling rate to ensure that key information can be accurately captured; when the data is relatively stable and the proportion of low-weight data points is large, reducing the sampling rate to reduce the data processing amount; at the same time, monitoring the CPU and memory resource usage of the system in real time to avoid exhausting system resources due to too high sampling rate;
[0092] S4. Bucket sampling: dividing the streaming data according to time windows or data characteristics, and independently sampling the data in each bucket; in each bucket, according to the characteristics of the data in the bucket and the pre-set sampling strategy, selecting an appropriate sampling method to ensure sampling accuracy, improve sampling efficiency, and facilitate separate analysis of data in different time periods or with different characteristics;
[0093] S5, constructing a data summary structure: on the basis of the sampled data, a data summary structure is constructed, which can store the main characteristics of the data with small space overhead, and is used for quickly answering approximate queries;
[0094] S6, query processing based on the summary structure: when receiving a query request, first perform query processing on the data summary structure, and according to the type and conditions of the query, use the characteristics of the summary structure to quickly calculate the approximate result; for range queries, the number of data satisfying the conditions can be estimated through the statistical information of the intervals in the histogram; for aggregation queries such as average value and sum, the approximate value can be quickly obtained by using the aggregation information of structures such as wavelet trees;
[0095] S7, result correction based on deep learning: a deep learning model (such as recurrent neural network RNN, long short-term memory network LSTM, etc.) is constructed, and the query conditions, the approximate result obtained by querying the original summary structure, and the context information of the data are used as the input of the model, and the corrected query result is output;
[0096] The specific steps of S7 are as follows:
[0097] S71, data preparation stage: first, data integration, then data preprocessing, and then data annotation and label generation;
[0098] The specific steps of S71 are as follows:
[0099] S711, data integration: collecting multi-source data related to approximate queries, including original streaming data, approximate query results generated based on the summary structure, query condition information, and context information of the data;
[0100] S712, data preprocessing: clean the integrated data, remove the error values and missing values, and process the missing data by interpolation or filling method; at the same time, normalize or standardize the data, and convert the data of different scales and different distributions into a unified format, which is convenient for deep learning model learning;
[0101] S713, data labeling and label generation: for data samples with accurate results, they are used as labeled data, and the difference between approximate query results and accurate results is used as label for model training;
[0102] S72, model construction and training phase: first, model selection and architecture design, then model training, and after training, model optimization and evaluation;
[0103] The specific steps of S72 are as follows:
[0104] S721, model selection and architecture design: according to the data characteristics and query task requirements, select appropriate deep learning model architecture; for time series data query result correction, recurrent neural network (RNN) and its variants long short-term memory network (LSTM) and gated recurrent unit (GRU) can be used to capture the time dependence of data; for multi-dimensional data correction, convolutional neural network (CNN) can be used to extract data features; for complex integrated tasks, Transformer architecture can also be used. For example, when processing the approximate query result correction of stock price time series, LSTM model is selected, and network architecture is constructed by stacking multiple LSTM units;
[0105] S722, model training: divide the preprocessed data into training set, validation set and test set, then use the training set to train the model, and in the training process, define appropriate loss function (such as mean square error loss function for regression task, cross entropy loss function for classification task), and update the parameters of the model through back propagation algorithm;
[0106] S723, model optimization and evaluation: use the validation set to evaluate the model during training to monitor the performance of the model, and optimize the performance of the model by adjusting the hyperparameters of the model;
[0107] S73, real-time data input: when receiving new approximate query results, the query conditions, approximate results and related context data are arranged according to the data format of the training phase, and are input into the trained deep learning model in real time;
[0108] S74, correction result generation: the model performs forward propagation calculation according to the input data, and outputs the correction value or corrected result of the approximate query result;
[0109] S75, correction effect evaluation: compare the corrected result with the actual accurate result, use the same performance index as in the training phase to evaluate the correction effect; at the same time, analyze the application effect of the corrected result in the actual business scenario;
[0110] S76, continuous optimization: according to the evaluation result, if the correction effect is not ideal, more data can be collected to retrain the model, or the model architecture can be fine-tuned, to continuously optimize the deep learning-based result correction model and improve the accuracy of the approximate query result;
[0111] By setting the deep learning-based result correction, the powerful learning ability of deep learning can be used to mine the potential relationship and rule in the data, and the approximate result obtained based on the summary structure can be finely corrected. Under the premise of not significantly increasing the query time, the accuracy of the query result is further improved, the approximate error caused by the complex data features is reduced, and the query result is more in line with the actual demand;
[0112] S8, error estimation and feedback adjustment: while returning the approximate query result, the error of the result is estimated; the error range is given by comparing part of the original data or based on a statistical model; if the user is not satisfied with the error, the sampling strategy or query algorithm can be automatically adjusted according to the error feedback, and the query is performed again to improve the accuracy of the result.
[0113] Although the present application has been described with reference to the embodiments above in the foregoing description, it is to be understood that various modifications can be made without departing from the scope of the present application. In particular, the features of the disclosed embodiments can be used in any combination without departing from the scope of the present application, and the combinations of these features are not exhaustively described in the specification only for the purpose of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for adaptive sampling and approximate query of streaming data, characterized in that, The specific steps include the following: S1, weight distribution based on data characteristics: analyze the characteristics of streaming data and assign corresponding weights to each data point; give higher weight to data points with sharp changes and large fluctuations; give lower weight to data points with smooth changes; S2, anomaly detection mechanism based on machine learning: train an anomaly detection model using historical data to perform real-time anomaly detection on streaming data; S3, dynamically adjust the sampling rate: dynamically adjust the sampling rate according to the weight distribution of the data and the system resource status; when there are more high-weight data points in the data, appropriately increase the sampling rate to ensure that key information can be accurately captured; when the data is relatively stable and the proportion of low-weight data points is large, reduce the sampling rate to reduce the data processing amount; at the same time, monitor the CPU and memory resource usage of the system in real time to avoid system resource exhaustion due to high sampling rate; S4. Bucket sampling: divide the streaming data into buckets according to time windows or data characteristics, and sample the data in each bucket independently; In each bucket, according to the characteristics of the data in the bucket and the pre-set sampling strategy, select the appropriate sampling method to ensure the sampling accuracy, improve the sampling efficiency, and facilitate the separate analysis of data in different time periods or with different characteristics; S5, construct data summary structure: based on the sampled data, construct a data summary structure, which can store the main characteristics of the data with small space overhead, and be used for fast approximate query; S6, query processing based on summary structure: when receiving a query request, first perform query processing on the data summary structure, and according to the type and condition of the query, use the characteristics of the summary structure to quickly calculate the approximate result; S7, result correction based on deep learning: construct a deep learning model, input the query condition, the approximate result obtained by querying the original summary structure, and the context information of the data into the model, and output the corrected query result; S8, error estimation and feedback adjustment: estimate the error of the approximate query result; give the error range by comparing part of the original data or using a statistical model; if the user is not satisfied with the error, the sampling strategy or query algorithm can be automatically adjusted according to the error feedback to re-query and improve the accuracy of the result. 2.The method for adaptive sampling and approximate query of streaming data according to claim 1, characterized in that, The specific steps of S2 are as follows: S21, data preparation stage: first collect data, then clean the collected data, then extract effective features related to anomaly detection, and then perform data labeling; S22, model selection and training stage: first select the model, then train the model, and after training, evaluate the model; S23, real-time detection stage: first perform real-time data access, then perform real-time prediction; S24, result processing stage: first perform anomaly labeling and recording, then perform anomaly alert. 3.The method for adaptive sampling and approximate query of streaming data according to claim 2, characterized in that, The specific steps of S21 are as follows: S211, data collection: collect historical data from the source of streaming data; S212, data cleaning: clean the collected historical data to remove noise data, duplicate data, and incomplete data; S213, feature engineering: analyze the original data and extract effective features related to anomaly detection; S214, data labeling: for supervised anomaly detection algorithms, label the data according to specific rules to determine which data belongs to normal data and which data belongs to abnormal data; for unsupervised anomaly detection algorithms, record the real abnormal situation in the data preparation stage.
4. The method for adaptive sampling and approximate query of streaming data according to claim 2, characterized in that, The specific steps of S22 are as follows: S221, model selection: select a suitable machine learning anomaly detection model according to the data characteristics and business requirements; S222, model training: divide the prepared data into training set and test set, then use the training set to train the selected model, and adjust the parameters of the model during the training process to optimize the model performance, and through multiple iterations of training, the model can learn the distribution pattern and feature rule of normal data; S223, model evaluation: evaluate the trained model using the test set, and use accuracy, recall, F1 value, and AUC indicators to measure the performance of the model; if the model performance does not meet the requirements, return to adjust the model parameters or reselect the model until the desired effect is achieved.
5. The method for adaptive sampling and approximate query of streaming data according to claim 2, characterized in that, The specific steps of S23 are as follows: S231, real-time data access: access the real-time generated streaming data to the trained anomaly detection model to ensure that the data format and features are consistent with the training data; S232, real-time prediction: the model predicts the real-time data accessed and determines whether the data is abnormal; and according to the output result of the model, determines whether each data point is an abnormal point.
6. The method for adaptive sampling and approximate query of stream data according to claim 2, characterized in that, The specific steps of S24 are as follows: S241, abnormal marking and recording: mark the detected abnormal data points and record relevant information, then store the abnormal information in the database for subsequent analysis and tracing; S242, abnormal alarm: according to the preset alarm rule, when abnormal data is detected, trigger the alarm mechanism.
7. The method for adaptive sampling and approximate query of streaming data according to claim 1, characterized in that, The specific steps of S7 are as follows: S71, data preparation stage: first, integrate the data, then perform data preprocessing, and then perform data labeling and label generation; S72, model construction and training stage: first, select the model and design the architecture, then train the model, and after training, optimize and evaluate the model; S73, real-time data input: when receiving new approximate query results, organize the query conditions, approximate results, and related context data according to the data format of the training stage, and input them into the trained deep learning model in real time; S74, correction result generation: the model performs forward propagation calculation based on the input data, and outputs the correction value or corrected result of the approximate query result; S75, correction effect evaluation: compare the corrected result with the actual accurate result, and use the same performance indicators as in the training stage to evaluate the correction effect; at the same time, analyze the application effect of the corrected result in the actual business scenario; S76, continuous optimization: according to the evaluation results, if the correction effect is not ideal, more data can be collected to retrain the model, or the model architecture can be fine-tuned to continuously optimize the deep learning-based result correction model, and improve the accuracy of the approximate query result. 8.The method for adaptive sampling and approximate query of stream data according to claim 7, characterized in that, The specific steps of S71 are as follows: S711, data integration: collect multi-source data related to approximate query, including original streaming data, approximate query results generated based on summary structure, query condition information and data context information; S712, data preprocessing: clean the integrated data, remove error values and missing values, and process missing data through interpolation or padding method; at the same time, normalize or standardize the data, convert different scales and different distributions of data into a unified format for deep learning model learning; S713, data annotation and label generation: for data samples with clear and accurate results, use the difference between approximate query results and accurate results as labels for model training.
9. The method for adaptive sampling and approximate query of stream data according to claim 7, characterized in that, The specific steps of S72 are as follows: S721, model selection and architecture design: select appropriate deep learning model architecture according to data characteristics and query task requirements; S722, model training: divide the preprocessed data into training set, validation set and test set, then use the training set to train the model, and define appropriate loss function during training, and update the parameters of the model through back propagation algorithm; S723, model tuning and evaluation: use the validation set to evaluate the model during training to monitor the performance indicators of the model, and optimize the performance of the model by adjusting the hyperparameters of the model.
Citation Information
Patent Citations
Streaming data-oriented non-repetitive sampling method
CN110609832A
Methods and systems for data collection, learning, and streaming of machine signals for analytics and maintenance using the industrial internet of things
CN112703457A