Model Training Method and Apparatus, Data Cleaning Method and Electronic Device

By performing multi-scale clustering, anomaly detection and multi-feature extraction on large-scale traffic data, we form an enhanced feature set and train an integrated model, which solves the problem of inefficient data cleaning in the existing technology, and achieves more efficient data cleaning and quality improvement.

CN119782831BActive Publication Date: 2025-06-13SHANGHAI DOUXIANG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510293490.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process and improve the cleaning efficiency of large-scale traffic data, and the traditional methods are inefficient in processing.

Method used

By obtaining the traffic data set that is extracted according to the time series, multi-scale clustering, anomaly detection and multi-feature extraction are performed, and the clustering data set, anomaly feature set and target feature set are fused to form an enhanced feature set, which is used to train the integrated model for data cleaning.

Benefits of technology

Improve data quality and availability, enhance data cleaning efficiency, and enable more efficient processing and identification of abnormal information in traffic data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782831B_ABST
    Figure CN119782831B_ABST
Patent Text Reader

Abstract

The present application provides a model training method and apparatus, a data cleaning method, and an electronic device. Among them, the model training method includes: obtaining a first data set; performing multi-scale clustering processing on the first data set to obtain a clustered data set; performing anomaly detection on the first data set to obtain an anomaly feature set; performing multi-feature extraction on the first data set according to a plurality of preset feature types respectively to obtain a target feature set corresponding to each feature type; performing feature fusion on the clustered data set, the anomaly feature set, and each target feature set to obtain an enhanced feature set; and using the enhanced feature set to train an ensemble model to obtain a trained ensemble model. Applying the model training method provided by the present application can not only obtain an ensemble module with higher accuracy, but also improve the efficiency of data cleaning when using this ensemble model for large-scale data cleaning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to a model training method and device, a data cleaning method, and an electronic device. Background Art

[0002] With the rapid development of Internet technology, the generation and collection of large-scale traffic data have become increasingly common. Such data is of great value for aspects such as network management, security analysis, and service quality optimization. However, raw traffic data often has problems such as noise, outliers, and missing values, seriously affecting the quality and usability of the data.

[0003] In the prior art, the way to remove problems such as noise, outliers, and missing values in raw traffic data is to perform data cleaning on the raw traffic data. Traditional data cleaning methods mainly rely on simple statistical techniques and rule filtering. However, for large-scale traffic data, traditional data cleaning methods are difficult to effectively process and have low processing efficiency. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a model training method and device, a data cleaning method, and an electronic device to improve the efficiency of data cleaning.

[0005] The embodiments of this application are implemented as follows:

[0006] In a first aspect, an embodiment of this application provides a model training method, including: obtaining a first data set, where the first data set is a traffic data set obtained by performing feature extraction according to a time series; performing multi-scale clustering processing on the first data set to obtain a clustering data set; performing anomaly detection on the first data set to obtain an anomaly feature set; performing multi-feature extraction on the first data set according to a preset plurality of feature types to obtain a target feature set corresponding to each feature type; fusing the clustering data set, the anomaly feature set, and each target feature set to obtain an enhanced feature set; and training an ensemble model using the enhanced feature set to obtain a trained ensemble model.

[0007] In the embodiment of this application, by capturing the similarity, complexity, and dependency relationships between different features of traffic data in the time dimension, abnormal information of traffic data is identified and processed, and then a high-precision ensemble model is trained. Applying this ensemble model to perform data cleaning on traffic data can not only improve the data quality and usability but also improve the efficiency of data cleaning.

[0008] In an alternative embodiment of the first aspect, the obtaining of the first data set includes: obtaining an initial data set; performing time alignment and segmentation processing on the initial data set according to a time series; constructing an index for the initial data set after the alignment and segmentation processing; and performing feature extraction on the initial data set with the constructed index according to a time series to obtain the first data set.

[0009] In an embodiment of the present application, the initial data set is a traffic data set composed of traffic data obtained from at least one data source. Since each traffic data in the initial data set may come from different data sources and different data sources may have different time granularities, in order to facilitate subsequent unified processing of each traffic data, each traffic data in the initial data set is subjected to time alignment and segmentation processing. The segmented traffic data can reflect the structural change points of the traffic data in the time series. An index is constructed for the initial data set after the alignment and segmentation processing to facilitate the management of each traffic data in the initial data set. Feature extraction is performed on the initial data set with the constructed index according to a time series to obtain the features of the traffic data in different time dimensions.

[0010] In an alternative embodiment of the first aspect, the performing of feature extraction on the initial data set with the constructed index according to a time series to obtain the first data set includes: decomposing the initial data set with the constructed index to obtain a second data set; extracting statistical feature of the second data set within a preset time window to obtain a third data set; performing correlation feature extraction on the third data set based on the statistical feature of the third data set to obtain a correlation feature set; performing time-frequency feature extraction on the correlation feature set and / or the second data set to obtain a time-frequency feature set, where the time-frequency feature is a three-dimensional feature used to characterize time-frequency-scale; performing complexity feature extraction on the time-frequency feature set and / or the second data set to obtain a complexity feature set, where the complexity feature is used to characterize the irregularity degree of the time series; and performing feature fusion on the complexity feature set and / or the time-frequency feature set and / or the correlation feature set to obtain the first data set.

[0011] In an embodiment of the present application, by means of decomposing the traffic data, extracting statistical features, video features, and complexity features, various aspects of the traffic data in the time series are considered. For example, by decomposition, the trend, periodicity, etc. of the traffic data in the time series are analyzed. By extracting statistical features, the local statistical characteristics and correlations of the traffic data are analyzed. By time-frequency feature extraction, the three-dimensional representation of the traffic data in time-frequency-scale is obtained. By complexity feature extraction, the complexity of the traffic data is obtained. After obtaining the complexity feature set, the time-frequency feature set, and the correlation feature set, feature fusion is performed on at least one of the sets to provide a rich information basis for subsequent feature selection, model construction, and training.

[0012] In an alternative embodiment of the first aspect, after fusing the complex feature set, the time-frequency feature set, and / or the correlation feature set to obtain the first data set, the method further includes: performing dimensionality reduction on the first data set, where performing dimensionality reduction on the first data set includes: screening the main features in the first data set to obtain a screened feature set, where the main features are data features highly correlated with the target variable, and the target variable is the main traffic indicator for prediction; applying an autoencoder to perform non-linear dimensionality reduction on the screened feature set to obtain a non-linearly dimensionally reduced feature set; performing high-order tensor decomposition on the non-linearly dimensionally reduced feature set to obtain a structural feature set; performing visualization processing on the structural feature set to obtain a visualized feature set; clustering the visualized feature set to obtain multiple similar feature groups; resampling each of the similar feature groups to obtain multiple sample sets respectively corresponding to each of the similar feature groups; for each sample set corresponding to the same similar feature group, calculating the similarity coefficient of each feature in the similar feature group in different sample sets; filtering the features in each of the similar feature groups with similarity coefficients lower than a preset threshold to obtain the dimensionally reduced first data set.

[0013] In the embodiments of the present application, feature screening of the first data set to obtain important features of traffic data and data reduction can improve the model accuracy in the subsequent model training process. Performing dimensionality reduction on the screened feature set after feature screening can achieve visualization of high-dimensional data and intuitively understand the data distribution and clustering results. The first data set after feature screening and dimensionality reduction not only retains the main information of the original data but also greatly reduces the dimension of the feature space.

[0014] In an alternative embodiment of the first aspect, performing anomaly detection on the first data set to obtain an anomaly feature set includes: applying a preset local outlier factor algorithm to perform anomaly detection on the first data set to obtain a local anomaly detection result; applying a preset seasonal autoregressive integrated moving average model to perform anomaly detection on the first data set to obtain a context anomaly detection result; applying a preset pruning exact linear time algorithm to perform anomaly detection on the first data set to obtain a structural change detection result; applying a preset variational autoencoder to perform anomaly detection on the first data set to obtain a reconstruction anomaly detection result; applying a preset online anomaly detection algorithm to perform anomaly detection on the first data set to obtain a dynamic anomaly detection result; performing evidence synthesis and uncertainty quantification processing on at least one of the local anomaly detection result, the context anomaly detection result, the structural change detection result, the reconstruction anomaly detection result, and the dynamic anomaly detection result to obtain an anomaly feature set.

[0015] In the embodiments of the present application, different algorithms are used to perform anomaly detection on the first data set, and evidence synthesis and uncertainty quantification processing are performed according to the results of each anomaly detection to obtain an anomaly feature set. In the subsequent model training process, the model can better learn the patterns and rules of abnormal data, and at the same time reduce the model's dependence on noise and irrelevant information, improving the generalization ability of the model.

[0016] In an alternative embodiment of the first aspect, training the ensemble model using the enhanced feature set to obtain a trained ensemble model includes: splitting the enhanced feature set into a training set, a test set, and a validation set; using the training set to construct a univariate model and a multivariate model; using the validation set to perform grid search and time series cross-validation on the univariate model and the multivariate model to optimize the model parameters of the univariate model and the multivariate model according to the validation results; using the test set to perform performance testing on the optimized univariate model and the optimized multivariate model; and synthesizing the optimized univariate model and the optimized multivariate model to obtain a trained ensemble model when the performance testing is passed.

[0017] In the embodiments of the present application, after constructing a univariate model and a multivariate model using the training set, the validation set and the test set are used to optimize and perform performance testing on the univariate model and the multivariate model respectively to improve the accuracy of the ensemble model. Using the high-precision ensemble model to perform data cleaning on the traffic data can not only improve the data quality and usability, but also improve the efficiency of data cleaning.

[0018] In a second aspect, the present application provides a data cleaning method, including: obtaining an original traffic data set; inputting the original traffic data set into a trained ensemble model to obtain a similar data set output by the ensemble model, where the ensemble model is trained by the above model training method; and the similar data set is obtained by deleting or filling the missing values in the original traffic data set, and / or formatting or replacing the outlier values in the original traffic data set.

[0019] In the embodiments of the present application, after obtaining the original traffic data set, only by using the trained ensemble model to perform data cleaning on each traffic data in the original traffic data set, a similar traffic data set similar to the original traffic data set after cleaning can be obtained, which can not only improve the quality and usability of the traffic data, but also improve the efficiency of data cleaning.

[0020] In an optional embodiment of the second aspect, it further includes: smoothing the similar data set to obtain a smoothed data set; marking the prominent data in the smoothed data set to obtain a marked data set; obtaining the data processing log corresponding to the marked data set; updating the difference information to the data processing log to obtain the updated data processing log, where the difference information is the difference information obtained by comparing the similar data set with the original traffic data set; and obtaining the corrected similar data set based on the updated data processing log and the smoothed data set, where the corrected similar data set is a smoothed data set carrying at least the data processing history record and the difference information in the updated data processing log.

[0021] In the embodiments of the present application, after obtaining the similar data set, in order to further improve the data quality, the similar data set can be further smoothed to determine whether it contains prominent data. Mark the prominent data so that the marked data set can quickly locate the prominent data. According to the difference information between the similar data set and the original traffic data set, update the data processing log corresponding to the marked data set, and obtain the corrected similar data set based on the updated data processing log and the smoothed data set. Through the corrected similar data set, not only the high-quality traffic data after data cleaning can be retained, but also the detailed data processing history record and difference information are included, which is convenient for subsequent analysis and traceability of the traffic data.

[0022] In a third aspect, the present application provides a model training device, including: an acquisition module, a clustering module, a detection module, a multi-feature extraction module, a feature fusion module, and a training module; the acquisition module is used to acquire a first data set, where the first data set is a traffic data set obtained by performing feature extraction according to a time series; the clustering module is used to perform multi-scale clustering processing on the first data set to obtain a clustering data set; the detection module is used to perform anomaly detection on the first data set to obtain an anomaly feature set; the multi-feature extraction module is used to perform multi-feature extraction on the first data set according to a preset plurality of feature types respectively to obtain a target feature set corresponding to each feature type; the feature fusion module is used to perform feature fusion on the clustering data set, the anomaly feature set, and each target feature set to obtain an enhanced feature set; and the training module is used to train an ensemble model by applying the enhanced feature set to obtain a trained ensemble model.

[0023] In a fourth aspect, the present application provides an electronic device, including: a memory and a processor, where the processor is connected to the memory; the memory is used to store a program; and the processor is used to call the program stored in the memory to execute the above data training method or the above data cleaning method.

[0024] Other features and advantages of the present application will be described in the subsequent specification. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification and the accompanying drawings. Brief Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings. As shown in the accompanying drawings, the above and other objectives, features, and advantages of the present application will become clearer.

[0026] Figure 1 The flowchart showing a model training method provided by an embodiment of the present application;

[0027] Figure 2 The flowchart showing the dimensionality reduction process provided by an embodiment of the present application;

[0028] Figure 3 The structural diagram related to the anomaly detection process provided by an embodiment of the present application;

[0029] Figure 4 The flowchart showing a data cleaning method provided by an embodiment of the present application;

[0030] Figure 5 The device structure diagram showing a model training device provided by an embodiment of the present application;

[0031] Figure 6 The device structure diagram showing a data cleaning device provided by an embodiment of the present application;

[0032] Figure 7 The structural diagram showing an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0033] The following will describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. The following embodiments can be used as examples to more clearly illustrate the technical solutions of the present application, but cannot be used to limit the protection scope of the present application. Those skilled in the art can understand that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0034] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, relational terms such as "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus.

[0035] Furthermore, the term "and / or" in the present application is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0036] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, the technical term "connection" may be a direct connection or an indirect connection through an intermediate medium.

[0037] Please refer to Figure 1 , Figure 1 which is a method flow chart of a model training method provided by an embodiment of the present application. This model training method can be applied to processors of various system platforms, computer terminals or various mobile devices. The method may include the following steps:

[0038] S101: Obtain a first data set.

[0039] S102: Perform multi-scale clustering processing on the first data set to obtain a clustered data set.

[0040] S103: Perform anomaly detection on the first data set to obtain an anomaly feature set.

[0041] S104: Perform multi-feature extraction on the first data set according to a plurality of preset feature types respectively to obtain a target feature set corresponding to each feature type.

[0042] S105: Perform feature fusion on the clustered data set, the anomaly feature set and each target feature set to obtain an enhanced feature set.

[0043] S106: Use the enhanced feature set to train an ensemble model to obtain a trained ensemble model.

[0044] In the model training method provided by the embodiments of the present application, in order to improve the cleaning efficiency of large-scale traffic data, a first data set is obtained, and the first data set is respectively subjected to multi-scale clustering, anomaly detection, and multi-feature extraction according to multiple feature types to obtain a clustering data set, an anomaly feature set, and multiple target feature sets. Each set is fused to obtain an enhanced feature set, and the enhanced feature set is applied to train a model to obtain a trained integrated model, and then the large-scale traffic data is cleaned by the integrated model.

[0045] The above steps are described in detail below.

[0046] S101: Obtain a first data set.

[0047] Among them, the first data set is a traffic data set obtained by performing feature extraction according to a time series. Each traffic data in the traffic data set is traffic data obtained from multiple data sources, and each data collection point may include a network traffic monitoring device, a log server, etc. The collected traffic data has time continuity, and the data types of the traffic data cover IP addresses, timestamps, packet sizes, protocol types, etc.

[0048] When collecting traffic data from a data source, in order to ensure that the collected traffic data has time continuity and representativeness, the data collection frequency and sampling strategy corresponding to each data source can be set. For example, for high-traffic network nodes, it may be necessary to sample multiple times per second, while for low-traffic nodes, it can be sampled once per minute.

[0049] In the embodiments of the present application, the obtained traffic data set is used as an initial data set, and the initial data set is processed to obtain a first data set. Among them, the specific process of obtaining the first data set is as follows: obtain the initial data set; perform time alignment and segmentation processing on the initial data set according to a time series; construct an index for the initial data set that has been aligned and segmented; perform feature extraction on the initial data set with the constructed index according to a time series to obtain a first data set.

[0050] Optionally, before performing time alignment and segmentation processing on the initial data set, the initial data set can also be subjected to format unification processing to obtain an initial data set with a unified format. The format unification processing includes at least one of the unification of time stamp formats (such as converting to UTC time), the unification of data units (such as converting all traffic data to bytes / second), and the unification of data types (such as converting all numerical data to floating-point numbers).

[0051] In this application, since the data in the initial dataset may come from different data sources and the time granularities of the data sources are different, interpolation methods can be applied to align the time of the initial dataset, unifying traffic data with different time granularities or different timestamps to the same time point. For example, cubic spline interpolation can be applied to achieve time alignment; if the initial dataset belongs to a multivariate time series, the DTW (Dynamic Time Warping) algorithm can be used to handle the time alignment problem between different variables. In a time series, data in different time periods may have different statistical characteristics (such as mean, variance, trend, etc.). By finding the structural change points and segmenting the data, for example, the PELT (Pruned Exact Linear Time) algorithm can be adopted. This algorithm can find the optimal segmentation by minimizing the following cost function. The PELT algorithm is:

[0052] ;

[0053] where, is the minimum segmentation cost up to time point t, is the cost function of the subsequence from s + 1 to t, is the complexity penalty term.

[0054] After time alignment and segmentation of the initial dataset, in order to ensure the data, the initial dataset after alignment and segmentation processing can be saved to a dedicated time series database in a distributed storage system (such as InfluxDB (Influx Database), TimescaleDB (Timescale Database), etc.). And in order to facilitate quick query of traffic data within a specific time range or for a specific device, an index is constructed for the initial dataset after alignment and segmentation processing. When querying data in the distributed storage system, queries can be made through the timestamp and / or the device information of the data source to which the traffic data belongs. Among them, the process of time alignment, segmentation, and index construction of the initial dataset is a preprocessing process. After preprocessing the initial dataset, feature extraction is performed on the preprocessed initial dataset (i.e., the initial dataset with an index constructed) according to the time series to obtain the first dataset.

[0055] Optionally, after time alignment, the initial dataset can also be processed for missing values. Using the modified Z-score method, the missing values in the initial dataset are identified and processed to obtain the filled initial dataset. Among them, if the traffic data in the initial dataset belongs to short-term missing, linear interpolation or spline interpolation methods can be used. For long-term missing, a sequence-to-sequence (Seq2Seq) model based on LSTM (Long Short-Term Memory) can be used for missing value filling.

[0056] Based on the preprocessing process of time alignment, segmentation, and index construction of the initial dataset described in the embodiments of the present application, the following application embodiments can be obtained:

[0057] Suppose the traffic data of a large network is being processed. The original data may contain information on millions of packets per second. By setting an appropriate sampling strategy (such as sampling once every 10 seconds), the data volume is reduced to a manageable level. Using the modified Z-score method, it may be found that the traffic at certain time points is extremely high, which may indicate a network attack or device failure. During the process of format unification, it may be necessary to unify the traffic units reported by different devices (such as bits per second and bytes per second) to bytes per second. During the time alignment process, it may be found that the clocks of some devices have slight deviations, and the DTW algorithm needs to be used for correction. For missing traffic data, the LSTM model is used to predict and fill based on historical patterns. When segmenting the time series, the PELT algorithm may find that the network traffic patterns are significantly different on weekdays and weekends, and thus automatically divide the data into weekday segments and weekend segments. Finally, these cleaned and preprocessed data are stored in TimescaleDB, enabling subsequent analysis to quickly query the traffic data for a specific time range or a specific device.

[0058] In the embodiments of the present application, after preprocessing the initial dataset, the process of feature extraction according to the time series can be as follows:

[0059] Decompose the indexed initial dataset to obtain a second dataset; extract the statistical features of the second dataset within a preset time window to obtain a third dataset; based on the statistical features of the third dataset, perform relevant feature extraction on the third dataset to obtain a relevant feature set; perform time-frequency feature extraction on the relevant feature set and / or the second dataset to obtain a time-frequency feature set; perform complexity feature extraction on the time-frequency feature set and / or the second dataset to obtain a complex feature set; fuse the complex feature set and / or the time-frequency feature set and / or the relevant feature set to obtain a first dataset. Among them, the time-frequency feature is a three-dimensional feature used to characterize time-frequency-scale, and the complexity feature is used to characterize the irregularity of the time series.

[0060] Among the various traffic data included in the initially constructed indexed dataset, the time series characteristics of some data may be univariate time series, while those of another part of the data may be multivariate time series. Before decomposing the initially constructed indexed dataset, the initially constructed indexed dataset can also be analyzed. Specifically, it can be dimension detection and correlation analysis of the initially constructed indexed dataset to determine the traffic data of the univariate time series and the traffic data of the multivariate time series in the initially constructed indexed dataset. Among them, dimension detection of the initially constructed indexed dataset refers to checking the number of observations at each time point. For example, a single traffic value is univariate, and multiple measurement indicators are multivariate (variables here such as numerical variables (traffic volume, memory usage rate, etc.), categorical variables (protocol type), etc.). Correlation analysis of the initially constructed indexed dataset is to calculate the correlation coefficient between variables through the mutual information coefficient method to judge the independence and correlation degree of variables. For example, a correlation threshold (such as 0.7) is set, and the dependence relationship between variables is judged based on the threshold.

[0061] For traffic data with different time series characteristics, different decomposition methods are used to decompose it.

[0062] For the traffic data of the univariate time series, the additive model decomposition method is applied, such as the Prophet (time series prediction) algorithm. By fitting the trend, seasonality, and holiday effects, the trend, seasonality, and residual components are obtained. The expression of the Prophet algorithm is:

[0063] ;

[0064] Among them, represents the trend function, represents the periodic change, represents the holiday effect, represents the error term.

[0065] For the traffic data of the multivariate time series, matrix decomposition and singular value decomposition are applied to obtain the multivariate decomposition result. Among them, matrix decomposition is used to handle the linear relationship between multivariate variables; singular value decomposition is used for dimensionality reduction and extraction of the main change patterns. Since the MSSA (singular spectrum analysis) method involves the construction of the trajectory matrix and singular value decomposition, the MSSA method can be applied to implement the process of matrix decomposition and singular value decomposition, and its corresponding expression is:

[0066] ;

[0067] Among them, X is the trajectory matrix, U and V are the left and right singular vector matrices, is the diagonal matrix of singular values, and V T is the transpose of V.

[0068] Since the decomposition methods for the flow data of univariate time series and multivariate time series are different, after obtaining the trend, seasonal, and residual components as well as the multivariate decomposition results, slope calculation, period identification, and statistic extraction are performed on the trend, seasonal, and residual components as well as the multivariate decomposition results to obtain a second data set. Among them, slope calculation can be to calculate the overall change trend of the time series using a linear regression algorithm to identify the rate of increase or decrease of the flow data; Fourier transform can be used for period identification, and the extracted statistics can include mean, variance, and median, etc.

[0069] After obtaining the second data set, a time window can be set for the second data set according to a preset time rule, and statistics are calculated within each time window, and a time series is generated by sliding the time window to capture the local dynamic statistic features, resulting in a corresponding statistical feature set (i.e., the third data set). Then, relevant feature extraction is performed on each statistic and feature in the third data set to obtain a relevant feature set, and the strength of the mutual relationship between the relevant feature sets is characterized to understand the dependence between the features of the flow data. Among them, relevant feature extraction can be to calculate the correlation coefficient between each statistic feature in the third data set and the mutual information based on kernel density estimation using the Pearson correlation coefficient, and use the kernel density estimation method to approximate the probability distribution density function of the variable, thereby calculating the mutual information between two variables to measure the non-linear correlation between variables to capture the possible non-linear dependence relationship between variables. Time-frequency feature extraction is performed on the relevant feature set and / or the second data set to obtain a time-frequency feature set. Among them, through wavelet function convolution and scale transformation, a three-dimensional representation of time-frequency-scale is generated to obtain time-frequency features, which combine the information of the three dimensions of time, frequency, and scale. Finally, sample entropy, approximate entropy, and multi-scale entropy are calculated for the video feature set and / or the second data set to achieve the extraction of the complexity features of the video feature set and / or the second data set, obtaining a complexity feature set. Among them, the main function of sample entropy is to measure the irregularity of the time series, and its corresponding calculation formula is:

[0070] ;

[0071] where m is the pattern length, r is the similarity threshold, N is the time series length, and A(m,r) is the pattern matching count. The main function of approximate entropy is to measure the regularity of the time series; the main function of multi-scale entropy is to analyze complexity on multiple time scales. Each type of entropy provides complexity information in different dimensions, and the results of the three types of entropy complement each other to jointly form a complexity feature set to characterize the irregularity of the time series through the complexity features.

[0072] In an alternative embodiment, after performing feature fusion on the complexity feature set and / or the video feature set and / or the correlation feature set to obtain the first data set, the first data set may also be processed by standardization and normalization to obtain the processed first data set.

[0073] If any one of the complexity feature time-frequency feature sets is a set obtained by directly processing the second data set, the feature fusion process may select at least one of the complexity feature set, the video feature set, and the correlation feature set for feature fusion processing according to business requirements, and perform standardization and normalization processing on the set obtained after feature fusion to obtain the first data set, where the feature fusion process includes at least one of feature splicing, feature weighting, and feature selection. If the time-frequency feature set is a set obtained by performing time-frequency feature extraction on the correlation feature set, and the complexity feature set is a set obtained by performing complexity feature extraction on the time-frequency feature set, then the feature fusion process is a feature selection process. Also, since the time-frequency feature set is obtained by processing the correlation feature set, and the complexity feature set is obtained by processing the video feature set, the first data set obtained by feature selection is the complexity feature set. Performing standardization and normalization processing on the first data set is actually performing standardization and normalization processing on the complexity feature set. Therefore, if the time-frequency feature set is a set obtained by performing time-frequency feature extraction on the correlation feature set, and the complexity feature set is a set obtained by performing complexity feature extraction on the time-frequency feature set, then the feature fusion process may not be performed, and the complexity feature set may be directly processed by standardization and normalization to obtain the processed first data set.

[0074] In an alternative embodiment, after performing feature fusion on the complex feature set and / or the time-frequency feature set and / or the correlation feature set to obtain the first data set, dimensionality reduction processing may also be performed on the first data set; or, after performing feature fusion on the complex feature set and / or the time-frequency feature set and / or the correlation feature set to obtain the first data set, and performing standardization and normalization processing on the first data set, dimensionality reduction processing may also be performed on the processed first data set.

[0075] Reference Figure 2 , performing dimensionality reduction processing on the first data set may include the following steps:

[0076] S201: Screen the main features in the first data set to obtain the screened feature set.

[0077] Among them, the main feature is a data feature that has a high correlation with the target variable, and the target variable is the main traffic metric used for prediction, such as traffic volume, duration, etc.

[0078] In the embodiments of the present application, the first data set is the first data set obtained after performing feature fusion in the above S101, or the first data set obtained after performing feature fusion, normalization, and normalization processing, or the first data set obtained after performing feature fusion, normalization, normalization processing, and dimensionality reduction processing.

[0079] In an alternative embodiment, during the process of screening the main features of the first data set, the preliminary screening of the main features is performed by calculating the correlation between features and evaluating the redundancy in the first data set. After the preliminary screening, further screening of the main features can be performed through random feature permutation and importance comparison, and then the linear correlation between features is quantified to evaluate the multicollinearity of the features. According to the evaluation results, features are gradually added or deleted from the data to obtain a refined feature set. Finally, by means of regularized principal component extraction, important feature combinations in the refined feature set are retained to obtain the screened feature set.

[0080] Among them, the way to perform the preliminary screening of the main features of the first data set can be to use the mRMR (Minimum Redundancy Maximum Relevance) algorithm to calculate the correlation and redundancy between each feature in the first data set, determine the correlation between each feature and the target variable according to the calculation results, retain the features with high correlation with the target variable, and delete the features with high redundancy between features to obtain the initially screened feature set after preliminary screening.

[0081] Apply the Boruta (Boruta Feature Selection) algorithm to perform feature selection on the initially screened feature set. The Boruta algorithm creates "shadow features" and conducts multiple rounds of random forest training with the original features. By comparing the importance scores of the original features and the shadow features, important features in the initially screened feature set are selected according to the scores to obtain the important feature set. Among them, comparing the importance of the original features and the shadow features is not simply comparing with a single shadow feature, but comparing with the MZSA (Maximum Importance Score) among all shadow features. Specifically: The Boruta algorithm first creates shadow features for each original feature, shuffles the shadow features to eliminate their correlation with the target variable, then merges the original features and the shadow features. After the merged features are processed by the model, the importance scores of the original features and the shadow features are calculated. Finally, the features with the importance scores of the original features greater than those of the shadow features are selected as important features. Among them, during the process of comparing the importance scores, this process needs to be repeated in multiple rounds of iterations, and statistical tests are used to evaluate the significance of the importance difference; and considering the relationship between the features and the target variable and the mutual influence between the features, a final decision is made based on the results of multiple rounds to select the important feature set.

[0082] Use VIF (Variance Inflation Factor) to evaluate the multicollinearity of each feature in the important feature set. The obtained VIF value is the evaluation result of the multicollinearity of the feature. Among them, the calculation formula of VIF is:

[0083] ;

[0084] Among them, is the coefficient of determination of the multiple regression of the i-th feature on all other features. The larger the VIF value, the stronger the collinearity between this feature and other features. For the VIF value corresponding to each feature of the traffic data in the important feature set, gradually add new features or delete features to the traffic data. During the process of gradually adding new features or deleting features, provide an empty model, which provides a performance benchmark and can be compared with the benchmark each time a feature is added or deleted to evaluate the VIF value of each feature in the traffic data after adding or deleting the feature. For example: if a feature needs to be deleted, when the VIF values of multiple features are all greater than 10, give priority to deleting the feature with the largest VIF value, and re-evaluate the VIF values of the remaining features after deletion, and proceed step by step until the VIF values of all features are within an acceptable range; if a feature needs to be added, the addition of the new feature must significantly improve the model performance, and the addition cannot cause the VIF value of any feature to exceed the threshold. If the addition causes a significant increase in the VIF value of other features, even if it does not exceed the threshold, it needs to be carefully considered. After gradually adding or deleting features until the model performance no longer improves significantly, obtain a refined feature set.

[0085] Optionally, use the F-test (F-test, joint hypothesis test), or AIC (Akaike Information Criterion) / BIC (Bayesian Information Criterion) to evaluate whether the newly added feature significantly improves the model performance.

[0086] Optionally, when adding or deleting features, give priority to considering the data features related to the target variable.

[0087] The screened feature set obtained by performing regularized principal component extraction on the refined feature set is a linearly reduced feature set. Specifically, apply the Sparse PCA (Sparse Principal Component Analysis) method to perform regularized principal component extraction on the refined feature set to balance the reconstruction error and the sparsity of the loading matrix, reduce the number of features while maintaining the main information, and make the loading matrix of the principal components become sparse. It is not just simply retaining the main component structure, but through an optimization algorithm to obtain a more concise and easier-to-interpret feature combination while maintaining the main information of the data.

[0088] S202: Apply an autoencoder to perform non - linear dimensionality reduction on the screened feature set to obtain a non - linear dimensionality reduction feature set.

[0089] Among them, since the screened feature set obtained by the above S201 is a linearly dimensionality - reduced feature set, in order to capture the non - linear relationships between features, an LSTM autoencoder can be used for non - linear dimensionality reduction. Specifically, the LSTM autoencoder includes an encoder and a decoder. The encoder encodes the feature sequence of the screened feature set, compresses the input feature sequence into a vector of a fixed length, and then the decoder decodes and reconstructs it to reconstruct a new feature sequence from the vector.

[0090] S203: Perform high - order tensor decomposition on the non - linear dimensionality reduction feature set to obtain a structural feature set.

[0091] This application can use Tucker decomposition to perform high - order tensor decomposition on the non - linear dimensionality reduction feature set. The expression of Tucker decomposition is:

[0092] ;

[0093] where X is the original data tensor, G is the core tensor, A, B, and C are factor matrices, represents the n - mode tensor product. Through Tucker decomposition, the structural information of multi - dimensional data can be retained to obtain a structural feature set.

[0094] S204: Perform visualization processing on the structural feature set to obtain a visualization feature set.

[0095] In the embodiments of this application, the t - SNE (t - distributed Stochastic Neighbor Embedding) algorithm is applied to minimize the KL divergence between the similarity of data points in the high - dimensional space and the similarity of corresponding points in the low - dimensional space, realizing the visualization of high - dimensional data to intuitively understand the data distribution and clustering structure.

[0096] S205: Cluster the visualization feature set to obtain multiple similar feature groups.

[0097] In the embodiments of this application, the spectral clustering algorithm is applied to calculate the similarity between each data in the visualization feature set, construct a similarity graph, then calculate the Laplacian matrix of this similarity graph, obtain the eigenvectors of each data through the Laplacian matrix, and cluster each data according to the eigenvectors to obtain multiple similar feature groups. The features of the data in each similar feature group are similar to each other.

[0098] S206: Resample each similar feature group to obtain multiple sample sets corresponding to each similar feature group respectively.

[0099] In the embodiments of the present application, Bootstrap (resampling method) is mainly applied to resample the similar feature groups to obtain multiple sample sets. Specifically, for each similar feature group, at least one different data is repeatedly extracted from the similar feature group as a sample set.

[0100] S207: For each sample set corresponding to the same similar feature group, calculate the similarity coefficient of each feature in the feature similarity group in different sample sets.

[0101] Since there is similarity between the features of each data in the same similar feature group, after obtaining each sample set corresponding to the similar feature group, the features of the sample set are selected to calculate the similarity with the features of other sample sets, and the similarity coefficient of the feature in different sample sets is obtained. Among them, the similarity coefficient is the Jaccard coefficient.

[0102] S208: Filter the features with similarity coefficients lower than the preset threshold in each similar feature group to obtain the first reduced-dimensional dataset.

[0103] In the present application, if the Jaccard coefficient is lower than the preset threshold, it indicates that the similarity of the feature in the group of the same similar feature is low, so it is filtered.

[0104] Optionally, if the Jaccard coefficient difference between the features in any sample set and other sample sets is relatively large, the feature can also be filtered.

[0105] Based on the dimensionality reduction process corresponding to S201 to S208 above, first perform feature screening on the first dataset to streamline each feature in the first dataset, then capture the non-linear relationship between the features, perform encoding and decoding reconstruction on the screened feature set obtained after screening, then perform high-order tensor decomposition to retain the structural information of the multi-dimensional data, then perform visualization processing to map the high-dimensional data to two dimensions, and finally perform clustering and similarity calculation on the visualization feature set and other processing, and finally obtain the first dataset with stable structure and reduced dimensions. Applying the method of S201 to S208 above not only retains the important information of the data, but also reduces the dimensionality of the feature space, facilitating the relevant calculations in the subsequent model training process.

[0106] S102: Perform multi-scale clustering processing on the first dataset to obtain a clustering dataset.

[0107] In the embodiments of the present application, in order to discover the traffic behavior patterns of each traffic data in the first data set at different time scales, the process of performing multi-scale clustering processing on the first data set is as follows: Apply the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm to calculate the density of the data and the optimal number of clusters at different time scales, obtain an initial clustering result, and cluster the first data set according to the initial clustering result to obtain an initial clustering data set. Apply the Gaussian kernel function to process the initial clustering data set to obtain a similarity matrix related to the initial clustering structure. Apply the formula: (Λ is the diagonal matrix of eigenvalues) to perform eigenvalue decomposition on the similarity matrix to determine the optimal hierarchical clustering structure of the initial clustering data set. Calculate the encoding length and select the optimal hierarchy for the hierarchical clustering structure. Using the MDL (Minimum Description Length) principle, select the optimal hierarchy by minimizing the encoding length of the clustering structure, so as to obtain an optimized clustering result. Then, calculate the internal compactness and external separation degree for this optimized clustering result to obtain a clustering quality evaluation report, that is, quantify the clustering quality by using the Silhouette coefficient. The calculation formula of the Silhouette coefficient is:

[0108] ;

[0109] where, is the average distance between data i and other data in the same cluster, is the average distance between data i and the nearest non-cluster sample. The closer the Silhouette coefficient is to 1, the better the clustering effect.

[0110] In the present application, the clustering quality evaluation report is used for the construction of clustering-related features, calculating the weights affecting the Mahalanobis distance and membership degree, and determining the reliability of the clustering center. It can affect the extraction method and weight setting of clustering-related features. When constructing a model, it can be used to guide the sample weight allocation and model parameter adjustment. In the final data cleaning, the evaluation result of the clustering quality evaluation report can be saved as part of the metadata for tracking and verifying the cleaning effect.

[0111] After obtaining the clustering quality evaluation report, extract clustering-related features from the first data set based on the clustering quality evaluation report to form a clustering data set.

[0112] S103: Perform anomaly detection on the first data set to obtain an anomaly feature set.

[0113] ReferenceFigure 3 Schematic structural diagram related to the abnormal detection process shown, the specific implementation process of performing abnormal detection on the first data set to obtain an abnormal feature set is as follows:

[0114] Apply the preset Local Outlier Factor algorithm to perform abnormal detection on the first data set to obtain the local abnormal detection result. Apply the preset Seasonal ARIMA (Autoregressive Integrated Moving Average) model to perform abnormal detection on the first data set to obtain the context abnormal detection result. Apply the preset PELT (Pruned Exact Linear Time) algorithm to perform abnormal detection on the first data set to obtain the structural change detection result. Apply the preset variational autoencoder to perform abnormal detection on the first data set to obtain the reconstruction abnormal detection result. Apply the preset online abnormal detection algorithm to perform abnormal detection on the first data set to obtain the dynamic abnormal detection result. Perform evidence synthesis and uncertainty quantification processing on at least one of the local abnormal detection result, context abnormal detection result, structural change detection result, reconstruction abnormal detection result, and dynamic abnormal detection result to obtain an abnormal detection report; extract features from the abnormal detection report according to the preset abnormal indicators to obtain an abnormal feature set.

[0115] Among them, through the Local Outlier Factor algorithm, by performing local density estimation and calculation of relative density deviation on the first data set, the local abnormal detection result corresponding to the first data set is obtained. Through the Seasonal ARIMA model, seasonal decomposition, autoregressive, and moving average modeling are performed on the first data set to obtain the context abnormal detection result. The process of the PELT algorithm performing abnormal detection on the first data set is as follows: perform dynamic programming segmentation on the first data set through the PELT algorithm to divide the time series into multiple segments with similar characteristics, and detect the change points of the time series by minimizing the cost function to obtain the structural change detection result. The variational autoencoder performs encoding-decoding reconstruction and reconstruction error calculation on the first data set to obtain the reconstruction abnormal detection result. The online abnormal detection algorithm performs real-time data stream processing and dynamic threshold update on the first data set to obtain the dynamic abnormal detection result. Then, through the Dempster-Shafer evidence theory, evidence synthesis and uncertainty quantification are performed on all the obtained abnormal detection results to obtain the final abnormal detection report. Detailed feature extraction is performed on the abnormal detection report according to abnormal indicators such as the occurrence frequency, duration, and severity of abnormal samples to obtain an abnormal feature set.

[0116] Based on the abnormal detection process corresponding to the above S103 embodiment, there are the following specific application scenarios:

[0117] Suppose there is a network traffic dataset containing multiple features such as packet size, protocol type, source and destination IPs, etc. Through the HDBSCAN algorithm, traffic patterns existing at different time scales may be discovered, such as normal traffic during daily working hours, low traffic periods on weekends, etc. Using the LOF algorithm, traffic with abnormally large packet sizes may be discovered, and through the seasonal ARIMA model, traffic at a certain time point may be found to deviate significantly from the expected seasonal pattern. The PELT algorithm may help discover sudden changes in traffic patterns caused by network configuration changes, and complex abnormal patterns that are difficult to describe with simple rules may be discovered through variational autoencoders. Finally, sudden DDoS (Distributed Denial of Service) attacks are detected in real time through an online algorithm. Through the Dempster-Shafer theory, evidence from these different sources is integrated to obtain a comprehensive anomaly detection report, and more detailed anomaly features are obtained through the anomaly detection report, and the anomaly features are extracted to obtain an anomaly feature set.

[0118] S104: Perform multi-feature extraction on the first dataset according to multiple preset feature types to obtain a target feature set corresponding to each feature type.

[0119] Among them, different feature types respectively include: time series features, multi-scale time series similarity features, causal features, dependency structure features, domain-specific advanced statistics, and symbolic dynamics representation and complexity features.

[0120] In this application, statistics, frequency domain features, and information-theoretic metrics are calculated for the first dataset to obtain an initial time series feature set, and then the initial time series feature set is evaluated and screened for feature importance. By calculating the correlation and significance between the features and the target variable, a screened time series feature set is obtained. The time series of each traffic data in the first dataset is aligned and compared at different time scales to extract the similarity features between each data at different time scales, and a similarity feature set is obtained. The multi-head self-attention mechanism is applied to calculate the attention weights and context vectors between different time steps for the first dataset to obtain an attention feature set. Nonlinear causal relationship identification is performed on the first dataset, and by constructing a nonlinear autoregressive model and calculating the prediction error, a causal feature set is obtained. Multivariate dependency structure modeling is performed on the first dataset, and by constructing a binary copula function combination in a tree structure, a dependency structure feature set is obtained. Domain-specific advanced statistics extraction is performed on the first dataset, and by calculating specific metrics such as protocol distribution entropy, connection duration distribution, etc., a domain feature set is obtained. Symbolic dynamics representation and complexity feature extraction are performed on the first dataset, and by converting the time series into a symbolic sequence and calculating the permutation entropy and multi-scale entropy, a dynamic symbol and complexity feature set is obtained.

[0121] Among them, the screened time series feature set, similarity feature set, attention feature set, causal feature set, dependency structure feature set, domain feature set, and dynamic symbol and complexity feature set are all target feature sets corresponding to different feature types.

[0122] S105: Perform feature fusion on the clustering data set, anomaly feature set, and each target feature set to obtain an enhanced feature set.

[0123] In the embodiments of the present invention, the process of feature fusion can be to concatenate the features of all sets, or a complex feature fusion algorithm can be used for feature fusion, such as: feature fusion in deep learning and multimodal feature fusion, etc.

[0124] S106: Use the enhanced feature set to train the ensemble model to obtain a trained ensemble model.

[0125] In the embodiments of this application, the process of training the ensemble model is as follows: split the enhanced feature set into a training set, a test set, and a validation set; use the training set to construct a univariate model and a multivariate model; use the validation set to perform grid search and time series cross-validation on the univariate model and the multivariate model to optimize the model parameters of the univariate model and the multivariate model according to the validation results; use the test set to perform performance testing on the optimized univariate model and the optimized multivariate model; in the case of passing the performance test, synthesize the optimized univariate model and the optimized multivariate model to obtain a trained ensemble model.

[0126] Each traffic data included in the enhanced feature set includes traffic data of univariate time series and traffic data of multivariate time series. After splitting the enhanced feature set into a training set, a test set, and a validation set, each set still includes traffic data of univariate time series and traffic data of multivariate time series. Therefore, when using the training set to construct a univariate model and a multivariate model, specifically, use the traffic data of univariate time series in the training set to construct the univariate model, and use the traffic data of multivariate time series in the training set to construct the multivariate model. Among them, the enhanced feature set is split into a training set, a test set, and a validation set in chronological order, and the data volume of the training set is greater than that of the test set and the validation set (for example, the first 70% of the enhanced feature set is used as the training set, the next 15% is used as the validation set, and the last 15% is used as the test set).

[0127] Specifically, autoregressive modeling and exponential smoothing are performed on the training set to obtain at least one univariate model. Among them, the ARIMA model can be used for autoregressive modeling, and the Holt-Winters method can be applied for exponential smoothing. Vector autoregressive (VAR) model and recursive neural network (such as LSTM) modeling are performed on the training set to obtain at least one multivariate model. Among them, the VAR model can capture the linear dependence relationship between multiple variables, and its general form is:

[0128] ;

[0129] Among them, is a k-dimensional random vector, c is a k-dimensional constant vector, is a k×k coefficient matrix, is a k-dimensional white noise process.

[0130] Grid search and time series cross-validation are applied to at least one univariate model and at least one multivariate model using the validation set to adjust the model parameters of the univariate model and the multivariate model. For example, adjusting the model parameters of the univariate model is to adjust different combinations of p, d, and q values in the ARIMA model; adjusting the model parameters of the multivariate model is to adjust hyperparameters such as the number of layers, the number of neurons, and the learning efficiency in the LSTM.

[0131] Then, the test set should be used to perform performance testing on the univariate model and the multivariate model. If the test is passed, the model synthesis is performed to obtain the trained integrated model. If the test is not passed, the model parameters of the univariate model and the multivariate model are adjusted again according to the above model parameter adjustment method until the test is passed.

[0132] After the univariate model and the variable model that pass the test, the multi-level prediction results of each model are combined, and the Stacking method is used to construct the integrated model. In Stacking, the XGBoost (eXtreme Gradient Boosting) method is used as the meta-learner, and its objective function is:

[0133] ;

[0134] Among them is the training loss function, is the regularization term, and are regularization parameters, T is the number of leaf nodes, and w is the leaf weight vector. This method can effectively combine the advantages of different models and improve the prediction stability.

[0135] Optionally, after obtaining the integrated model, the integrated model can be further enhanced and optimized. The specific enhancement and optimization process is as follows: Adjust the time scale of the input data of the integrated model, such as aggregating hourly data into daily data, and injecting Gaussian noise to enhance the robustness of the model, obtaining an enhanced integrated model. Finally, use the EWMA (Exponentially Weighted Moving Average) method to perform sliding window updates and weight adjustments on the enhanced integrated model. The calculation formula of EWMA is:

[0136] ;

[0137] where is the smoothed value at time t, is the actual observed value at time t, is the smoothing coefficient (0 < α < 1). By dynamically adjusting the value, the model can better adapt to the changing trend of the data, and finally obtain an adaptive integrated model. For example, in network traffic prediction, dynamically adjust the value according to the fluctuation degree of the traffic. Increase the value during rapid traffic changes for quick response, and decrease the value during stable traffic to maintain the stability of the prediction.

[0138] In the model training method provided by the embodiments of the present application, after obtaining the initial data set, time series feature extraction is performed on the initial data set. During the process of time series feature extraction, data decomposition, statistic feature extraction, correlation feature extraction, time-frequency feature extraction, complexity feature extraction, and dimensionality reduction are performed on the data set. Through the above processing methods, the features of large-scale traffic data are finely screened, providing a corresponding basis for the subsequent model training process. After obtaining the first data set, multi-scale clustering processing is performed on the first data set to obtain a clustering data set corresponding to the multi-scale clustering evaluation report. At the same time, anomaly detection is performed on the first data set, and the anomaly detection results under different anomaly detection conditions are analyzed. According to the anomaly detection results, a corresponding anomaly detection report is obtained, and then the anomaly feature set in the first data set is obtained by combining the anomaly detection report and related anomaly indicators. In addition, although the first data set is obtained after time series feature extraction, statistic feature extraction, correlation feature extraction, time-frequency feature extraction, complexity feature extraction, and dimensionality reduction, further feature extraction can be continued on it to obtain target feature sets corresponding to different feature types such as the screened time series feature set, similarity feature set, attention feature set, causal feature set, dependency structure feature set, domain feature set, and dynamic symbol and complexity feature set. Finally, the obtained clustering data set, anomaly feature set, and each target feature set are feature fused to obtain an enhanced feature set; the enhanced feature set is used to train an ensemble model to obtain a trained ensemble model. By applying the method provided by the embodiments of the present application, multiple extractions and optimization processes of different features are performed on large-scale traffic data to train an ensemble model, ensuring that the ensemble model has high-precision processing capabilities. When applying the ensemble model for data cleaning, not only can the data quality and usability be improved, but also the data cleaning efficiency can be increased.

[0139] Based on the model training method provided in the above embodiments, correspondingly, referring to the Figure 4 flow chart shown, the embodiments of the present application also provide a data cleaning method. The data cleaning can be applied to the processors of various system platforms, computer terminals, or various mobile devices. The method includes the following steps:

[0140] S401: Obtain the original traffic data set.

[0141] Optionally, after obtaining the original traffic data set, format unification processing can also be performed on the original traffic data set to obtain a data set with a unified format; specifically, it involves unifying the timestamp format (such as converting to UTC time), unifying the data unit (such as converting all traffic data to bytes / second), and unifying the data type (such as converting all numerical data to floating-point numbers).

[0142] S402: Input the original traffic dataset into the trained ensemble model to obtain the similar dataset output by the ensemble model.

[0143] Among them, the ensemble model is trained by the model training method provided in the above embodiment. For the specific implementation process of this model training method, refer to the steps corresponding to S101 to S106 above, which will not be elaborated here.

[0144] Among them, the similar dataset deletes or fills the missing values in the original traffic dataset, and / or formats or replaces the outliers in the original traffic dataset.

[0145] In this application, the ensemble model is used to identify the missing values and / or outliers in the original traffic dataset. When a missing value is identified, the missing value is deleted or filled. When an outlier is identified, the outlier is replaced or formatted.

[0146] After the ensemble model identifies the missing values and / or outliers in the original traffic dataset, it marks the missing values and / or outliers. Among them, the way for the ensemble model to identify the missing values and outliers in the traffic dataset can be to predict the missing values using the modified Z-score method. The expression of the modified Z-score method is:

[0147] ;

[0148] Among them, X is the original data, is the median, is the median absolute deviation, and k is a constant (usually taken as 1.4826). When |Z| exceeds a certain threshold (such as 3.5), this data is marked as a missing value or an outlier. Among them, when using the modified Z-score method to predict outliers and missing values, different thresholds can be set for distinguishing outliers and missing values. For example, set a first threshold and a second threshold. When |Z| exceeds the first threshold, the data point is marked as a missing value. When |Z| exceeds the second threshold, the data point is marked as an outlier.

[0149] When a missing value in the original traffic data is identified, the missing value can be processed by linear interpolation or spline interpolation. If it is a long-term missing value, a sequence-to-sequence model based on LSTM can be used to fill the missing value.

[0150] Optionally, the way for the ensemble model to identify outliers in the original traffic dataset can also be to identify the outliers in the original traffic dataset based on the anomaly detection report, and this anomaly detection report is obtained by executing the implementation process corresponding to step S103 in the above model training method.

[0151] In the embodiment of the present application, after the integrated model outputs the similar data set, the similar data set can be further processed. The specific processing process is as follows:

[0152] Perform smoothing processing on the similar data set to obtain a smoothed data set; mark the prominent data in the smoothed data set to obtain a marked data set; obtain the data processing log corresponding to the marked data set; update the difference information to the data processing log to obtain an updated data processing log, where the difference information is the difference information obtained by comparing the similar data set with the original traffic data set; based on the updated data processing log and the smoothed data set, obtain a corrected similar data set, and the corrected similar data set is a smoothed data set that at least carries the data processing history record and the difference information in the updated data processing log.

[0153] Apply moving average and wavelet transform to smooth the similar data set. The moving average uses the EWMA method, and the formula corresponding to the EWMA method can refer to the expression sampled when applying EWMA to optimize the integrated model after obtaining the integrated model in the above embodiment, which will not be repeated here. The wavelet transform is a discrete wavelet transform, specifically, it can be to apply the Daubechies wavelet basis function for multi-scale decomposition and reconstruction to remove the high-frequency noise in the similar data set.

[0154] Marking the prominent data in the smoothed data set is specifically to identify the prominent data in the smoothed data set after performing anomaly clustering and threshold division on the smoothed data set. The prominent data can be the abnormal data in the smoothed data set or the data that is quite different from other data. Among them, DBSCAN algorithm is used for density clustering when performing anomaly clustering on the smoothed data set, and then an anomaly threshold is set based on the Mahalanobis distance. Each data in the cluster that exceeds the anomaly threshold in each cluster is the prominent data.

[0155] After marking the prominent data, obtain the data processing log, which contains the processing records of the prominent data after marking the prominent data. After marking the prominent data, perform conditional probability calculation and graph model reasoning on the prominent data to obtain causally adjusted data. Conditional probability calculation uses a Bayesian network to construct the probability dependence relationship between variables. Graph model reasoning uses a causal discovery algorithm such as the PC algorithm to infer the causal direction between variables. Based on the inference result, perform causal adjustment on the data to obtain causally adjusted data. Calculate statistical indicators and consistency measures for the causally adjusted data. The statistical indicators include mean, variance, skewness, kurtosis, etc., and the consistency measure uses Cronbach's α coefficient, and the calculation formula is:

[0156] ;

[0157] where k is the number of items, is the variance of the i-th item, is the total variance.

[0158] According to the above processes of conditional probability calculation and graphical model inference for the prominent data in the labeled dataset, as well as the processes of calculating statistical indicators and consistency metrics, a data processing log for processing the labeled dataset is generated. The data processing log includes information such as the processing history of smoothing the similar dataset, the processing history of processing the labeled dataset, parameter settings, data sources, version numbers, descriptions of processing methods, timestamps, and operators.

[0159] While processing the labeled dataset, the similar dataset is compared with the original traffic dataset to obtain the difference information between the two. The edit distance algorithm (Levenshtein) is applied to calculate the difference between the two. The calculation formula of Levenshtein is:

[0160]

[0161] where a and b are two compared data versions, and i and j are character positions.

[0162] After obtaining the data processing log and the difference information, the difference information is updated to the data processing log. At least the data processing history and the difference information in the updated data processing log are added to the smoothed dataset to obtain a corrected similar dataset. The data processing history includes the processing history of smoothing the similar dataset and the processing history of processing the labeled dataset.

[0163] In the embodiments of the present application, after obtaining the similar dataset by data cleaning the original traffic dataset using the integrated model, the similar dataset can be further smoothed and the prominent data therein can be marked. According to the data processing history and the difference information, subsequent analysis and traceability can be performed on the corrected similar dataset.

[0164] Based on the same inventive concept of the above model training method, an embodiment of the present application also provides a model training apparatus 500. Refer to Figure 5 , Figure 5 is the structural diagram of a model training apparatus 500 provided by an embodiment of the present application. The model training apparatus 500 can be applied to the processors of various system platforms, computer terminals, or various mobile devices. The model training apparatus 500 includes: an acquisition module 501, a clustering module 502, a detection module 503, a multi-feature extraction module 504, a feature fusion module 505, and a training module 506.

[0165] Among them, the acquisition module 501 is used to acquire a first data set, and the first data set is a traffic data set obtained by performing feature extraction according to a time series;

[0166] The clustering module 502 is used to perform multi-scale clustering processing on the first data set to obtain a clustering data set;

[0167] The detection module 503 is used to perform anomaly detection on the first data set to obtain an anomaly feature set;

[0168] The multi-feature extraction module 504 is used to perform multi-feature extraction on the first data set according to a plurality of preset feature types respectively to obtain a target feature set corresponding to each feature type;

[0169] The feature fusion module 505 is used to perform feature fusion on the clustering data set, the anomaly feature set, and each target feature set to obtain an enhanced feature set;

[0170] The training module 506 is used to train an ensemble model by applying the enhanced feature set to obtain a trained ensemble model.

[0171] In an alternative embodiment, the acquisition module 501 acquires the first data set, and specifically is used for:

[0172] Acquire an initial data set; perform time alignment and segmentation processing on the initial data set according to a time series; construct an index for the initial data set that has been aligned and segmented; perform feature extraction on the initial data set with the constructed index according to a time series to obtain a first data set.

[0173] In an alternative embodiment, the acquisition module 501 performs feature extraction on the initial data set with the constructed index according to a time series to obtain a first data set, and specifically is used for:

[0174] Decompose the initial data set with the constructed index to obtain a second data set; extract the statistical feature of the second data set within a preset time window to obtain a third data set; perform relevant feature extraction on the third data set based on the statistical feature of the third data set to obtain a relevant feature set; perform time-frequency feature extraction on the relevant feature set and / or the second data set to obtain a time-frequency feature set, and the time-frequency feature is a three-dimensional feature used to characterize time-frequency-scale; perform complexity feature extraction on the time-frequency feature set and / or the second data set to obtain a complex feature set, and the complexity feature is used to characterize the irregularity degree of the time series; perform feature fusion on the complex feature set and / or the time-frequency feature set and / or the relevant feature set to obtain a first data set.

[0175] In an alternative embodiment, the model training device 500 further includes a dimensionality reduction module; the dimensionality reduction module is configured to perform dimensionality reduction processing on the first data set after the acquisition module 501 fuses the complex feature set and / or the time-frequency feature set and / or the correlation feature set to obtain the first data set.

[0176] The dimensionality reduction module performing dimensionality reduction processing on the first data set specifically includes:

[0177] Screening the main features in the first data set to obtain a screened feature set, where the main features are data features highly correlated with the target variable, and the target variable is the main traffic indicator for prediction; applying an autoencoder to perform non-linear dimensionality reduction processing on the screened feature set to obtain a non-linear dimensionality reduction feature set; performing high-order tensor decomposition on the non-linear dimensionality reduction feature set to obtain a structural feature set; performing visualization processing on the structural feature set to obtain a visualization feature set; clustering the visualization feature set to obtain multiple similar feature groups; resampling each of the similar feature groups to obtain multiple sample sets respectively corresponding to each of the similar feature groups; for each sample set corresponding to the same similar feature group, calculating the similarity coefficient of each feature in the similar feature group in different sample sets; filtering the features in each of the similar feature groups with a similarity coefficient lower than a preset threshold to obtain the dimensionality-reduced first data set.

[0178] In an alternative embodiment, the detection module 503 performing anomaly detection on the first data set to obtain an anomaly feature set specifically includes:

[0179] Applying a preset local outlier factor algorithm to perform anomaly detection on the first data set to obtain a local anomaly detection result; applying a preset seasonal autoregressive integrated moving average model to perform anomaly detection on the first data set to obtain a context anomaly detection result; applying a preset pruning exact linear time algorithm to perform anomaly detection on the first data set to obtain a structural change detection result; applying a preset variational autoencoder to perform anomaly detection on the first data set to obtain a reconstruction anomaly detection result; applying a preset online anomaly detection algorithm to perform anomaly detection on the first data set to obtain a dynamic anomaly detection result; performing evidence synthesis and uncertainty quantification processing on at least one of the local anomaly detection result, the context anomaly detection result, the structural change detection result, the reconstruction anomaly detection result, and the dynamic anomaly detection result to obtain an anomaly detection report; extracting features from the anomaly detection report according to a preset anomaly index to obtain an anomaly feature set.

[0180] In an alternative embodiment, the training module 506 trains the ensemble model using the enhanced feature set to obtain a trained ensemble model, specifically for:

[0181] Split the enhanced feature set into a training set, a test set, and a validation set; use the training set to construct univariate models and multivariate models; apply grid search and time series cross-validation to the univariate models and the multivariate models using the validation set to optimize the model parameters of the univariate models and the multivariate models according to the validation results; perform performance testing on the optimized univariate models and the optimized multivariate models using the test set; in the case of passing the performance test, synthesize the optimized univariate models and the optimized multivariate models to obtain a trained ensemble model.

[0182] The model training device provided in the embodiments of the present application has the same implementation principle and the same technical effects as those in the foregoing embodiments of the model training method. For the sake of brevity, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments.

[0183] Based on the same inventive concept as the above data cleaning method, an embodiment of the present application also provides a data cleaning device 600. Refer to Figure 6 , Figure 6 which is a structural diagram of a data cleaning device 600 provided in an embodiment of the present application. The data cleaning device 600 can be applied to processors of various system platforms, computer terminals, or various mobile devices. The data cleaning device 600 includes: a raw traffic data acquisition module 601 and a data cleaning module 602.

[0184] Among them, the raw traffic data acquisition module 601 is used to acquire a raw traffic data set;

[0185] The data cleaning module 602 is used to input the raw traffic data set into the trained ensemble model to obtain a similar data set output by the ensemble model. The ensemble model is trained using the above model training method;

[0186] Among them, the similar data set deletes or fills in missing values in the raw traffic data set, and / or formats or replaces outliers in the raw traffic data set.

[0187] In an alternative embodiment, the data cleaning device 600 further includes an update module for:

[0188] Smoothing the similar data set to obtain a smoothed data set; marking prominent data in the smoothed data set to obtain a marked data set; obtaining a data processing log corresponding to the marked data set; updating difference information to the data processing log to obtain an updated data processing log, wherein the difference information is difference information obtained by comparing the similar data set with the original traffic data set; based on the updated data processing log and the smoothed data set, obtaining a revised similar data set, wherein the revised similar data set is a smoothed data set that carries at least the data processing history record in the updated data processing log and the difference information.

[0189] The data cleaning device 600 provided in the embodiment of the present application has the same implementation principle and technical effects as those of the aforementioned data cleaning method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned data cleaning method embodiment.

[0190] like Figure 7 As shown, Figure 7 The electronic device 700 provided in the embodiment of the present application is shown in a structural block diagram. The electronic device 700 includes: a transceiver 701 , a memory 702 , a communication bus 703 and a processor 704 .

[0191] The transceiver 701, the memory 702, and the processor 704 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses 703 or signal lines. The transceiver 701 is used to send and receive data. The memory 702 is used to store computer programs, such as storing Figure 5 or Figure 6 The software function module shown in, namely, the model training device 500 or the data cleaning device 600. Among them, the model training device 500 or the data cleaning device 600 includes at least one software function module that can be stored in the memory 702 in the form of software or firmware or solidified in the operating system (OS) of the electronic device 700. The processor 704 is used to execute the executable module stored in the memory 702, such as the software function module or computer program included in the model training device 500 or the data cleaning device 600. For example, the processor 704 is used to execute the model training method or data cleaning method provided in the above embodiment.

[0192] Among them, the memory 702 can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc.

[0193] The processor 704 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), an Accelerated Processing Unit, a Multimedia Application Processor (MAP), a microprocessor, etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. Or the processor 704 can also be any conventional processor, etc.

[0194] Among them, the above-mentioned electronic device 700 includes, but is not limited to, mobile terminals such as mobile phones, laptop computers, tablet computers, etc., and fixed terminals such as digital TVs, desktop computers, etc.

[0195] The embodiments of the present application also provide a non-volatile computer-readable storage medium (hereinafter referred to as the storage medium). A computer program is stored on this storage medium. When the computer program is run by a computer such as the above-mentioned electronic device 700, it executes the model training method or data cleaning method provided in the above-mentioned embodiments.

[0196] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0197] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0198] In addition, in each embodiment of this application, the various functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0199] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which can be a personal computer, a laptop, a server, or an electronic device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned computer-readable storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0200] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that: include: Acquire a first data set, where the first data set is a flow data set obtained by extracting features according to a time series; Performing multi-scale clustering processing on the first data set to obtain a clustered data set; Performing anomaly detection on the first data set to obtain an abnormal feature set; Performing multi-feature extraction on the first data set according to a plurality of preset feature types respectively to obtain a target feature set corresponding to each of the feature types; Performing feature fusion on the clustering data set, the abnormal feature set and each of the target feature sets to obtain an enhanced feature set; Applying the enhanced feature set to train the integrated model to obtain a trained integrated model; The trained integrated model is used to clean the traffic data; The obtaining of the first data set includes: Get the initial data set; Performing time alignment and segmentation processing on the initial data set according to time series; constructing an index for the aligned and segmented initial data set; Perform feature extraction on the indexed initial data set according to time series to obtain a first data set; The step of extracting features from the indexed initial data set according to the time series to obtain the first data set includes: Decomposing the indexed initial data set to obtain a second data set; Extracting statistical features of the second data set within a preset time window to obtain a third data set; Based on the statistical features of the third data set, extract relevant features from the third data set to obtain a relevant feature set; Extracting time-frequency features from the relevant feature set and / or the second data set to obtain a time-frequency feature set, wherein the time-frequency feature is a three-dimensional feature used to characterize time-frequency-scale; Performing complexity feature extraction on the time-frequency feature set and / or the second data set to obtain a complex feature set, wherein the complexity feature is used to characterize the degree of irregularity of the time series; The complex feature set and / or the time-frequency feature set and / or the related feature set are subjected to feature fusion to obtain a first data set; wherein the feature fusion includes at least one of feature concatenation, feature weighting and feature selection.

2. The method according to claim 1, characterized in that After fusing the complex feature set and / or the time-frequency feature set and / or the related feature set to obtain the first data set, the method further includes: Performing dimensionality reduction processing on the first data set, wherein the performing dimensionality reduction processing on the first data set includes: Screening main features in the first data set to obtain a screened feature set, wherein the main features are data features that have a high correlation with a target variable, and the target variable is a main traffic indicator used for prediction; Applying an automatic encoder to perform nonlinear dimensionality reduction processing on the screening feature set to obtain a nonlinear dimensionality reduction feature set; Performing high-order tensor decomposition on the nonlinear dimensionality reduction feature set to obtain a structural feature set; Performing visualization processing on the structural feature set to obtain a visualization feature set; Clustering the visualization feature set to obtain multiple similar feature groups; Resampling each of the similar feature groups to obtain a plurality of sample sets corresponding to each of the similar feature groups; For each sample set corresponding to the same similar feature group, calculate the similarity coefficient of each feature in the similar feature group in different sample sets; The features whose similarity coefficients in each of the similar feature groups are lower than a preset threshold are filtered to obtain the first data set after dimensionality reduction.

3. The method according to claim 1, characterized in that The performing anomaly detection on the first data set to obtain an abnormal feature set includes: Applying a preset local anomaly factor algorithm to perform anomaly detection on the first data set to obtain a local anomaly detection result; Applying a preset seasonal autoregressive integrated moving average model to perform anomaly detection on the first data set to obtain a contextual anomaly detection result; Applying a preset pruning exact linear time algorithm to perform anomaly detection on the first data set to obtain a structural change detection result; Applying a preset variational autoencoder to perform anomaly detection on the first data set to obtain a reconstructed anomaly detection result; Applying a preset online anomaly detection algorithm to perform anomaly detection on the first data set to obtain a dynamic anomaly detection result; Performing evidence synthesis and uncertainty quantification processing on at least two of the local anomaly detection result, the context anomaly detection result, the structural change detection result, the reconstruction anomaly detection result, and the dynamic anomaly detection result to obtain an anomaly detection report; Feature extraction is performed on the anomaly detection report according to preset anomaly indicators to obtain an anomaly feature set.

4. The method according to claim 1, characterized in that: The applying the enhanced feature set to train the integrated model to obtain a trained integrated model includes: Splitting the enhanced feature set into a training set, a test set, and a validation set; Using the training set to construct univariate models and multivariate models; Applying the validation set to perform grid search and time series cross validation on the univariate model and the multivariate model, so as to optimize model parameters of the univariate model and the multivariate model according to the validation results; Using the test set to perform performance testing on the optimized univariate model and the optimized multivariate model; When the performance test is passed, the optimized univariate model and the optimized multivariate model are synthesized to obtain a trained integrated model.

5. A data cleaning method, characterized in that: include: Get the original traffic data set; Inputting the original traffic data set into a trained integrated model to obtain a similar data set output by the integrated model, wherein the integrated model is trained using the model training method according to any one of claims 1 to 4; The similar data set is to delete or fill missing values ​​in the original traffic data set, and / or to format or replace abnormal values ​​in the original traffic data set.

6. The method according to claim 5, characterized in that Also includes: Performing smoothing on the similar data set to obtain a smoothed data set; Marking the prominent data in the smoothed data set to obtain a marked data set; Get the data processing log corresponding to the labeled dataset; Updating the difference information to the data processing log to obtain the updated data processing log, wherein the difference information is the difference information obtained by comparing the similar data set with the original traffic data set; Based on the updated data processing log and the smoothed data set, the revised similar data set is obtained, and the revised similar data set is a smoothed data set that carries at least the data processing history records in the updated data processing log and the difference information.

7. A model training device, characterized in that: include: An acquisition module, used to acquire a first data set, where the first data set is a flow data set obtained by extracting features according to a time series; A clustering module, used for performing multi-scale clustering processing on the first data set to obtain a clustered data set; A detection module, used to perform anomaly detection on the first data set to obtain an abnormal feature set; A multi-feature extraction module, used to perform multi-feature extraction on the first data set according to a plurality of preset feature types, to obtain a target feature set corresponding to each of the feature types; A feature fusion module, used for fusing the clustering data set, the abnormal feature set and each target feature set to obtain an enhanced feature set; A training module, used to apply the enhanced feature set to train the integrated model to obtain a trained integrated model; the trained integrated model is used to clean the traffic data; The acquisition module is specifically used to acquire an initial data set; perform time alignment and segmentation processing on the initial data set according to a time series; construct an index for the aligned and segmented initial data set; perform feature extraction on the indexed initial data set according to a time series to obtain a first data set; The acquisition module is specifically used to decompose the indexed initial data set to obtain a second data set; extract the statistical features of the second data set within a preset time window to obtain a third data set; based on the statistical features of the third data set, extract relevant features of the third data set to obtain a relevant feature set; extract time-frequency features of the relevant feature set and / or the second data set to obtain a time-frequency feature set, wherein the time-frequency features are three-dimensional features used to characterize time-frequency-scale; Extracting complexity features from the time-frequency feature set and / or the second data set to obtain a complex feature set, wherein the complexity features are used to characterize the degree of irregularity of the time series; fusing the complex feature set and / or the time-frequency feature set and / or the related feature set to obtain a first data set; wherein the feature fusion includes feature splicing , feature weighting and feature selection.

8. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is connected to the memory; The memory is used to store programs; The processor is used to call the program stored in the memory to execute the model training method as described in any one of claims 1 to 4, or to execute the data cleaning method as described in claim 5 or 6.

Citation Information

Patent Citations

  • AI training processing method based on big data cleaning and artificial intelligence training system

    CN115422179A

  • Dark web traffic classification method and system based on multi-modal fusion

    CN117911782A