Method for mining key features of time series data

CN122817633APending Publication Date: 2026-09-25HUBEI POLYTECHNIC UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610882064.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]申请号为202110054385.X的发明专利申请中公开了一种基于动态网格划分的时序数据趋势特征提取方法,该申请旨在解决“现有技术在等间隔抽取数据时,只能保证主体的趋势而容易遗漏时序数据的关键特征

Benefits of technology

本发明通过对高维时序数据做滑动窗口划分并动态识别、拟合替换噪声点,完整留存数据原始时序结构,可自适应调整窗口长度校验时序结构相似度,避免降噪处理破坏数据内在时序关联,基于数据时序跨度与维度分布生成动态多尺度特征基,尺度参数可随数据特性自主适配变化,并采用轻量化编码方式并行处理多尺度特征,能实时平衡计算资源消耗与特征表达能力,同时,依靠时序关联网络挖掘不同时刻、不同维度特征的深层依赖关系,经分层遍历量化特征关联显著性,平稳完成关键特征筛选,最后应用时序任务性能指标形成闭环反馈,动态微调特征挖掘相关参数实现迭代优化,筛选所得关键特征可融入任务模型注意力机制,合理引导模型关注高关联特征组合,从而稳步提升时序预测与分类任务的实际表现。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817633A_ABST
    Figure CN122817633A_ABST
Patent Text Reader

Abstract

The application discloses a key feature mining method for time series data and relates to the field of data mining, and comprises the following steps: performing sliding window division on input high-dimensional time series original data, identifying noise points in each window by dynamically calculating a data fluctuation threshold, replacing the noise points with time series trend fitting values of non-noise data in the window, retaining the original time series structure of the data, and outputting denoising reconstructed time series data; the application can accurately remove time series data noise and completely retain the original time series structure, adaptively adapt to multi-time and multi-dimensional analysis scales, balance the operation efficiency and feature expression capability, deeply mine the internal correlation of data across time and across dimensions, and efficiently screen high-value core features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data mining technology, specifically to a method for mining key features of time-series data. Background Technology

[0002] Time series data is widely found in fields such as industrial monitoring, financial transactions, and IoT sensing. It consists of observations arranged in chronological order and exhibits characteristics such as time dependence, non-stationarity, and multi-scale. Key feature mining is the core of time series data processing, aiming to extract representative information from massive amounts of redundant data to support subsequent tasks such as prediction and classification. Its methods have evolved from traditional statistical models to end-to-end deep learning models.

[0003] Patent application No. 202110054385.X discloses a method for extracting trend features of time-series data based on dynamic grid partitioning. This application aims to solve the problem that "existing technologies, when extracting data at equal intervals, can only guarantee the main trend and are prone to missing key features of time-series data. Patent application No. CN108804731A discloses a trend feature extraction method based on dual evaluation factors of important points. It is based on piecewise linearity and supplemented by distance factors and trend factors. It can extract the main trend and key features of time-series data, but it is easy to ignore secondary key features. Moreover, the threshold and weight in this method need to be identified according to the specific dataset, which has certain limitations."

[0004] However, existing time series feature mining methods are either poorly adaptable due to manual design or suffer from an imbalance between computational efficiency and feature discrimination power in automatic mining models, and are difficult to fully mine the deep dynamic correlations of high-dimensional noisy time series data.

[0005] To address this, we propose a key feature mining method for time-series data. Summary of the Invention

[0006] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a key feature mining method for time series data, which can effectively solve the problems of the existing technology.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions; This invention discloses a key feature mining method for time-series data, including: Step 1: Perform sliding window partitioning on the input high-dimensional time-series raw data. Within each window, identify noise points by dynamically calculating the data fluctuation threshold. Replace the noise points with the time-series trend fitting value of the non-noise data within the window, preserving the original time-series structure of the data, and output denoised and reconstructed time-series data. Step 2: Generate a multi-scale feature base set based on the time-series span and dimensional distribution of the denoised and reconstructed time-series data. Each feature base corresponds to a different combination of time granularity and dimension, and the scale parameter of the feature base is dynamically adjusted according to the data characteristics. Step 3: Perform parallel encoding of the feature bases at each scale using a lightweight time-series encoding unit. During the encoding process, a dynamic parameter adjustment mechanism is used to balance the consumption of computational resources and the feature expression capability in real time, dynamically adapting the computational load and feature expression capability. Step 4: Based on the predefined criteria for discriminative power, the encoded feature vectors are obtained; Step 5: Based on the encoded feature vectors, a temporal correlation network is constructed to mine the deep dependencies between features at different times and in different dimensions. The network nodes are the feature vectors at each time point, and the edge weights are determined by calculating the combined value of the temporal evolution gradient and dimensional co-coefficient between nodes; Step 6: A hierarchical traversal is performed on the temporal correlation network to calculate the correlation significance score of the features of each node, and a predefined number of key features are selected from high to low scores; Step 7: The selected key features are applied to temporal prediction or classification tasks, and the feature mining parameters are dynamically adjusted based on the task performance indicators fed back in real time. If the performance does not meet the predefined criteria, the process returns to Step 2 to re-optimize the feature scale.

[0008] Furthermore, in step 1, when outputting the denoised and reconstructed time-series data, the high-dimensional time-series original data is divided into continuous sliding windows of equal length, and a dynamic fluctuation threshold is calculated for the data within each sliding window: ; In the formula: This is the median of the absolute values ​​of the first-order differences of all data points within the current sliding window. , , This is a preset dimensionless adjustment coefficient; The dimensionless local trend curvature of the data within the current sliding window; This represents the local trend abrupt change entropy of the data within the current sliding window. When the absolute value of the difference between a data point within the window and its two adjacent data points is greater than the dynamic fluctuation threshold When this happens, the data point is determined to be a noise point; Noise points are replaced by the time-series trend fitting value of non-noise data within the window. The time-series trend fitting value is dynamically selected based on the number of non-noise data points within the window: when the number of non-noise data points within the window is greater than or equal to 3, a quadratic time-series trend polynomial is fitted using all non-noise data points within the window with the local time step index within the window as the independent variable. The coefficients of each term of the quadratic polynomial are obtained by solving the problem using the unweighted least squares method. The local time step index within the window corresponding to the noise point is substituted into the quadratic polynomial, and the quadratic extrapolation replacement value is calculated as the time-series trend fitting value; when the number of non-noise data points within the window is 2, the linear interpolation result of two non-noise data points is used as the time-series trend fitting value; when the number of non-noise data points within the window is 1, the eigenvalue of the non-noise data point is used as the time-series trend fitting value.

[0009] Furthermore, step 1 includes a timing structure verification step before outputting the denoised and reconstructed timing data: The temporal structure similarity between the denoised reconstructed time series data and the original high-dimensional time series data is calculated. The temporal structure similarity is obtained by calculating the reciprocal of the dynamic time warp distance between the two data sequences. When the time series structure similarity is greater than a preset similarity threshold, output denoised and reconstructed time series data; When the time sequence structure similarity is less than or equal to the preset similarity threshold, the length of the sliding window is adaptively adjusted according to the intensity of local data fluctuations. The more intense the local fluctuations, the shorter the sliding window length is; the more gradual the local fluctuations, the longer the sliding window length is. The window division, noise identification, and noise point replacement steps are then re-executed.

[0010] Furthermore, in step 2, all possible combinations of time granularity and dimension are traversed, and the time granularity of each group is calculated. Combination with dimensions Matching degree ; In the formula: To denoise and reconstruct the time series autocorrelation coefficient of the time series data at this time granularity; To denoise and reconstruct the dimensional mutual information of time series data for this dimensional combination; This is a preset dimensionless minimum value; These are the time-series-dimensional coupling coefficients. ; in, This represents the gradient vector of the time series autocorrelation coefficient as the time granularity changes. This is the gradient vector of dimensional mutual information when the dimensional combination changes; Select matching degree All time granularities and dimension combinations greater than the preset matching degree threshold generate a corresponding multi-scale feature base set, and the scale parameters of the feature base are updated in real time with the changes in the temporal autocorrelation coefficient and dimensional mutual information of the input data.

[0011] Furthermore, in step 3, the feature bases at each scale are encoded in parallel using lightweight temporal coding units, and the dynamic parameter adjustment mechanism is used to balance computational resource consumption and feature representation capability in real time during the encoding process. Specifically, the steps are as follows: The lightweight temporal coding unit adopts a depthwise separable temporal convolutional structure, and the dynamic parameter adjustment mechanism is as follows: Real-time monitoring of the remaining computing power percentage of the current computing resource unit and the information entropy of the encoded feature vector; When the remaining computing power ratio is greater than the preset computing power threshold and the feature information entropy is less than the preset entropy threshold, the dilation coefficient of the temporal convolution is increased. When the remaining computing power percentage is less than the preset computing power threshold and the feature information entropy is greater than the preset entropy threshold, the dilation coefficient of the temporal convolution is reduced. The dimension of the feature vector of the encoded output remains unchanged throughout the parameter adjustment process; The dilation coefficient of the temporal convolution is adjusted in a stepwise manner. The step size for each adjustment is a preset step size, and the value range of the dilation coefficient is limited to between the preset minimum dilation coefficient and the preset maximum dilation coefficient. When increasing the dilation coefficient, the preset step size is increased each time until the dilation coefficient reaches the preset maximum dilation coefficient or the information entropy of the encoded feature vector increases to the preset entropy threshold. When decreasing the dilation coefficient of the temporal convolution, the preset step size is decreased each time until the dilation coefficient reaches the preset minimum dilation coefficient or the remaining computing power ratio of the current computing resource unit increases to the preset computing power threshold. Each scale feature base is input into an independent lightweight temporal coding unit for parallel processing, and all coding units share the same dynamic parameter adjustment mechanism.

[0012] Furthermore, in step 4, the encoded feature vector at each time step is used as a node in the temporal correlation network, and the results of any two nodes are calculated. and Edge weights between The edge weight ; In the formula: For nodes To the node The temporal evolution gradient; For nodes and The dimensional nonlinear dependence coefficient; It serves as the global temporal evolution consistency factor. When edge weight When the weight is greater than the preset edge weight threshold, at the node With nodes A directed edge is generated between the nodes, with the edge pointing from the earlier node to the later node.

[0013] Furthermore, a hierarchical traversal is performed on the temporal correlation network to calculate the association saliency score of each node's features: The temporal correlation network is divided into several levels according to the time sequence, and each level corresponds to a preset time interval; Based on a bottom-up hierarchical traversal strategy, we first traverse all nodes corresponding to the lowest time interval and calculate the local association significance score of each node. By traversing upwards layer by layer, the local association saliency scores of lower-level nodes are passed to the upper-level associated nodes, ultimately obtaining the global association saliency score for each node.

[0014] Furthermore, the initial local association saliency score of the node is taken as the geometric mean of the product of the weights of all incoming edges and the product of the weights of all outgoing edges of the node. The global association significance score of a node is equal to the final local association significance score of that node multiplied by the proportion of the time span of the node's level to the total time series span.

[0015] Furthermore, the feature mining parameters are dynamically adjusted based on real-time feedback of task performance metrics. The task performance metrics include the mean absolute error of a time-series prediction task or the classification accuracy of a time-series classification task. When the task performance indicators do not meet the preset standards, adjust the matching threshold of the feature base. If the task performance still fails to meet the preset standard after adjusting the matching degree threshold, then adjust the edge weight threshold of the temporal correlation network. If the task performance still fails to meet the preset standard after adjusting the edge weight threshold, the multi-scale feature base set is regenerated, and the subsequent encoding, temporal correlation network construction, and feature selection steps are repeated.

[0016] Furthermore, when applying the selected key features to time-series prediction or classification tasks, the selected key features are concatenated into a feature sequence according to the time sequence. At the same time, the edge weights between the nodes corresponding to the key features in the time-series association network are used as prior weights and embedded into the attention mechanism of the model used in the time-series prediction or classification task. This allows the model to allocate different attention resources according to the prior weights during the inference process, giving priority to feature combinations with higher correlation significance.

[0017] Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects: This invention divides high-dimensional time-series data into sliding windows and dynamically identifies, fits, and replaces noise points, preserving the original time-series structure of the data. It adaptively adjusts the window length to verify the similarity of the time-series structure, avoiding noise reduction that could damage the inherent time-series correlations. Based on the data's time-series span and dimensional distribution, it generates a dynamic multi-scale feature base, with scale parameters adapting autonomously to data characteristics. A lightweight encoding method is used to process multi-scale features in parallel, balancing computational resource consumption and feature representation capabilities in real time. Simultaneously, it relies on a time-series correlation network to mine deep dependencies between features at different times and dimensions. Through hierarchical traversal, it quantifies the significance of feature correlations, smoothly completing the selection of key features. Finally, it applies time-series task performance metrics to form a closed-loop feedback loop, dynamically fine-tuning feature mining parameters for iterative optimization. The selected key features can be integrated into the task model's attention mechanism, reasonably guiding the model to focus on highly correlated feature combinations, thereby steadily improving the actual performance of time-series prediction and classification tasks. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0019] Figure 1 This is a flowchart illustrating a key feature mining method for time-series data. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0021] The present invention will be further described below with reference to embodiments.

[0022] Example: The key feature mining method for time-series data in this embodiment, such as Figure 1 As shown, it includes: Step 1: Perform sliding window partitioning on the input high-dimensional time series raw data. Within each window, identify noise points by dynamically calculating the data fluctuation threshold. Replace the noise points with the time series trend fitting value of the non-noise data within the window, preserve the original time series structure of the data, and output the denoised and reconstructed time series data. Step 1: When outputting denoised and reconstructed time-series data, the high-dimensional original time-series data is divided into continuous sliding windows of equal length, and a dynamic fluctuation threshold is calculated for the data within each sliding window. ; In the formula: This is the median of the absolute values ​​of the first-order differences of all data points within the current sliding window. , , This is a preset dimensionless adjustment coefficient; The dimensionless local trend curvature of the data within the current sliding window is calculated as follows: normalize the feature values ​​of all data points within the window to the interval [0,1], construct a planar point set with the time step index as the horizontal axis and the normalized feature value as the vertical axis, take the first and last points and the middle point to form a triangle, and calculate the reciprocal of the radius of the circumcircle of the triangle. The local trend abrupt change entropy of the data within the current sliding window is calculated by: counting the frequency of the first-order difference symbol of all continuous non-noise data points within the window, and then calculating it according to the Shannon entropy formula; The eigenvalue normalization uses the minimum-maximum normalization algorithm, and the normalization formula is as follows: ,in , These are the maximum and minimum eigenvalues ​​of all data points within the current sliding window, respectively. When the number of data points in the sliding window is odd, the index of the data point in the window sequence is selected as the midpoint; when the number of data points is even, the last data point in the first half of the window is selected as the midpoint. The time step index is identified by consecutive positive integers and used as the horizontal coordinate of the plane coordinate system. The normalized eigenvalues ​​are used as the dimensionless values ​​of the vertical coordinate to construct a standard Cartesian plane point set. First-order differencing uniformly adopts the backward differencing method, that is, taking the eigenvalue of the later time step and subtracting the eigenvalue of the previous time step from the eigenvalue of two adjacent consecutive data points; the first-order differencing result is marked as positive if greater than 0, negative if less than 0, and stationary if equal to 0; the Shannon entropy calculation formula is... In the formula The frequency of each symbol appearing in all difference results within the window. The total number of symbol categories; The above formula uses the median of the absolute value of the first difference of the data within the sliding window as the threshold benchmark. By leveraging the convergence property of the hyperbolic tangent function, it integrates two core time series features: local trend curvature and local trend abrupt change entropy. It also uses three dimensionless coefficients to adjust the influence weights of each feature and the benchmark offset. This allows it to adaptively generate a dynamic fluctuation threshold based on the fluctuation characteristics of the data within the window, thus overcoming the limitation that fixed thresholds cannot adapt to the local fluctuation differences of time series data and accurately adapting to the noise judgment criteria of different time series segments. When the absolute value of the difference between a data point within the window and its two adjacent data points is greater than the dynamic fluctuation threshold When this happens, the data point is determined to be a noise point; Noise points are replaced by the fitted time-series trend values ​​of non-noise data points within a window. The fitted time-series trend values ​​are dynamically selected based on the number of non-noise data points within the window: When the number of non-noise data points is greater than or equal to 3, a quadratic time-series trend polynomial is fitted using all non-noise data points within the window and their local time step indices as independent variables. The coefficients of the quadratic polynomial are obtained using unweighted least squares. The local time step index corresponding to the noise point is substituted into this quadratic polynomial to calculate the quadratic extrapolation replacement value, which is then used as the fitted time-series trend value. When the number of non-noise data points within the window is 2, the linear interpolation result of two non-noise data points is used as the fitted time-series trend value. When the number of non-noise data points within the window is 1, the eigenvalue of that non-noise data point is used as the fitted time-series trend value. The form of the quadratic time-series trend polynomial is: Where t is the time step index, Here, represents the eigenvalues ​​for the corresponding time steps, where a is the coefficient of the quadratic term, b is the coefficient of the linear term, and c is the constant term. Among them, the local trend curvature adjustment coefficient α is determined by statistically analyzing the noise recognition accuracy of multiple sets of standard time series datasets under different curvature weights, and taking the value that makes the average recognition accuracy reach the peak; the local trend mutation entropy adjustment coefficient β is determined by statistically analyzing the noise recall of multiple sets of time series datasets containing trend mutation noise, and taking the value that makes the average recall reach the peak; the dynamic fluctuation threshold benchmark offset coefficient γ is determined by statistically analyzing the false positive rate of multiple sets of clean and noise-free time series datasets, and taking the minimum value that makes the average false positive rate lower than the preset false positive rate threshold. Step 1 includes a timing structure verification step before outputting the denoised and reconstructed timing data: The temporal structure similarity between the denoised and reconstructed time series data and the original high-dimensional time series data is calculated. The temporal structure similarity is obtained by calculating the reciprocal of the dynamic time warp distance between the two data sequences. When the time series structure similarity is greater than a preset similarity threshold, output denoised and reconstructed time series data; When the time sequence structure similarity is less than or equal to the preset similarity threshold, the length of the sliding window is adaptively adjusted according to the intensity of local data fluctuations. The more intense the local fluctuations, the shorter the sliding window length is; the more gentle the local fluctuations, the longer the sliding window length is. The window division, noise identification, and noise point replacement steps are then re-executed. The severity of local fluctuations in the data is characterized by the degree of dispersion of the first-order difference of the time series data within the corresponding data interval; Step 2: Generate a multi-scale feature base set based on the temporal span and dimensional distribution of the denoised and reconstructed time series data. Each feature base corresponds to a different combination of time granularity and dimension, and the scale parameter of the feature base is dynamically adjusted according to the data characteristics. In step 2, during the execution phase, all possible combinations of time granularity and dimension are traversed, and the time granularity of each combination is calculated. Combination with dimensions Matching degree ; The above formula performs basic coupling operations on the time series autocorrelation coefficient and dimensional mutual information, introduces a minimum value to avoid the calculation failure of zero denominator, and then combines the time series-dimensional coupling coefficient for weighted correction. At the same time, it comprehensively evaluates the adaptation effect of time granularity and dimensional combination from three dimensions: time series self-correlation, dimensional information correlation, and spatiotemporal dimensional linkage matching. This breaks through the one-sidedness of single index evaluation and achieves accurate screening of multi-scale feature basis combination. In the formula: To denoise and reconstruct the time series autocorrelation coefficient of the time series data at this time granularity; To denoise and reconstruct the dimensional mutual information of time series data for this dimensional combination; This is a preset dimensionless minimum value; These are the time-series-dimensional coupling coefficients. ; The above formula uses the gradient vector dot product to represent the synergy between the changes in temporal autocorrelation and dimensional mutual information. It completes the normalization process by superimposing the minimum value of the gradient magnitude, thus eliminating the problem of singular values ​​in numerical calculation. It uses the similarity logic of vector space to characterize the dynamic evolution and linkage of temporal granularity and dimensional combination, effectively capturing the deep coupling relationship between temporal characteristics and dimensional features, and making up for the shortcomings of traditional methods in being unable to characterize the synergistic changes of the gradients of the two. in, This represents the gradient vector of the time series autocorrelation coefficient as the time granularity changes. This is the gradient vector of dimensional mutual information when the dimensional combination changes; Select matching degree All time granularities and dimension combinations greater than the preset matching threshold are used to generate a corresponding multi-scale feature base set, and the scale parameters of the feature base are updated in real time with the changes in the temporal autocorrelation coefficient and dimensional mutual information of the input data. in, Determined in the following ways: Denoising and reconstructing time-series data according to the current time granularity Perform equal-interval resampling to obtain a length of Resampled time series Each resampling point The arithmetic mean of the feature values ​​of all data points within the corresponding time granularity window is given; the Pearson autocorrelation coefficient of the resampled time series at a lag of 1 step is calculated, i.e. , For resampled time series The arithmetic mean, The length of the resampled time sequence; This formula resamples time series data according to a set time granularity and reconstructs the sequence with the interval mean. It uses the covariance and variance ratio with a lag of one step to solve the Pearson autocorrelation coefficient. It removes the interference caused by the overall numerical shift by using the sequence mean, and accurately measures the strength of the linear correlation between time series before and after time series at different time granularities. It provides a reliable quantitative basis for the dynamic update of multi-scale feature base scale parameters. Step 3: The feature bases at each scale are encoded in parallel using a lightweight temporal coding unit. During the encoding process, a dynamic parameter adjustment mechanism is used to balance the consumption of computing resources and the feature representation capability in real time, dynamically adapting to the preset standards of computational load and feature discriminative power, and obtaining the encoded feature vector. Step 3 involves parallel encoding of feature bases at various scales using lightweight temporal coding units. The specific steps involved in balancing computational resource consumption and feature representation capability in real-time through a dynamic parameter adjustment mechanism are as follows: The lightweight temporal coding unit adopts a depthwise separable temporal convolutional structure, and the dynamic parameter adjustment mechanism is as follows: Real-time monitoring of the remaining computing power percentage of the current computing resource unit and the information entropy of the encoded feature vector; When the remaining computing power ratio is greater than the preset computing power threshold and the feature information entropy is less than the preset entropy threshold, the dilation coefficient of the temporal convolution is increased. When the remaining computing power percentage is less than the preset computing power threshold and the feature information entropy is greater than the preset entropy threshold, the dilation coefficient of the temporal convolution is reduced. The dimension of the feature vector of the encoded output remains unchanged throughout the parameter adjustment process; The depthwise separable temporal convolutional structure consists of channel-wise temporal convolutional layers and pointwise convolutional layers stacked sequentially. The channel-wise temporal convolutional layer performs a one-dimensional temporal convolution operation independently on each input channel, with a kernel size of k×1 (k is the preset temporal kernel length). It slides only along the time dimension to extract independent temporal local features of each channel. The pointwise convolutional layer uses a 1×1 convolutional kernel to linearly fuse the output features of all channels to integrate cross-channel feature information. By decomposing the standard temporal convolution into two independent convolutional steps, the number of computational parameters and computational complexity are significantly reduced while ensuring feature extraction capabilities. Wherein, the temporal convolution kernel length Only odd-numbered values ​​are selected, with the range limited to 3, 5, and 7; the length of the convolution kernel does not exceed the time step length of a single division of the sliding window, and maintains an integer multiple matching relationship with the minimum time granularity of the multi-scale feature base. Small-sized convolution kernels are preferred for stationary time series data, while large-sized convolution kernels are preferred for time series data with abrupt changes. The dilation coefficient of temporal convolution is adjusted in a stepwise manner. The step size for each adjustment is a preset step size, and the value range of the dilation coefficient is limited to the preset minimum dilation coefficient and the preset maximum dilation coefficient. When increasing the dilation coefficient, the preset step size is increased each time until the dilation coefficient reaches the preset maximum dilation coefficient or the information entropy of the encoded feature vector increases to the preset entropy threshold. When decreasing the dilation coefficient of temporal convolution, the preset step size is decreased each time until the dilation coefficient reaches the preset minimum dilation coefficient or the remaining computing power ratio of the current computing resource unit increases to the preset computing power threshold. Each scale feature base is input into an independent lightweight temporal coding unit for parallel processing. All coding units share the same dynamic parameter adjustment mechanism to ensure the consistency of coding quality for features at different scales. The temporal convolution dilation coefficient is limited to a range of 1 to 16, with a preferred minimum dilation coefficient of 1 and a preferred maximum dilation coefficient of 16. The preset adjustment step size for the dilation coefficient is fixed at 1. The preset threshold for remaining computing power is 30% to 70%, with a preferred critical threshold of 40%. The preset entropy threshold for feature vector information entropy is 1.2 to 1.8, with a preferred baseline threshold of 1.5. The upper limit of the range is selected for high-frequency sampled temporal data, and the lower limit of the range is selected for low-frequency stable temporal data, thereby achieving adaptive parameter matching. The information entropy of the encoded feature vector is calculated using the histogram estimation method. Each dimension of the feature vector is divided into a preset number of equal-width intervals, the frequency of occurrence of feature values ​​in each interval is counted, and then calculated according to the Shannon entropy formula. The number of equal-width intervals for a single dimension of the feature vector ranges from 5 to 15, with a preferred fixed division of 10 equal-width intervals. The start and end boundaries of the intervals are defined by the maximum and minimum values ​​under the corresponding feature vector dimension as the global boundary, and the range of each interval is evenly divided. The interval boundary values ​​are assigned to the adjacent interval on the right to complete the frequency statistics. Step 4: Construct a temporal correlation network based on the encoded feature vectors to explore the deep dependencies between features at different times and in different dimensions. The network nodes are the feature vectors at each time point, and the edge weights are determined by calculating the combined value of the temporal evolution gradient and the dimensional co-coefficient between nodes. Step 4, during execution, uses the encoded feature vector at each time step as a node in the temporal correlation network, and calculates the values ​​of any two nodes. and Edge weights between edge weight ; In the formula: For nodes To the node The temporal evolution gradient is calculated using the following formula: , and They are nodes and The L2 norm of the eigenvectors, with dimensions of eigenvalues / time; For nodes and The nonlinear dependence coefficient of the dimension is calculated using the following formula: ,in The mutual information of two feature vectors. and These are the information entropies of the two feature vectors, both of which are dimensionless. The global temporal evolution consistency factor is calculated using the following formula: ,in, It is the average vector of the temporal evolution gradients of all adjacent node pairs, with the dimension being eigenvalue dimension / time; The above formula combines the absolute value of the temporal evolution gradient, the dimensional nonlinear dependency coefficient, and the global temporal evolution consistency factor to comprehensively define the network node edge weights from three aspects: feature temporal evolution rate, cross-dimensional nonlinear correlation, and the degree of fit between local and global evolution trends. It abandons the simple method of weighting by single distance or correlation coefficient, and allows the weights to truly reflect the deep dependency relationship between features at different times and in different dimensions. When edge weight When the weight is greater than the preset edge weight threshold, at the node With nodes A directed edge is generated between the nodes, with the direction of the edge pointing from the earlier node to the later node. Step 5: Perform a hierarchical traversal on the temporal correlation network, calculate the correlation significance score of each node's features, and select a preset number of key features from high to low scores. The temporal correlation network is subjected to a hierarchical traversal, and the correlation significance score of each node's features is calculated in the following stages: The temporal correlation network is divided into several levels according to the time sequence, and each level corresponds to a preset time interval; The temporal correlation network adopts a temporal span equal division rule. The number of layers is preset according to the total duration of the overall temporal series, and the total duration is divided into layer units with equal time spans. The time span of a single element in a hierarchy shall be no less than 10 times the data sampling period and no more than one-tenth of the total duration of the overall time series. If the time series data has business cycle characteristics, the hierarchy can be divided according to the natural business cycle instead of the equal division rule. Based on a bottom-up hierarchical traversal strategy, we first traverse all nodes corresponding to the lowest time interval and calculate the local association significance score of each node. By traversing upwards layer by layer, the local association saliency scores of lower-level nodes are passed to the upper-level associated nodes, and finally the global association saliency score of each node is obtained. The specific process for transferring the initial local association saliency score of lower-level nodes to upper-level associated nodes is as follows: For each lower-level node, identify all upper-level associated nodes that have directed edges connected to it and are located in a later time interval; calculate the normalized value of the edge weights between the node and each upper-level associated node, where the normalized value is the ratio of the corresponding edge weight to the sum of the edge weights of all edges pointing to upper-level associated nodes from the node; allocate the initial local association saliency score of the node to the corresponding upper-level associated nodes according to the normalized value; each upper-level associated node arithmetically sums its own initial local association saliency score with the scores allocated from all lower-level nodes to obtain the final local association saliency score of the upper-level associated node. The initial local association significance score of a node is the geometric mean of the product of the weights of all incoming edges and the product of the weights of all outgoing edges of that node. The global association significance score of a node is equal to the final local association significance score of that node multiplied by the proportion of the time span of the node's level to the total time series span. Step 6: Apply the selected key features to time series prediction or classification tasks, dynamically adjust the feature mining parameters based on real-time feedback of task performance indicators, and return to Step 2 to re-optimize the feature scale if the performance does not meet the preset standard. The feature mining parameters are dynamically adjusted based on real-time feedback of task performance metrics. Task performance metrics include the mean absolute error for time series prediction tasks or the classification accuracy for time series classification tasks; When the task performance indicators do not meet the preset standards, adjust the matching degree threshold of the feature base; If the task performance still fails to meet the preset standard after adjusting the matching degree threshold, then adjust the edge weight threshold of the temporal correlation network. If the task performance still fails to meet the preset standard after adjusting the edge weight threshold, the multi-scale feature base set is regenerated, and the subsequent encoding, temporal correlation network construction and feature selection steps are repeated. The adjustment logic for the matching degree threshold of the feature base is as follows: if the task performance index is lower than the preset standard, the matching degree threshold is reduced to increase the number of selected time granularity and dimension combinations, and expand the coverage of the multi-scale feature base set; if the task performance index is higher than the preset standard and the computational resource consumption exceeds the preset upper limit, the matching degree threshold is increased to reduce the number of selected time granularity and dimension combinations, and reduce the overall computational complexity. The threshold values ​​for temporal structure similarity are 0.8 to 0.95, the threshold values ​​for multi-scale feature base matching are 0.6 to 0.8, the threshold values ​​for temporal association network edge weights are 0.35 to 0.5, and the preset threshold for the false positive rate of clean temporal data is fixed at less than 5%. For high-dimensional, highly volatile temporal data, the lower limit of each threshold range is selected, and for low-dimensional, stable temporal data, the upper limit of the range is selected. The threshold settings can be iteratively fine-tuned based on the feature mining accuracy of similar historical temporal datasets. The adjustment logic for the edge weight threshold of the temporal correlation network is as follows: if the task performance index is still lower than the preset standard after adjusting the matching degree threshold of the feature base, the edge weight threshold is reduced to retain more weak correlation edges and explore more comprehensive dependencies between features at different times and in different dimensions; if the task performance index is higher than the preset standard but there is overfitting, the edge weight threshold is increased to filter out false weak correlation edges and enhance the discriminative power of the selected key features. When the selected key features are applied to time series prediction or classification tasks, the selected key features are concatenated into a feature sequence in time series order. At the same time, the edge weights between the nodes corresponding to the key features in the time series association network are used as prior weights and embedded into the attention mechanism of the model used in the time series prediction or classification task. This allows the model to allocate different attention resources according to the prior weights during the inference process, giving priority to feature combinations with higher correlation significance.

[0023] The methods described in the above embodiments can accurately remove noise from time-series data while fully preserving the original time-series structure. They can adaptively adapt to multiple time and dimensional analysis scales, balance computational efficiency and feature representation capabilities, deeply mine the intrinsic correlations between data across time and dimensions, efficiently screen high-value core features, and dynamically optimize parameters based on task performance to avoid overfitting. This effectively improves the accuracy of time-series prediction and classification, reduces computational overhead, adapts to various time-series application scenarios, and effectively enhances the accuracy of data analysis and its practical application.

[0024] Referring to the methods in the above embodiments, the following is an example of the application of the methods described in the above embodiments: High-dimensional time-series data on the entire lifecycle operation monitoring of wind turbine main shaft bearings in the field of intelligent manufacturing were selected as the practical application object. The monitoring data simultaneously collected eight indicators, including vibration acceleration, operating temperature, main shaft speed, mechanical load, radial displacement, lubrication flow, ambient humidity, and axial impact force. The equipment completed data sampling at a frequency of once per minute, continuously collecting data for 30 days, and accumulating 43,200 high-dimensional time-series raw data. The key feature mining method for time-series data of this invention was used to carry out key feature mining on the monitoring data. The key features mined were finally applied to two practical engineering tasks: early fault classification of wind turbine main shaft bearings and time-series prediction of remaining service life.

[0025] First, the input high-dimensional time-series raw data of wind turbine bearings is divided into sliding windows. Initially, the sliding windows are set to a continuous, equal-length division mode, with each window containing 60 time steps of data points. Within each sliding window, a dynamic fluctuation threshold of 0.86 is automatically calculated. Based on this threshold, each data point within the window is compared, ultimately identifying 216 noise data points that exceed the fluctuation range. Depending on the actual number of non-noise data points within different windows, the system adaptively uses the corresponding time-series trend fitting method to replace noise points. For windows with at least three non-noise data points, a quadratic time-series trend polynomial is fitted using the local time step index within the window, and the quadratic extrapolated fitted replacement value corresponding to the noise point is directly obtained after solving the polynomial. For windows with only two non-noise data points, the fitted replacement result is directly obtained through two-point linear interpolation. When there is only one non-noise data point within the window, the feature value of that valid data point is directly used as the fitted replacement value. After noise replacement is completed and the denoised and reconstructed time series data is obtained, the time series structure is further verified. The structural similarity between the denoised and reconstructed data and the original time series data is calculated to be 0.92, which is higher than the preset similarity threshold of 0.85. Therefore, there is no need to adaptively adjust the sliding window length, and the denoised and reconstructed time series data completed in this process is directly output.

[0026] Next, based on the overall temporal span and multi-dimensional distribution characteristics of the denoised and reconstructed time series data, all available temporal granularity and dimension combination schemes are traversed, and the temporal autocorrelation coefficient, dimensional mutual information, and temporal-dimensional coupling coefficient corresponding to each combination are calculated one by one. After simultaneous calculation, the matching degree value of each combination is obtained, and 12 temporal granularity and dimension combinations with matching degree higher than the preset matching degree threshold of 0.72 are selected. Based on these 12 effective combinations, a multi-scale feature basis set is constructed. The set contains feature bases with 5 different temporal granularities and 8 types of dimensional cross combinations, and the scale parameters of all feature bases can be dynamically updated and adjusted in real time according to the changes in the temporal autocorrelation coefficient and dimensional mutual information of the real-time input data.

[0027] Subsequently, independent lightweight temporal coding units were configured for each of the 12 feature bases in the multi-scale feature base set, and parallel encoding processing was carried out on each feature base using a depthwise separable temporal convolutional structure. During the encoding process, the system monitored the remaining computing power of the local computing resource unit in real time, which was 68%. At the same time, the information entropy value of the encoded feature vector was calculated to be 1.23. This information entropy value was lower than the preset entropy threshold of 1.50. The system increased the temporal convolution dilation coefficient to 5 according to the preset fixed step size. In the subsequent operation phase, if the remaining computing power ratio dropped to 35% and the information entropy of the feature vector rose to 1.58, the system automatically reduced the dilation coefficient to 2. Throughout the entire dynamic parameter adjustment process, the dimension of the encoded output feature vector remained fixed at 64 dimensions. The lightweight temporal coding unit is composed of a stack of channel-wise temporal convolutional layers and pointwise convolutional layers. It first extracts the temporal local features of each channel independently, and then fuses cross-channel feature information. The feature encoding is completed while controlling the amount of computational parameters and computational complexity. Moreover, all coding units share the same set of dynamic parameter adjustment mechanisms to ensure that the encoding effect of features at different scales is consistent.

[0028] After obtaining standardized feature vectors through feature encoding, a temporal correlation network is constructed based on these vectors. Each time step's encoded feature vector is treated as an independent node in the network. The temporal evolution gradient, dimensional nonlinear dependency coefficient, and global temporal evolution consistency factor between any two nodes at different times are calculated sequentially. The weights of directed edges between nodes are then directly derived from the combined calculation of these three indicators. A preset threshold of 0.45 for network edge weights is set, retaining node associations with edge weights exceeding this threshold. Directed connections are generated from earlier nodes to later nodes in chronological order. This ultimately constructs a temporal correlation network containing 43,200 network nodes and 186,000 directed edges, fully uncovering the deep dependencies hidden between different operating times and monitoring dimensions of wind turbine bearings.

[0029] A hierarchical traversal operation was performed on the constructed temporal correlation network. The network was divided into 120 hierarchical units, with each 6-hour period as an independent time interval. A bottom-up hierarchical traversal strategy was adopted. First, all network nodes in the lowest-level time interval were calculated to obtain an initial local correlation significance score for each node, with all node scores ranging from 0.15 to 0.94. Then, the network was traversed layer by layer from bottom to top. The correlation scores of lower-level nodes were distributed and passed to upper-level nodes according to the edge weight normalization ratio. The upper-level nodes accumulated their initial scores and the scores passed from the lower levels to obtain the final local correlation significance score. Finally, the global correlation significance score of all nodes was calculated by combining the proportion coefficient of each level's time span to the entire time series span. Based on the global correlation significance scores, the top 30 features were selected as the core key features of the wind turbine bearing's operating status, mainly including highly discriminative monitoring features such as high-frequency vibration amplitude, temperature change rate, load fluctuation coefficient, and radial displacement offset.

[0030] Thirty key features selected from the screening were applied to the wind turbine bearing fault classification and remaining life prediction task. These key features were strictly assembled into a complete feature sequence according to temporal order. Simultaneously, the edge weights between corresponding nodes of key features in the temporal correlation network were used as prior weights and embedded into the attention mechanism of the fault classification and temporal prediction model. This allows the model to allocate attention resources rationally based on the prior weights during inference, prioritizing feature combinations with higher correlation significance. During the task execution phase, the classification accuracy and the mean absolute error of the prediction task were used as performance feedback indicators. Initially, the bearing fault classification accuracy was 87.2%, failing to meet the preset performance standard of 90%. The system first lowered the multi-scale feature basis matching threshold to 0.65, broadening the selection range of temporal granularity and dimensional combinations. After adjustment, the classification accuracy improved to 88.5%, still not meeting the preset requirement. Then, the edge weight threshold of the temporal correlation network was lowered to 0.38, retaining more weakly correlated feature dependencies and fully exploring potential feature associations. After adjustment, the classification accuracy rose to 91.3%, meeting the preset performance standard. If model overfitting or abnormal drop in classification accuracy occurs later, the system will reversely increase the matching degree threshold and edge weight threshold to filter out meaningless false weak association edges. If the mean absolute error of the remaining lifetime time series prediction exceeds the preset allowable range, the system will automatically return to the multi-scale feature base generation stage, iterate and optimize the feature scale and combination again, and repeat the entire process of encoding, network construction and feature selection until the performance indicators of the two tasks are stable and meet the preset standards in the long term.

[0031] In summary, the methods described in the above embodiments divide high-dimensional time-series data into sliding windows and dynamically identify, fit, and replace noise points, thus preserving the original time-series structure of the data. The methods can adaptively adjust the window length to verify the similarity of the time-series structure, avoiding the destruction of the inherent time-series correlations during noise reduction. Based on the time-series span and dimensional distribution of the data, dynamic multi-scale feature bases are generated, with scale parameters adapting autonomously to data characteristics. Lightweight encoding is used to process multi-scale features in parallel, balancing computational resource consumption and feature representation capabilities in real time. Simultaneously, deep dependencies between features at different times and dimensions are mined using a time-series correlation network. The significance of feature correlations is quantified through hierarchical traversal, smoothly completing the selection of key features. Finally, time-series task performance indicators are applied to form a closed-loop feedback loop, dynamically fine-tuning feature mining parameters to achieve iterative optimization. The selected key features can be integrated into the task model's attention mechanism, reasonably guiding the model to focus on highly correlated feature combinations, thereby steadily improving the actual performance of time-series prediction and classification tasks.

[0032] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A key feature mining method for time-series data, characterized in that, include: Step 1: Perform sliding window partitioning on the input high-dimensional time series raw data. Within each window, identify noise points by dynamically calculating the data fluctuation threshold. Replace the noise points with the time series trend fitting value of the non-noise data within the window, preserve the original time series structure of the data, and output the denoised and reconstructed time series data. Step 2: Generate a multi-scale feature base set based on the temporal span and dimensional distribution of the denoised and reconstructed time series data. Each feature base corresponds to a different combination of time granularity and dimension, and the scale parameter of the feature base is dynamically adjusted according to the data characteristics. Step 3: The feature bases at each scale are encoded in parallel using a lightweight temporal coding unit. During the encoding process, a dynamic parameter adjustment mechanism is used to balance the consumption of computing resources and the feature representation capability in real time, dynamically adapting to the preset standards of computational load and feature discriminative power, and obtaining the encoded feature vector. Step 4: Construct a temporal correlation network based on the encoded feature vectors to explore the deep dependencies between features at different times and in different dimensions. The network nodes are the feature vectors at each time point, and the edge weights are determined by calculating the combined value of the temporal evolution gradient and the dimensional co-coefficient between nodes. Step 5: Perform a hierarchical traversal on the temporal correlation network, calculate the correlation significance score of each node's features, and select a preset number of key features from high to low scores. Step 6: Apply the selected key features to time series prediction or classification tasks, dynamically adjust the feature mining parameters based on real-time feedback of task performance indicators, and return to Step 2 to re-optimize the feature scale if the performance does not meet the preset standard.

2. The key feature mining method for time-series data according to claim 1, characterized in that, In step 1, when outputting denoised and reconstructed time-series data, the high-dimensional time-series original data is divided into continuous sliding windows of equal length, and a dynamic fluctuation threshold is calculated for the data within each sliding window. ; In the formula: This is the median of the absolute values ​​of the first-order differences of all data points within the current sliding window. , , This is a preset dimensionless adjustment coefficient; The dimensionless local trend curvature of the data within the current sliding window; This represents the local trend abrupt change entropy of the data within the current sliding window. When the absolute value of the difference between a data point within the window and its two adjacent data points is greater than the dynamic fluctuation threshold When this happens, the data point is determined to be a noise point; Noise points are replaced by the time-series trend fitting value of non-noise data within the window. The time-series trend fitting value is dynamically selected based on the number of non-noise data points within the window: when the number of non-noise data points within the window is greater than or equal to 3, a quadratic time-series trend polynomial is fitted using all non-noise data points within the window with the local time step index within the window as the independent variable. The coefficients of each term of the quadratic polynomial are obtained by solving the problem using the unweighted least squares method. The local time step index within the window corresponding to the noise point is substituted into the quadratic polynomial, and the quadratic extrapolation replacement value is calculated as the time-series trend fitting value; when the number of non-noise data points within the window is 2, the linear interpolation result of two non-noise data points is used as the time-series trend fitting value; when the number of non-noise data points within the window is 1, the eigenvalue of the non-noise data point is used as the time-series trend fitting value.

3. The key feature mining method for time-series data according to claim 1, characterized in that, Step 1, before outputting the denoised and reconstructed time-series data, also includes a time-series structure verification step: The temporal structure similarity between the denoised reconstructed time series data and the original high-dimensional time series data is calculated. The temporal structure similarity is obtained by calculating the reciprocal of the dynamic time warp distance between the two data sequences. When the time series structure similarity is greater than a preset similarity threshold, output denoised and reconstructed time series data; When the time sequence structure similarity is less than or equal to the preset similarity threshold, the length of the sliding window is adaptively adjusted according to the intensity of local data fluctuations. The more intense the local fluctuations, the shorter the sliding window length is; the more gradual the local fluctuations, the longer the sliding window length is. The window division, noise identification, and noise point replacement steps are then re-executed.

4. The key feature mining method for time-series data according to claim 1, characterized in that, In the execution phase of step 2, all possible combinations of time granularity and dimension are traversed, and the time granularity of each group is calculated. Combination with dimensions Matching degree ; In the formula: To denoise and reconstruct the time series autocorrelation coefficient of the time series data at this time granularity; To denoise and reconstruct the dimensional mutual information of time series data for this dimensional combination; This is a preset dimensionless minimum value; These are the time-series-dimensional coupling coefficients. ; in, This represents the gradient vector of the time series autocorrelation coefficient as the time granularity changes. This is the gradient vector of dimensional mutual information when the dimensional combination changes; Select matching degree All time granularities and dimension combinations greater than the preset matching degree threshold generate a corresponding multi-scale feature base set, and the scale parameters of the feature base are updated in real time with the changes in the temporal autocorrelation coefficient and dimensional mutual information of the input data.

5. The key feature mining method for time-series data according to claim 1, characterized in that, In step 3, the feature bases at each scale are encoded in parallel using lightweight temporal coding units. The specific steps involved in balancing computational resource consumption and feature representation capability in real-time through a dynamic parameter adjustment mechanism during the encoding process are as follows: The lightweight temporal coding unit adopts a depthwise separable temporal convolutional structure, and the dynamic parameter adjustment mechanism is as follows: Real-time monitoring of the remaining computing power percentage of the current computing resource unit and the information entropy of the encoded feature vector; When the remaining computing power ratio is greater than the preset computing power threshold and the feature information entropy is less than the preset entropy threshold, the dilation coefficient of the temporal convolution is increased. When the remaining computing power percentage is less than the preset computing power threshold and the feature information entropy is greater than the preset entropy threshold, the dilation coefficient of the temporal convolution is reduced. The dimension of the feature vector of the encoded output remains unchanged throughout the parameter adjustment process; The dilation coefficient of the temporal convolution is adjusted in a stepwise manner. The step size for each adjustment is a preset step size, and the value range of the dilation coefficient is limited to between the preset minimum dilation coefficient and the preset maximum dilation coefficient. When increasing the dilation coefficient, the preset step size is increased each time until the dilation coefficient reaches the preset maximum dilation coefficient or the information entropy of the encoded feature vector increases to the preset entropy threshold. When decreasing the dilation coefficient of the temporal convolution, the preset step size is decreased each time until the dilation coefficient reaches the preset minimum dilation coefficient or the remaining computing power ratio of the current computing resource unit increases to the preset computing power threshold. Each scale feature base is input into an independent lightweight temporal coding unit for parallel processing, and all coding units share the same dynamic parameter adjustment mechanism.

6. The key feature mining method for time-series data according to claim 1, characterized in that, When step 4 is executed, the encoded feature vector at each time step is used as a node in the temporal correlation network, and the calculation is performed between any two nodes. and Edge weights between The edge weight ; In the formula: For nodes To the node The temporal evolution gradient; For nodes and The dimensional nonlinear dependence coefficient; It serves as the global temporal evolution consistency factor. When edge weight When the weight is greater than the preset edge weight threshold, at the node With nodes A directed edge is generated between the nodes, with the edge pointing from the earlier node to the later node.

7. The key feature mining method for time-series data according to claim 1, characterized in that, The temporal correlation network is subjected to a hierarchical traversal, and the correlation significance score of each node's features is calculated in the following stages: The temporal correlation network is divided into several levels according to the time sequence, and each level corresponds to a preset time interval; Based on a bottom-up hierarchical traversal strategy, we first traverse all nodes corresponding to the lowest time interval and calculate the local association significance score of each node. By traversing upwards layer by layer, the local association saliency scores of lower-level nodes are passed to the upper-level associated nodes, ultimately obtaining the global association saliency score for each node.

8. The key feature mining method for time-series data according to claim 7, characterized in that, The initial local association significance score of the node is the geometric mean of the product of the weights of all incoming edges and the product of the weights of all outgoing edges of the node. The global association significance score of a node is equal to the final local association significance score of that node multiplied by the proportion of the time span of the node's level to the total time series span.

9. The key feature mining method for time-series data according to claim 1, characterized in that, The feature mining parameters are dynamically adjusted based on real-time feedback of task performance metrics. The task performance metrics include the mean absolute error of a time-series prediction task or the classification accuracy of a time-series classification task. When the task performance indicators do not meet the preset standards, adjust the matching degree threshold of the feature base; If the task performance still fails to meet the preset standard after adjusting the matching degree threshold, then adjust the edge weight threshold of the temporal correlation network. If the task performance still fails to meet the preset standard after adjusting the edge weight threshold, the multi-scale feature base set is regenerated, and the subsequent encoding, temporal correlation network construction, and feature selection steps are repeated.

10. The key feature mining method for time-series data according to claim 1, characterized in that, When applying the selected key features to time-series prediction or classification tasks, the selected key features are concatenated into a feature sequence according to the time sequence. At the same time, the edge weights between the nodes corresponding to the key features in the time-series association network are used as prior weights and embedded into the attention mechanism of the model used in the time-series prediction or classification task. This allows the model to allocate different attention resources according to the prior weights during the inference process, giving priority to feature combinations with higher correlation significance.

Citation Information

Patent Citations

  • Time series trend feature extraction method based on important point double evaluation factor

    CN108804731A

  • Time series data trend feature extraction method based on dynamic grid division

    CN112765562A