A Fuzzy Piecewise Time Series Classification Method and System Based on Consistency Learning

Through the combination of consistency learning and Con-Transformer, the context-aware problem of fuzzy boundaries in time series classification is solved, achieving a more robust and accurate classification effect.

CN117171649BActive Publication Date: 2025-07-11ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311209391.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-19
Publication Date
2025-07-11
Estimated Expiration
2043-09-19

AI Technical Summary

Technical Problem

Existing time series classification methods lack context-awareness when processing fuzzy boundary data, resulting in poor classification results.

Method used

Using a fuzzy segmented time series classification method based on consistency learning, continuous representations are extracted through Con-Transformer, combined with neighbor tag consistency discrimination and constraint prediction behavior, hierarchical training is performed to obtain a more robust and consistent model.

Benefits of technology

Effective modeling and precise classification of boundary fuzzy time series data is realized, and classification accuracy is improved in noise label environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117171649B_ABST
    Figure CN117171649B_ABST
Patent Text Reader

Abstract

The present invention discloses a fuzzy segmented time series classification method and system based on consistency learning, belonging to the technical fields of time series and pattern recognition. Given time series data is divided into several time periods; an encoder module is used to extract local representations of each time period, and based on the local representations, global correlation representations of each time period are encoded; a context-aware coherent prediction module is used to perform self-prediction and context prediction according to the global correlation representations of each time period, and the monotonicity of the context prediction is fitted according to the hyperbolic tangent function to obtain the constrained context prediction; the final prediction of each time period is obtained by combining the self-prediction and the constrained context prediction. The present invention designs consistency label learning based on noise label learning and curriculum learning techniques to coordinate labels from different sources, so as to obtain a more robust and consistent classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of time series and pattern recognition, and in particular, to a fuzzy segmented time series classification method and system based on consistency learning. Background Art

[0002] Time series classification (TSC) is a popular research direction in the field of data mining. In recent years, with the rapid growth of time series data, its importance has become increasingly prominent. For this reason, many algorithms have been proposed by researchers [Hassan 2019]. Traditional algorithms such as Rocket and its variants [Dempster 2020, Dempster 2021] and integrated methods such as HIVE-COTE [Lines 2018] have achieved high accuracy in TSC. They utilize random convolutional kernels with relatively low computational costs and weight individual classifiers at the same time. In addition, the improvement of the non-linear modeling ability of deep models has led to the increasing popularity of deep learning-based TSC algorithms. A variety of techniques have been adopted in TSC: methods based on recurrent neural networks (RNNs) [Rajan 2018, Shen 2020] capture temporal changes through state transitions; methods based on multi-layer perceptrons (MLPs) [Challu 2022, Zhang 2022] encode temporal dependencies as parameters of MLP layers; the latest TimesNet method [Wu 2023] transforms one-dimensional time series into two-dimensional space and achieves state-of-the-art performance on five mainstream tasks. In addition, Transformer models based on the attention mechanism [Nie 2023, Shao 2022, Wu 2022, Zhou 2022] have been widely used. Among them, we note that PatchTST [Nie 2023] and STEP [Shao 2022] use a segmentation mechanism similar to the model proposed in this paper. However, these models only utilize the dependencies between time points to incorporate context information during the modeling process and do not use it for the classification of time periods. Therefore, they lack context awareness to a certain extent.

[0003] In machine learning, learning with noisy labels is an important and challenging research topic because real-world data often relies on error-prone manual annotations. Early understanding work on noisy labels mainly focused on statistical learning [Angluin 1988, Han 2021, Lawrence 2001]. Sukhbaatar et al. in 2015 pioneered a new era in the learning of noisy label representations [Sukhbaatar 2014]. An important tool for dealing with the label noise problem is the label noise transition matrix, which represents the transition probability from clean labels to noisy labels [Han 2020]. Common loss correction techniques include forward and backward correction [Patrini 2017], and using prior knowledge to mask invalid class transitions is also an important method [Han 2018]. Modifying the objective function is another common strategy. For example, adding explicit or implicit regularization terms can reduce the model's sensitivity to noise, while reweighting mislabeled data can reduce its impact on the objective function [Azadi 2016, Liu 2022, You 2020]. Other methods involve training on small-loss instances and leveraging memory effects. MentorNet [Jiang 2018] pre-trains an auxiliary network to select clean instances for the main network's training. Co-teaching [Han 2018] and Co-teaching+ [Yu 2019] introduce two neural networks with different learning abilities to train simultaneously, and they filter noisy labels from each other. To our knowledge, the only work on learning with noisy labels in the time series domain is SREA [Castellani 2021], which uses shared embedding representations to gradually self-label mislabeled data samples and trains classifiers and autoencoders in a self-supervised manner.

[0004] In 2009, Bengio et al. proposed the concept of Curriculum Learning (CL) [Bengio 2009]. This method mimics the human learning process, starting from simple samples and gradually transitioning to complex samples. Based on this concept, curriculum learning can remove noisy data because learners are encouraged to train on simpler data and spend less time on noisy samples [Gong 2016, Wang 2021]. Currently, the mainstream methods include Self-paced Learning [Kumar 2010], where students arrange their learning according to their own learning progress; Transfer Teacher based on a predefined training plan [Weinshall 2018]; and RL Teacher that incorporates the feedback information of students into the framework [Graves 2017, Matiisen 2019]. Summary of the Invention

[0005] To solve the problem that the existing technology has poor classification effect on fuzzy boundary time series, the present invention proposes a fuzzy segmented time series classification method and system based on consistency learning. By using Con-Transformer to obtain a more continuous representation, using the technology of neighbor label consistency discrimination and constraint prediction behavior to obtain a more coherent prediction, and combining the training process of progressive consistency label learning, a more robust and consistent model is obtained, realizing accurate classification on fuzzy boundary time series.

[0006] The present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a fuzzy segmented time series classification method based on consistency learning, including:

[0008] Dividing the given time series data into several time periods;

[0009] Using an encoder module to extract the local representation of each time period, and encoding the global association representation of each time period based on the local representation;

[0010] Using a context-aware coherent prediction module to perform self-prediction and context prediction according to the global association representation of each time period, fitting the monotonicity of the context prediction according to the hyperbolic tangent function, and obtaining the constrained context prediction;

[0011] Combining the self-prediction and the constrained context prediction to obtain the final prediction of each time period.

[0012] Further, the encoder module and the context-aware coherent prediction module adopt a hierarchical training method, including:

[0013] Given a time series data, divide it into several subsequences according to different categories, divide each subsequence from the center point to the boundary into different levels from low to high, and sample the same number of time periods from each level of subsequences as samples;

[0014] Divide the training batches according to different levels from low to high, and update the coordinated labels for each round of training according to the original labels, self-prediction results, context-prediction results after constraint, and training progress of the samples in each time period;

[0015] For each round of training, calculate the cross-entropy loss function of self-prediction according to the coordinated labels of the current round and the self-prediction results obtained by the context-aware coherent prediction module; splice the global correlation representations of different time periods pairwise, judge whether the two time periods belong to the same category according to the spliced representations to obtain the consistency discrimination result, and calculate the cross-entropy loss function of context prediction by using the consistency discrimination result and the true consistency label; train the encoder module and the context-aware coherent prediction module based on the two cross-entropy loss functions.

[0016] Furthermore, the context-aware coherent prediction module uses the global correlation representations of each time period for context prediction, specifically:

[0017] Splice the global correlation representations of different time periods pairwise, perform neighbor label consistency discrimination on the spliced representations to obtain the consistency discrimination result, use the obtained consistency discrimination result as the weight vector, and take the product of the weight vector and the self-prediction result of the corresponding time period as the context prediction result of this time period.

[0018] Furthermore, the encoder module includes a CNN network, a CT1 network, an average pooling layer, and a CT2 network;

[0019] The CNN network adopts a multi-layer one-dimensional convolutional neural network, the CT1 network and the CT2 network are based on Transformer and implemented with a dual-branch self-attention layer, and the average pooling layer is used to implement the average pooling operation;

[0020] For the original time series containing several time periods, first use the CNN network to map each time period to the latent representation space to obtain the low-level latent representations of each time period, then use the CT1 network to encode the low-level latent representations of each time period to obtain the local encodings of each time period, use the average pooling layer to perform average pooling on the local encodings of each time period respectively to obtain the local representations of each time period, and use the CT2 network to obtain the global correlation representations of each time period according to the local representations of all time periods.

[0021] Furthermore, the dual-branch self-attention layer includes:

[0022] For the input representation of the double-branch self-attention layer, the first aggregated representation is calculated through a standard Transformer;

[0023] The input representation is transformed by different learnable matrices to obtain the scale parameter vector and the value vector of the Gaussian kernel function. The Gaussian kernel function is calculated and the calculated values are normalized to obtain the aggregation weights. The product of the aggregation weights and the value vector is the second aggregated representation;

[0024] The attention mechanism is used to adaptively fuse the first aggregated representation and the second aggregated representation, and the fused aggregated representation is output.

[0025] In a second aspect, the present invention provides a fuzzy piecewise time series classification system based on consistency learning for implementing the above-mentioned fuzzy piecewise time series classification method.

[0026] Compared with the prior art, the beneficial effects of the present invention are as follows: By hierarchically sampling the training data, extracting continuous representations, making coherent predictions, and training based on consistent labels, the present invention can effectively model and accurately classify the time series data with blurred boundaries from different sources. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic diagram of a model encoder and a Con-Transformer shown according to an exemplary embodiment;

[0028] Figure 2 is a schematic diagram of neighbor label consistency discrimination, constraint prediction behavior, and consistency label learning shown according to an exemplary embodiment. DETAILED DESCRIPTION

[0029] The present invention will be further described below in conjunction with the drawings and embodiments. The drawings are only schematic diagrams of the present invention. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0030] In this embodiment, the model used in the fuzzy segmented time series classification method based on consistency learning is denoted as "Con4m", which is a deep learning framework based on hierarchical sampling, context awareness, and consistent training labels. The overall framework of the present invention (Con4m) consists of four main modules: a preprocessing module for hierarchical sampling and segmentation of the original time series, an encoder module for encoding local and global correlation representations of time periods, a context-aware coherent prediction module for integrating neighbor prediction information, and a consistency label learning module for progressive coordination of training labels. In addition, those skilled in the art can also add other modules for achieving the purpose of the present invention, such as a data acquisition module, an inference module, a display module, etc. Specifically, the present invention first performs hierarchical sampling and segmentation on the original time series through the preprocessing module; then, uses the encoder module to extract local and global correlation representations of a single time period and its neighbors; then, adopts the coherent prediction module to integrate the prediction information of the time period and the context; at the same time, uses the consistency label learning module to coordinate the training labels to obtain a more robust model; finally, the inference module inputs the obtained representation and gives the prediction result of the fuzzy segmented time series. This embodiment will sequentially elaborate on the specific methods of each module in order.

[0031] I. Preprocessing module.

[0032] For the original complete time series, it is divided into different time intervals according to the continuous states belonging to the same category. Taking the epilepsy prediction task based on intracranial electroencephalogram (EEG) time series as an example, the EEG signals at different time indices include two categories: normal and seizure. According to the continuous states of each category, it is divided into different continuous time intervals. For example, a segment of intracranial EEG time series is divided into a normal segment with a length of K1, a seizure segment with a length of K2, and a normal segment with a length of K3. Since the intracranial EEG time series belongs to the fuzzy segmented time series, the normal segment here means that the data at the vast majority of time indices belongs to the normal state, and the seizure segment means that the data at the vast majority of time indices belongs to the seizure state. The present invention is not limited to intracranial EEG time series and is applicable to any time series containing different categories.

[0033] For a continuous time interval with a length of K, it is divided into 5 equal-length levels:

[0034]

[0035] Sample the same number and equal-length time series from data sets at different levels, and the time series sampled for different categories in each level are balanced. For each time series, set a sliding window with a length of W and a sliding distance of S, segment it, and take each time period as a sample; since each data point in a time period of length W corresponds to a label, select the label that occupies more than W / 2 of this period from the labels corresponding to the data points in the time period as the label of the sample for this time period..

[0036] II. Encoder Module

[0037] First, introduce the structural design of Con-Transformer.

[0038] Con-Transformer uses the standard Transformer as the backbone and is mainly divided into two calculation branches.

[0039] The first calculation branch is the calculation process of the standard Transformer;

[0040] The second calculation branch first learns the scale parameters of different Gaussian kernel functions for each time period representation, then calculates the distance between each time period and its neighbors under the learned Gaussian kernel functions, and after normalizing the distance, uses it as a weight to aggregate the context time periods.

[0041] Specifically, the l-th dual-branch self-attention layer Con-Attention is calculated as follows:

[0042]

[0043]

[0044]

[0045]

[0046] Among them, L is the length of the input sequence, d hidden is the dimension of the hidden representation, is the hidden representation of the (l - 1)-th layer, Q, K, V are all learnable matrices, t , V s are the query, key, and two value vectors respectively. σ is the scale parameter vector of the Gaussian kernel function, T l represents the aggregation weight of the standard Transformer, S l are the aggregation weights obtained by the Gaussian kernel function respectively, are the aggregation representation of the standard Transformer and the aggregation representation obtained by the Gaussian kernel function respectively, σ iis the scale parameter of the Gaussian kernel function for the i-th time step, SoftMax(.) is the Softmax function, and Rescale(.) is the function for row-wise normalization of the matrix.

[0047] Next, the attention mechanism is used to adaptively fuse these two representations:

[0048]

[0049]

[0050] where are all learnable parameters, [.||.] represents vector concatenation, represents the attention weights obtained from the attention weights of the standard Transformer and the Gaussian kernel function, and z l represents the fused aggregated representation.

[0051] By stacking Con-Attention layers, Con-Transformer can be defined. Assume the input sequence is The structure of the l-th layer is defined as follows:

[0052]

[0053]

[0054] where, represents the fused aggregated representation after layer normalization.

[0055] For the initial input, learnable absolute positional encoding (Positional Embedding) is used to embed c. In practice, a multi-head version of Con-Transformer with one Con-Attention layer is used.

[0056] Next, the model architecture of Con4m is introduced ( Figure 1 ). Figure 1 The leftmost part shows the details of Con-Attention, Figure 1 and the right part is the architecture of Con-Transformer and the entire Con4m model architecture.

[0057] Assume the original time series of the input contains L consecutive time periods of length W to be classified, and d is the dimension of the original features of each time period in the original time series. First, the non-linear encoder g enc (multi-layer 1D-CNN) is used to map each time period in x to the latent representation space to obtain the low-level latent representation is the sequence length after downsampling by g enc The sequence length after downsampling.

[0058] Then, Con-Transformer (CT1) is used to encode the low-level latent representations of the time periods to obtain local encodings for each time period. Mean pooling is performed on the local encodings of each time period respectively to obtain local representations After that, another Con-Transformer (CT2) is applied to capture the global patterns among the entire sequence. The encoder of Con4m can be formalized as:

[0059] z = [g enc (x i )] i∈1,…,L

[0060] c local = [Mean-Pooling(Con-Transformer_1(z i ))] i∈1,…,L

[0061] c = Con-Transformer_2(c local )

[0062] where, x i represents the i-th time period of the original time series, c represents the global correlation representation of all time periods, Con-Transformer_1(.) and Con-Transformer_2(.) represent two independent Con-Transformer encoders, and Mean-Pooling(.) represents the mean pooling operation. III. Context-Aware Coherent Prediction Module

[0063] First, neighbor label consistency discrimination is introduced.

[0064] To refer to the prediction results of the context time periods, a self-attention-like technique is adopted. Different from the original self-attention mechanism, a discriminative task is used to supervise the learning of the aggregation weights. If a simple discriminator can distinguish whether the neighbor labels are consistent, then the distances between the representations of samples belonging to the same class will be closer. Therefore, the prediction results of this discriminative task are used as the aggregation weights. This process is formalized as the following equation:

[0065]

[0066]

[0067]

[0068] where, is the probability that two time periods belong to the same class. Q, K, and V represent the query, key, and value respectively, and Q i , K j represents the global correlation representation of the i-th and j-th time periods. (·||·) represents the concatenation operation, represents the self-prediction of each time period, represents the context prediction of each time period, represents taking out the second value of the elements contained in the first two dimensions of

[0069] Define the two cross-entropy losses for model training as:

[0070]

[0071] where y e represents the coordination label for this round of training, represents the true consistency label, represents the indicator function, and y i , y j represent the true labels of the i-th and j-th time periods. Combine these two loss functions as the final loss function l = l1 + l2.

[0072] Next, introduce the prediction behavior constraints.

[0073] As Figure 2 shown, the left side describes the neighbor label consistency discrimination task and prediction behavior constraints based on function fitting, and the rightmost part gives the overall training and inference process based on curriculum learning and noisy label learning. For each class in the prediction results, there are only four different prediction behaviors between consecutive time periods, namely high confidence, low confidence, confidence decrease, and confidence increase.

[0074] To constrain the behavior, use function fitting to integrate the prediction results. Select the hyperbolic tangent function (i.e., Tanh) as the basis. Formally, introduce four adjustable parameters to fit the monotonicity:

[0075]

[0076] where represents the constrained context prediction, a, b, d, and e are four adjustable parameters. Parameter a constrains the value range of the function, b controls the transition slope of the function, d and e adjust the symmetry center of the function, and x is a free vector given in the x-axis. In the present invention, x is set to At the same time, use the MSE loss to fit the reference prediction as follows:

[0077]

[0078] Where l3 represents the MSE loss, and ||.|| represents the norm.

[0079] In the inference stage, the average values of and are used as the final coherent prediction, that is,

[0080] IV. Consistency Label Learning Module

[0081] As Figure 2 shown, in the training stage, Con4m learns the corresponding samples in a hierarchical order from low to high, and there is a lag of 5 rounds between adjacent levels.

[0082] To coordinate inconsistent labels, the idea of noise label learning technology is referred to, and the original labels are gradually changed. Specifically, given the original label y0, the label in the e-th round is updated according to the following formula:

[0083]

[0084]

[0085]

[0086] Where and are both stacked by the corresponding predictions in the recent 5 rounds, represents the self-prediction at each time period in the k-th round, represents the context prediction after constraint at each time period in the k-th round. is the weight vector of exponential averaging, used to aggregate the information in the recent 5 rounds. is the aggregation result of the self-predictions at each time period in the recent 5 rounds, represents the aggregation result of the context predictions after constraint at each time period in the recent 5 rounds; y e represents the coordinated label updated in the e-th round, and the dynamic weighting factor η is used to adjust the degree of label update. As Figure 2 shown, η increases linearly from 0 to 1, and the influence of the original label gradually weakens. Set the increasing number of rounds E = 30.

[0087] In a specific implementation of the present invention, the fuzzy segmented time series classification method based on consistency learning includes:

[0088] Step 1: Obtain the fuzzy segmented time series dataset. Divide the complete time series into different time intervals according to the continuous states of different categories. For each continuous state, divide it into 5 equal-length sub-segments from the center point to the boundary. Consider the sub-segments close to the center point as low-level and those far from the center point as high-level. Sample the same number and equal-length time series from the data sets of different levels, and the time series sampled for different categories in each level are balanced. For each time series, use a predefined sliding window to divide it into continuous time periods and determine the original labels of each time period;

[0089] Step 2: Based on the divided time periods, use the encoder module to extract the global correlation representation of the time periods;

[0090] Step 3: Based on the context-aware coherent prediction module, calculate the loss function of itself and the context according to the global representation of the time period, and obtain the self-prediction and the context prediction Fit the monotonicity of the context prediction according to the hyperbolic tangent function to obtain the constrained context prediction

[0091] Step 4: Based on the self-prediction and the constrained context prediction, use the consistency label learning module to calculate the coordinated labels with the number of training rounds, which are used to coordinate the labels from different sources, and finally obtain a more robust and consistent model;

[0092] Step 5: For the inference stage, cut the classified fuzzy segmented time series into the time period lengths described in Step 1, and use the Con4m model composed of the trained encoder module, the context-aware coherent prediction module, and the monotonicity fitting function to obtain the self-prediction and the constrained context prediction Average the self-prediction and the constrained context prediction to obtain the final classification prediction

[0093] To extract the representation of the time period in Step 2, the above-mentioned Con-Transformer-based encoder can be used. The specific steps are as follows:

[0094] Step 2.1: For each individual time period, use a multi-layer CNN and Con-Transformer to extract the representation of the time period in turn, and use average pooling operation to obtain the local representation of each time period;

[0095] Step 2.2: For the local representation of each time period, use another Con-Transformer to extract the global correlation representation of each time period.

[0096] ​Step 3 To perform context-aware coherent prediction, the above-mentioned neighbor label consistency discrimination task and constrained prediction behavior design can be adopted to achieve it. The specific steps are as follows:

[0097] Step 3.1 Based on the global correlation representation obtained in Step 2.2, a multi-layer perceptron is used to directly perform self-prediction for each time period, and the cross-entropy loss function is calculated using this prediction and the coordinated label of this round of training.

[0098] Step 3.2 Based on the global correlation representation obtained in Step 2.2, the global correlation representations of different time periods are pairwise concatenated, and another multi-layer perceptron is used to perform neighbor label consistency discrimination on the concatenated representation to obtain the consistency discrimination result, and the cross-entropy loss function is calculated using this discrimination result and the true consistency label.

[0099] Step 3.3 Use the consistency discrimination result obtained in Step 3.2 as the weight to perform weighted summation on the self-prediction obtained in Step 3.1 to obtain the prediction corresponding to each time period containing context information.

[0100] Step 3.4 Based on the prediction containing context information obtained in Step 3.3, use a learnable hyperbolic tangent function to fit and reorganize it, and use the output of the fitted hyperbolic tangent function as the constrained context prediction.

[0101] Step 4 To unify the training labels from different sources, the above-mentioned hierarchical progressive coordinated label y e can be used to achieve it.

[0102] The specific steps are as follows:

[0103] Step 4.1 Based on the time period samples at different levels obtained in Step 1, the samples at different levels are lagged in the order from low level to high level to participate in the training rounds.

[0104] Step 4.2 Based on the time period self-prediction and the constrained context prediction obtained in Step 3.1 and Step 3.4, the training labels of each round are updated by means of linear weighted summation.

[0105] In this embodiment, a fuzzy segmented time series classification system based on consistency learning is also provided. This system is used to implement the above embodiment, and the parts that have been described will not be repeated here. The following terms such as "module" and "unit" can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.

[0106] The described fuzzy segmented time series classification system includes:

[0107] A time series segmentation module for dividing given time series data into several time periods;

[0108] An encoder module for extracting local representations of each time period and encoding global correlation representations of each time period based on the local representations;

[0109] A context-aware coherent prediction module for performing self-prediction and context prediction according to the global correlation representations of each time period, fitting the monotonicity of the context prediction according to the hyperbolic tangent function, and obtaining the constrained context prediction;

[0110] A time period classification module for combining self-prediction and the constrained context prediction to obtain the final prediction of each time period.

[0111] In this embodiment, it further includes a hierarchical training module, and the hierarchical training module includes:

[0112] A training data sampling module for dividing given time series data into several subsequences according to different categories, dividing each subsequence from the center point to the boundary into different levels from low to high, and sampling the same number of time periods from the subsequences at each level as samples;

[0113] A label update module for dividing training batches according to different levels from low to high, and updating the coordinated labels for each round of training according to the original labels, self-prediction results, constrained context prediction results, and training progress of the time period samples;

[0114] A loss calculation and parameter update module for calculating the cross-entropy loss function of self-prediction according to the coordinated labels of the current round and the self-prediction results obtained by the context-aware coherent prediction module for each round of training; splicing the global correlation representations of different time periods pairwise, judging whether two time periods belong to the same category according to the spliced representations to obtain a consistency discrimination result, and calculating the cross-entropy loss function of context prediction using the consistency discrimination result and the true consistency label; training the encoder module and the context-aware coherent prediction module based on the two cross-entropy loss functions and updating the parameters.

[0115] For the implementation processes of the functions and roles of the various modules in the above system, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here. For the system embodiments, since they basically correspond to the method embodiments, relevant parts can refer to the descriptions in the method embodiments. The system embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0116] Embodiments of the system of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.

[0117] The methods described above are merely illustrative. The modules described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0118] The technical effects of the present invention are verified by experiments below.

[0119] (1) Datasets. In this embodiment, in this work, two public and one private boundary-blurred time series data are used to measure the performance of the model.

[0120] fNIRS. The Tufts fNIRS to Mental Workload (Tufts fNIRS2MW [Huang 2021]) data contains records of brain activities and other data of adults performing controlled cognitive workload tasks. They labeled each part of the experiment with one of four possible levels of n-back working memory intensity. In this embodiment, following Huang et al., only the 0-back and 2-back tasks are classified.

[0121] Sleep. The SleepEDF [Kemp 2000] dataset contains polysomnographic sleep recordings of 197 subjects over the entire night, including EEG, EOG, mental electromyogram, event markers, and in some cases, related data such as respiration and body temperature. In this embodiment, following Kemp et al., the Fpz-Cz channel of the EEG and the horizontal channel of the EOG were used.

[0122] SEEG. The private SEEG data used in this embodiment is anonymous and provided by a top hospital. These data include electroencephalogram signal recordings collected from epilepsy patients. A total of 52 to 153 channels were used to record signals at a sampling rate of 1000 Hz or 2000 Hz. A total of 847 hours and 1.2 TB of SEEG data were collected. In this embodiment, the data was sampled to 250 Hz, and epileptic seizures in individual channels were identified.

[0123] Table 1 Overview of the Datasets

[0124]

[0125] (2) Label perturbation. For the two publicly available datasets, their labels are relatively consistent. Following related work on noisy label learning, a certain degree of perturbation was introduced to the original labels of the publicly available data to simulate scenarios of label inconsistency. Specifically, first, boundary points between different classes were found in a complete long time series. Then, with a probability of 0.5, it was randomly determined whether each boundary point moved forward or backward. Finally, a new boundary point position was randomly selected from r% of the time interval length in the boundary movement direction. In this way, the boundary labels can be perturbed to simulate label inconsistency. At the same time, the larger the r% value, the higher the degree of label inconsistency. In this work, experiments with r values of 0, 20, and 40 were conducted on fNIRS and Sleep data.

[0126] (3) Baseline models. In this embodiment, Con4m was compared with the latest models from various fields, including a time series classification (TSC) model SREA [Castellani 2021] with noisy labels, two image classification models with noisy labels: Co-teaching [Han 2018] and SIGUA [Han 2020], a supervised TSC model MiniRocket [Dempster 2021], a time series prediction model PatchTST [Nie 2023], and a time series backbone model TimesNet [Wu 2023].

[0127] (4) Application details. In this embodiment, cross-validation is used to evaluate the ability of the model to generalize new data by dividing the dataset into non-overlapping subsets for training and testing. As shown in Table 1, for fNIRS and SEEG data, this embodiment divides the dataset into 4 groups according to the subjects and conducts experiments according to the 2-training-1-validation-1-test (2-1-1) setting. The Sleep data is divided into 3 groups and the 1-1-1 experimental setting is adopted. Therefore, this embodiment reports the average values of 12 and 6 cross-validation results for the fNIRS and Sleep datasets respectively. Note that for SEEG data, there are already fuzzy boundaries in the original data. To obtain a high-quality test group, this embodiment selects a group for accurate labeling and uses an expensive majority voting procedure to determine the boundaries. Then the test group is set aside and only the validation group is changed, and the average value of 3 experiments is reported.

[0128] (5) Evaluation metrics. This embodiment uses test accuracy (Acc.) and Macro-F1 (F1) scores as evaluation metrics. Macro-F1 is the average F1 score of all classes in the test set. The macro metric is selected because the test set is balanced, ensuring that each class is given the same weight in the evaluation. At the same time, the F1 score is selected to balance the accuracy and recall rate.

[0129] (6) Main classification experiment. The average results of the cross-validation experiments of different methods are shown in Table 2. In Table 2, the best results are shown in bold and the second-best results are underlined. Generally speaking, Con4m is basically better than all baseline models on all datasets and at all interference levels.

[0130] Table 2 Comparison of test accuracy (%) and F1 score (%) of the state-of-the-art methods on three datasets

[0131]

[0132] Results of different methods. For fNIRS, the accuracy and F1 score of Con4m are similar to or better than those of SREA. It is worth noting that SREA is specifically designed for time series data and can better identify inconsistent time periods in a self-supervised manner. Taking the Sleep dataset as an example, when the interference ratios are 0%, 20%, and 40% respectively, Con4m improves the F1 score by 9.71%, 7.38%, and 10.16% compared with the best baseline respectively. The improvement on the SEEG dataset is the largest, and it improves the F1 score by nearly 15.41% compared with the best baseline. This can be attributed to the high inconsistency in the SEEG data, indicating that the model of this embodiment effectively solves the problem of label inconsistency.

[0133] Results for different r%. As r% increases, the performance of Con4m degrades the least. In contrast, for the noisy label learning baselines, with an interference rate less than 20%, the performance is comparable. However, when the ratio is increased to 40%, the performance drops significantly. Specifically, between r values of 20 and 40, the F1 score of Con4m drops slightly by only 2.37%, while the F1 score of the best noisy label learning model drops by 3.01%. For most TSC models, the performance continues to degrade as r% increases. MiniRocket, which is based on non-deep learning, exhibits more stable and robust performance. The F1 score of PatchTST on the fNIRS dataset is very unstable. This may be due to overfitting caused by rapid convergence during the training phase. The stable performance of Con4m means that the model of this embodiment is more suitable for time series data with blurred boundaries because of its consistent context-aware design.

[0134] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A fuzzy segmented time series classification method based on consistency learning for epilepsy prediction tasks, where the prediction results include normal states and seizure states; characterized in that, The classification method includes: Dividing the given intracranial EEG time series data into several time periods; Using an encoder module to extract the local representations of each time period, and encoding the global correlation representations of each time period based on the local representations; the encoder module includes a CNN network, a CT1 network, an average pooling layer, and a CT2 network; the CNN network uses a multi-layer one-dimensional convolutional neural network, and the CT1 network and the CT2 network are based on Transformer as the backbone and use a dual-branch self-attention layer to implement, and the average pooling layer is used to implement the average pooling operation; for the original time series containing several time periods, first use the CNN network to map each time period to the latent representation space to obtain the low-level latent representations of each time period, then use the CT1 network to encode the low-level latent representations of each time period to obtain the local encodings of each time period, use the average pooling layer to perform average pooling on the local encodings of each time period respectively to obtain the local representations of each time period, and according to the local representations of all time periods, use the CT2 network to obtain the global correlation representations of each time period; Using a context-aware coherent prediction module to perform self-prediction and context prediction according to the global correlation representations of each time period, and fitting the monotonicity of the context prediction according to the hyperbolic tangent function to obtain the constrained context prediction; The context prediction process is: splicing the global correlation representations of different time periods in pairs, performing neighbor label consistency discrimination on the spliced representations to obtain the consistency discrimination result, using the obtained consistency discrimination result as the weight vector, and taking the product of the weight vector and the self-prediction result of the corresponding time period as the context prediction result of this time period; Combining the self-prediction and the constrained context prediction to obtain the final prediction of each time period.

2. The method according to claim 1, wherein The encoder module and the context-aware coherent prediction module adopt a hierarchical training method, including: Given a time series data, dividing it into several subsequences according to different categories, dividing each subsequence from the center point to the boundary into different levels from low to high, and sampling the same number of time periods from the subsequences at each level as samples; Dividing the training batches according to different levels from low to high, and updating the coordinated labels of each round of training according to the original labels, self-prediction results, constrained context prediction results and training progress of each time period sample; For each round of training, calculating the cross-entropy loss function of the self-prediction according to the coordinated label of the current round and the self-prediction result obtained by the context-aware coherent prediction module; splicing the global correlation representations of different time periods in pairs, judging whether the two time periods belong to the same category according to the spliced representations to obtain the consistency discrimination result, and using the consistency discrimination result and the true consistency label to calculate the cross-entropy loss function of the context prediction; training the encoder module and the context-aware coherent prediction module based on the two cross-entropy loss functions.

3. The fuzzy segmented time series classification method based on consistency learning according to claim 2, wherein The step of dividing each subsequence from the center point to the boundary into different levels from low to high and sampling the same number of time periods from the subsequences at each level as samples includes: For a subsequence of length K, dividing it into 5 equal-length levels: Sample the same number and equal-length time series from subsequences at each level; For each sampled time series, set a sliding window with a length of W and a sliding distance of S, segment it, and use each time period as a sample.

4. The fuzzy segmented time series classification method based on consistency learning according to claim 2, characterized in that, The training batches are divided according to different levels from low to high. According to the original labels, self-prediction results, context prediction results after constraint, and training progress of samples in each time period, update the coordinated labels for each round of training, specifically: Use the lowest level as the initial training batch, and the remaining levels are added in order from low to high. Train 5 rounds before adding each level, and update the coordinated labels according to the following formula: Among them, represents the self-prediction for each time period in the last 5 rounds, represents the context prediction after constraint for each time period in the last 5 rounds, represents the weight vector of exponential averaging, is the aggregation result of the self-prediction for each time period in the last 5 rounds, represents the aggregation result of the context prediction after constraint for each time period in the last 5 rounds, y0 represents the original label, y e represents the coordinated label after the e-th round of update, η is the dynamic weighting factor used to adjust the degree of label update; Rescale(.) is the function for row-wise normalization of the matrix.

5. The fuzzy segmented time series classification method based on consistency learning according to claim 1, characterized in that The described double-branch self-attention layer includes: For the input representation of the double-branch self-attention layer, calculate the first aggregated representation through a standard Transformer; Transform the input representation through different learnable matrices to obtain the scale parameter vector and value vector of the Gaussian kernel function, calculate the Gaussian kernel function and normalize the calculated value to obtain the aggregation weight, and the product of the aggregation weight and the value vector is the second aggregated representation; Use the attention mechanism to adaptively fuse the first aggregated representation and the second aggregated representation, and output the fused aggregated representation.

6. The fuzzy segmented time series classification method based on consistency learning according to claim 1, wherein The described fitting of the monotonicity of context prediction according to the hyperbolic tangent function is specifically: Introduce four adjustable parameters a, b, d, and e, and use the hyperbolic tangent function to fit the monotonicity of the context prediction to obtain the constrained context prediction Among them, the parameter a constrains the value range of the function, b controls the transition slope of the function, d and e adjust the symmetry center of the function, and x is a free vector given in the x-axis; Use the MSE loss l3 to optimize the fitting result: where ||.|| represents the norm, represents context prediction.

7. A fuzzy segmented time series classification system based on consistency learning for epilepsy prediction tasks, where the prediction results include normal state and seizure state; characterized in that, The classification system includes: A time series segmentation module, which is used to divide the given intracranial electroencephalogram time series data into several time periods; An encoder module, which is used to extract the local representation of each time period and encode the global correlation representation of each time period based on the local representation; the encoder module includes a CNN network, a CT1 network, a mean pooling layer, and a CT2 network; the CNN network uses a multi-layer one-dimensional convolutional neural network, and the CT1 network and the CT2 network are implemented with a Transformer as the backbone and a double-branch self-attention layer. The mean pooling layer is used to implement the mean pooling operation; for the original time series containing several time periods, first use the CNN network to map each time period to the latent representation space to obtain the low-level latent representation of each time period, then use the CT1 network to encode the low-level latent representation of each time period to obtain the local encoding of each time period, use the mean pooling layer to perform mean pooling on the local encoding of each time period respectively to obtain the local representation of each time period, and according to the local representation of all time periods, use the CT2 network to obtain the global correlation representation of each time period; A context-aware coherent prediction module, which is used to perform self-prediction and context prediction according to the global correlation representation of each time period, and obtain the context prediction after constraint by fitting the monotonicity of context prediction according to the hyperbolic tangent function; The context prediction process is as follows: pairwise splice the global correlation representations of different time periods, perform neighbor label consistency discrimination on the spliced representations to obtain a consistency discrimination result, use the obtained consistency discrimination result as a weight vector, and take the product of the weight vector and the self-prediction result of the corresponding time period as the context prediction result of that time period; A time period classification module, which is used to combine self-prediction and the context prediction after constraint to obtain the final prediction of each time period.

8. The fuzzy segmented time series classification system based on consistency learning according to claim 7, characterized in that, It further includes a hierarchical training module, and the hierarchical training module includes: A training data sampling module, which is used to divide a given time series data into several subsequences according to different categories, divide each subsequence from the center point to the boundary into different levels from low to high, and sample the same number of time periods from the subsequences at each level as samples; A label update module, which is used to divide the training batches according to different levels from low to high, and update the coordinated labels of each round of training according to the original labels, self-prediction results, context prediction results after constraint, and training progress of the time period samples; A loss calculation and parameter update module, which is used to calculate the cross-entropy loss function of self-prediction for each round of training according to the coordinated labels of the current round and the self-prediction results obtained by the context-aware coherent prediction module; pairwise splice the global correlation representations of different time periods, judge whether two time periods belong to the same category according to the spliced representations to obtain a consistency discrimination result, and calculate the cross-entropy loss function of context prediction by using the consistency discrimination result and the true consistency label; train the encoder module and the context-aware coherent prediction module based on the two cross-entropy loss functions and update the parameters.

Citation Information

Patent Citations

  • Time sequence classification method, system, medium and device based on multi-representation learning

    CN112925822A

  • Electroencephalogram signal self-supervised representation learning method and system and storage medium

    CN115005839A