A semi-supervised algorithm in a hybrid online data stream scenario

By combining online learning and semi-supervised learning algorithms in hybrid online data flow scenarios, using Gaussian linking GC and local density peak Local-DPC to learn the geometric structural characteristics of data, the problem of labeled and unlabeled data utilization in mixed data flow is solved, and rapid modeling and efficient learning are achieved.

CN115796301BActive Publication Date: 2025-07-11GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211382596.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-07
Publication Date
2025-07-11
Estimated Expiration
2042-11-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize labeled and unlabeled data in hybrid online data flow scenarios, and it is difficult to deal with the correlation between the features of mixed types, resulting in high training costs, low timeliness and few data labels.

Method used

A semi-supervised algorithm in a hybrid online data flow scenario is adopted, combining online learning and semi-supervised learning, and the geometric structural characteristics of data are learned through Gaussian joint GC and local density peak Local-DPC are used to learn the geometric structural characteristics of data, and an online combination algorithm with accelerated convergence is used to process any type of data set and quickly converge.

Benefits of technology

It realizes the efficient use of labeled and unlabeled data in hybrid online data streams, quickly modeling mixed features, and improves the learning efficiency and accuracy of data streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796301B_ABST
    Figure CN115796301B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of semi-supervised learning and online learning, and particularly to a semi-supervised algorithm in the scenario of arbitrary online data streams. The algorithm framework mainly includes four parts: constructing arbitrary data streams, learning latent rules through Gaussian connection GC, learning geometric structure features of data through local density peak Local-DPC, and an online combination algorithm for accelerating convergence. The construction of arbitrary data streams is for datasets with two situations of mixture and missingness that occur in real online application scenarios; learning latent rules through GC is to construct a latent data feature space by using the observed data space through the marginal distribution function; learning the geometric structure of the data feature space through Local-DPC to construct pseudo-labels for missing labels. Finally, an online combination algorithm for accelerating convergence is constructed for models under different data distribution spaces. The semi-supervised algorithm model in the scenario of mixed online data streams not only effectively solves the problem of filling missing data, but also solves the problem of missing labels for missing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of online learning and semi-supervised learning, and specifically to a semi-supervised algorithm in a hybrid online data stream scenario. Background Art

[0002] Online learning with dual-stream input is a newly emerging paradigm for data stream analysis. Different from traditional online learning, this new learning paradigm can only process data streams residing in a fixed feature space and endeavors to build an incremental model related to the stream data and stream features. This allows for a more flexible learning environment in which new features can arbitrarily appear and be incorporated into the model training process, while pre-existing features may become unobservable or disappear from the model over different time spans.

[0003] Under this flexible learning paradigm, applications in various fields have started to model their data in the form of dual streams. For example, consider a people sensing application where mobile users collectively submit their data to train an incremental model for detecting local air pollution. The dual-stream characteristic is manifested in the data stream of people sensing - new users joining the sensing effort with upgraded or brand-new devices (such as mobile phones, sensor kits) will generate new features, while any departing users (or some devices dropping offline due to network issues) will cause the unobservability of features. To learn from such a data stream, a common practice in previous studies was to establish correlations between features so that the incremental model could 1) initialize the learning coefficients of any new features with informed guesses and use a jump start to accelerate convergence when these new features were not described by sufficient data instances; 2) utilize the reconstruction information of unobserved features and their learning coefficients to improve the prediction performance through online ensembles.

[0004] First, the incremental model is trained under full supervision, which means that each arriving data instance must have a class label. Unfortunately, due to limited human resources and time stretched by large volumes and high speeds of data streams, annotating labels is generally difficult. Second, all features flowing into the model are stipulated to share the same data type, which is often violated in practical applications. For example, features captured by various types of sensor devices are naturally of different data types, including boolean (such as raining or not raining), ordinal (such as PM2.5 level), and continuous (such as outdoor temperature). Establishing correlations between such mixed-type features is very challenging and cannot be achieved by an online parametric model that presumes a Gaussian correlation matrix beforehand.

[0005] In view of the problems such as high cost, low timeliness, and few data labels in offline data training, this paper specifically proposes a semi-supervised algorithm for any data stream scenario. This algorithm combines the advantages of online learning and semi-supervised learning. While using labeled data, it can also take into account a large amount of unlabeled data and can well handle online mixed data streams. Summary of the Invention

[0006] (1) Technical problems to be solved

[0007] In view of the deficiencies of the prior art, the present invention provides a semi-supervised algorithm for a hybrid online data stream scenario, which has the characteristics of being able to utilize both labeled and unlabeled samples and having online learning, combines the advantages of both, and solves the problems such as high cost, low timeliness, and few data labels in offline data training. This paper specifically proposes a semi-supervised algorithm for any data stream scenario. This algorithm combines the advantages of online learning and semi-supervised learning. While using labeled data, it can also take into account a large amount of unlabeled data and can well handle the problem of online mixed data streams.

[0008] (2) Technical solutions

[0009] To achieve the above object of being able to utilize both labeled and unlabeled samples and having the characteristics of online learning and combining the advantages of both, the present invention provides the following technical solutions: A semi-supervised algorithm for an online arbitrary data stream scenario, including constructing an arbitrary data stream, learning latent rules through Gaussian connection GC, learning the geometric structure features of data through local density peak Local-DPC, and an online combination algorithm for accelerating convergence. It is characterized in that: the construction of the arbitrary data stream is to construct a corresponding arbitrary type of data set for the online data application scenario (mixed and missing data streams); learning latent rules through GC is to use the GC model to learn the marginal distribution features between different variables from the observation space (missing), and through online maximum expectation Online-EM, find the filling values of the missing values in the observation space (missing); learning the geometric structure features of data through Local-DPC is to use Local-DPC to learn the data geometric structure distribution features of the observation space (complete) and the latent space; the online combination algorithm for accelerating convergence is to construct a fast-converging online learning combination algorithm for different distribution feature spaces of the observation space (complete) and the latent space.

[0010] Preferably, any data stream construction is to construct corresponding datasets of any type for online data application scenarios (mixed and missing data streams). The characteristics of any data stream referred to in this algorithm include data types such as ordinal, binary, continuous, and discrete. In addition, there are missing values in any data stream, and problems such as uncertain missing ratios exist.

[0011] Preferably, learning latent rules through GC is to use the GC model to learn the marginal distribution characteristics between different variables from the observation space (missing), and find the filling values of the missing values in the observation space (missing) through Online-EM (Online Expectation-Maximization). The content involved includes reconstruction of unobserved features and Online-EM parameter evaluation. Reconstruction of unobserved features refers to the reconstruction of missing values in the observed values. The purpose of Online-EM parameter evaluation is to ensure the maximum similarity between the filling space of missing values and the original observed data distribution space.

[0012] Preferably, learning the geometric structure features of data through Local-DPC is to use Local-DPC to learn the data geometric structure distribution features of the observation space (complete) and the latent space; Local-DPC selects the center points of different categories by constructing different clusters, and uses the distances from the center points to other surrounding nodes to construct different categories of clusters, forming the geometric space distribution structure of the corresponding categories.

[0013] Preferably, the online combination algorithm for accelerated convergence constructs a fast-converging online learning combination algorithm for different distribution feature spaces of the observation space (complete) and the latent space. The single data space distribution cannot meet the fast convergence of data. Considering the weights of models in different spaces, dynamically adjust the model weights in different spaces, thereby accelerating the convergence speed of the model.

[0014] (III) Beneficial Effects

[0015] Compared with the prior art, the present invention provides a semi-supervised algorithm in a mixed online data stream scenario, having the following beneficial effects:

[0016] The semi-supervised algorithm in this mixed online data stream scenario can solve the problem that it is difficult to model the mixed data features composed of discrete and continuous types in online learning. It models the observation space composed of mixed data streams through GC and maps it into a continuous latent space. Use Local-DP to explore the real structure of the data space and integrate this process into semi-supervised learning to make full use of unlabeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1Schematic diagram of the overall model of the present invention;

[0018] Figure 2 Schematic diagram of the overall process of the present invention. Detailed implementation manners

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] Please refer to Figure 1-2 , the present invention provides a technical solution: a semi-supervised algorithm in a hybrid online data stream scenario,

[0021] including arbitrary data stream construction, learning latent rules through Gaussian Copula (GC), learning geometric structure features of data through Local Density Peaks Clustering (Local-DPC), and an online combination algorithm for accelerating convergence. It is characterized in that: the arbitrary data stream construction is to construct corresponding datasets of any type for online data application scenarios (hybrid and missing data streams); learning latent rules through GC is to use the GC model to learn the marginal distribution features between different variables from the observation space (missing), and find the filled values of the missing values in the observation space (missing) through Online Expectation-Maximization (Online-EM); learning geometric structure features of data through Local-DPC is to use Local-DPC to learn the geometric structure distribution features of the data in the observation space (complete) and the latent space; the online combination algorithm for accelerating convergence is to construct a fast-converging online learning combination algorithm for different distribution feature spaces in the observation space (complete) and the latent space.

[0022] At the same time, it includes the following steps:

[0023] S1. Arbitrary data stream construction

[0024] The data types included are ordered numerical values, discrete values, binary values, and continuous values;

[0025] The construction of arbitrary data streams aims to construct corresponding datasets of any type in the scenario of online data applications (mixed and missing data streams). The characteristics of the arbitrary data streams referred to in this algorithm include data types such as ordinal, binary, continuous, and discrete values. In addition, there are missing values in the arbitrary data streams, and the missing ratio is uncertain (i.e., there is randomness).

[0026] The data streams in the construction of arbitrary data streams involve data types such as binary, ordinal values, continuous values, and discrete values, and there are uncertain missing situations in the arbitrary data streams themselves.

[0027] S2. Latent space learning

[0028] Use the GC model to learn the marginal distribution characteristics between different variables from the observation space (missing), and find the filling values of the missing values in the observation space (missing) through Online-EM (Online Expectation-Maximization).

[0029] Learning latent rules through GC is to use the GC model to learn the marginal distribution characteristics between different variables from the observation space (missing), and find the filling values of the missing values in the observation space (missing) through Online-EM. The contents involved include the reconstruction of unobserved features and the evaluation of Online-EM parameters. The reconstruction of unobserved features refers to the reconstruction of the missing values in the observed data. The purpose of Online-EM parameter evaluation is to ensure the maximum similarity between the filling space of the missing values and the original observed data distribution space.

[0030] Latent space learning is to use GC and Online-EM to repeatedly iterate multiple rounds in the observation space (missing) to construct the latent space, so as to finally obtain the data in the observation space (complete).

[0031] The contents included are as follows:

[0032] 1) Define the GC model in the scenario of online mixed data streams:

[0033]

[0034] Among them, cutoff(.) is the truncation function, z ∈ R is a continuous normal, the cumulative distribution function (CDF) is Fz and The latent space vector is z t :=g -1 (x t )=(g -1 (x C ), cutoff-1(x D ));

[0035] 2) Reconstruction of Missing Data Space

[0036]

[0037] where z O is the latent space corresponding to the x O observation space, and z M is the missing distribution space corresponding to that in z O , Σ M,O and Σ O,O represent sub - matrices of the correlation Σ, whose rows and columns correspond to the characteristic exponents of (x M , x O ) and (x O , x O ) respectively; is the maximum estimation function;

[0038] 3) Online - EM Parameter Evaluation

[0039] Define where Φ is a standard normal CDF, and F i is the true but unknown CDF corresponding to the i - th feature;

[0040]

[0041] where the scale H = |B| / (|B| + 1) ensures a finite output, and B is the buffer size of the online mixed data stream; for discrete features, the cut - off point S i is defined as a special case of by replacing its sample average with the probability mass

[0042]

[0043] where x t [i] represents the i - th (discrete) feature of the t - th input; to eliminate ambiguity, Σ (t-1) is denoted as the empirical correlation obtained in the previous round, and Σ is denoted as the target to be approximated in this round; for the objective in the current round, the log - likelihood function is expressed as:

[0044]

[0045] where Σ (0) is initialized as an initial matrix;

[0046] S3, Geometric Structure Learning;

[0047] Two metrics are used to describe each arriving instance x tfeatures; i.e., local density ρ t and distance δ t , are defined as:

[0048]

[0049]

[0050] where d(x t , x i ) measures the Euclidean distance between x t and x i in the reconstructed general feature space U t , and d cut is an adaptively adjusted cut-off distance;

[0051] Learning the geometric structure features of data by Local-DPC is to utilize the data geometric structure distribution features in the Local-DPC learning observation space (complete) and the latent space; Local-DPC constructs different categories of central points by different clusters, and uses the distances from the central points to other surrounding nodes to construct different categories of clusters, forming the geometric space distribution structure of the corresponding categories;

[0052] Geometric structure learning is to construct different categories of clusters for the mutual relationships of different nodes in the data, select the class center nodes, calculate the distances between the class center nodes and other nodes, and construct the spatial geometric structure between different nodes.

[0053] S4. Ensemble algorithm

[0054] Let y Z = <f Z , z t > represent the prediction of the latent representation of x t ; the ensemble prediction is where α1 + α2 = 1; the values of α1 and α2 respectively determine the importance of the two base classifiers f O and f Z ;

[0055]

[0056] where,[[]] is an adjusted parameter; the ensemble of the two base classifiers enables the learned classifier to have good performance.

[0057] The online combination algorithm for accelerating convergence is a fast-converging online learning combination algorithm proposed for the observation space (complete) and different distribution characteristic spaces in the latent space. The distribution of a single data space cannot meet the fast convergence of data. Considering the weights of models in different spaces, the weights of models in different spaces are dynamically adjusted to accelerate the convergence speed of the model;

[0058] The ensemble algorithm is for the models learned in different data distribution spaces in the data space, and uses dynamic adaptive model weight tuning to dynamically adjust the model weights in different data distribution spaces, accelerating the fitting and convergence speed of the model to the data.

[0059] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A semi-supervised algorithm in a hybrid online data stream scenario, characterized in that, It includes the following steps: S1. Construction of arbitrary data streams The data types included are ordered numerical values, discrete values, binary values, and continuous values; S2. Latent space learning Use the GC model to learn the marginal distribution characteristics between different variables from the missing values in the observation space, and find the filled values of the missing values in the missing observation space through Online-EM (Online Expectation-Maximization); The content included is: 1) Define the GC model in the online mixed data scenario: Among them, is a truncation function, z ∈ R is a continuous normal, and the cumulative distribution function CDF is Fz and ; the latent space vector is ; 2) Reconstruction of the missing data space Among them, is the latent space corresponding to the observation space, is the missing distribution space corresponding in and represents the correlation sub - matrix, the rows and columns of which respectively correspond to the characteristic exponents of ( , ) and ( , ), is the maximum estimation function; 3) Online-EM parameter evaluation Definition , where is a standard normal CDF , corresponding to the true but unknown for the CDF ; Among them, the scale \(H = |B| / (|B| + 1)\) ensures a finite output, where \(B\) is the buffer size of the online hybrid data stream; for discrete features, the cut-off point is defined as a special case of by replacing the sample mean with the probability mass of the th feature, and is defined as follows: Among them, represents the th feature of the th input; to eliminate ambiguity, is expressed as the experience correlation obtained in the previous round, and is expressed as the target to be approximated in this round; for the target in the current round, the log-likelihood function is expressed as: Among them, is initialized to an initial matrix; S3. Geometric structure learning; Two metrics are used to describe the characteristics of each arriving instance ; namely, the local density and the distance , which are defined as: Among them, Measure And In the reconstructed general feature space U t The Euclidean distance in, d cut Is the cut-off distance that is adaptively adjusted; S4. Ensemble algorithm Let =⟨ , ⟩ denote a prediction on the latent representation of ; the set of predictions is , where ; and values respectively determine the importance of two base classifiers and ; Among them, is an adjusted parameter; the set of two base classifiers gives our learner a nice property.

2. The semi-supervised algorithm in a hybrid online data stream scenario according to claim 1, wherein The construction of arbitrary data streams is to construct corresponding datasets of any type under the mixed and missing data streams in the online data application scenario. The characteristics of the arbitrary data streams referred to in this algorithm include data types such as ordinal (ordered numerical), binary, continuous, and discrete. In addition, there are missing numerical values in the arbitrary data streams, and the missing ratio is uncertain.

3. A semi-supervised algorithm in a hybrid online data stream scenario according to claim 1, characterized in that, Use the GC model to learn the marginal distribution characteristics between different variables from the missing observation space, and find the filled values of the missing values in the missing observation space through Online-EM. The content involved includes the reconstruction of unobserved features and the evaluation of Online-EM parameters. The reconstruction of unobserved features refers to the reconstruction of the missing numerical values in the observed numerical values. The purpose of Online-EM parameter evaluation is to ensure the maximum similarity between the filled space of the missing values and the original observed data distribution space.

4. A semi-supervised algorithm in a hybrid online data stream scenario according to claim 1, wherein Use Local-DPC to learn the data geometric structure distribution characteristics of the complete observation space and the latent space; Local-DPC selects the central points of different categories by constructing different clusters, and uses the distances from the central points to other surrounding nodes to construct different clusters of categories, forming the geometric space distribution structure of the corresponding categories.

5. A semi-supervised algorithm in a hybrid online data stream scenario according to claim 4, wherein, The online combination algorithm for accelerating convergence is to construct a fast-converging online learning combination algorithm for the different distribution feature spaces of the complete observation space and the latent space. The single data space distribution cannot meet the fast convergence of the data. Consider the weights of the models in different spaces and dynamically adjust the weights of the models in different spaces.

6. A semi-supervised algorithm in a hybrid online data stream scenario according to claim 1, characterized in that In S1, the data streams in the construction of arbitrary data streams involve data types such as binary, ordered numerical values, continuous values, and discrete values, and there are uncertain missing situations in the arbitrary data streams themselves.

7. A semi-supervised algorithm in a hybrid online data stream scenario according to claim 1, wherein In S2, latent space learning is to use GC and Online-EM to repeatedly iterate multiple rounds in the missing observation space to construct the latent space, so as to finally obtain the complete data in the observation space.

8. A semi-supervised algorithm in a hybrid online data stream scenario according to claim 1, characterized in that In S3, geometric structure learning is to construct different clusters for the mutual relationships of different nodes in the data, select the central nodes of the clusters, and calculate the distances between the central nodes of the clusters and other nodes to construct the spatial geometric structure between different nodes.

9. A semi-supervised algorithm in a hybrid online data stream scenario according to claim 1, characterized in that In S4, the integrated algorithm is a model learned for different data distribution spaces in the data space. It uses dynamic adaptive model weight tuning to dynamically adjust the model weights of different data distribution spaces, accelerating the fitting and convergence speed of the model to the data.

Citation Information

Patent Citations

  • Network traffic classification method and system based on federal semi-supervised learning

    CN113705712A

  • Drought disaster weather prediction method based on semi-supervised ensemble learning

    CN114841064A