An automatic data mining method based on time series data

By adopting an automatic data mining method based on deep learning in the financial industry, using Embedding and field and event aggregation technology to process time series data, the problems of low feature derivation efficiency and model performance bottlenecks in the existing technology are solved, and efficient time series data mining and model performance improvement are achieved.

CN114996325BActive Publication Date: 2025-05-16SHANGHAI PUDONG DEVELOPMENT BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210469145.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-05-16
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

The existing technology has low efficiency in the feature derivation of time series data in the financial industry, and relies on manual analysis, resulting in sparse features and low information value, making it difficult to establish an effective model, and it is easy to experience dimensional explosions, resulting in bottlenecks in model performance.

Method used

The automatic data mining method based on deep learning is adopted to process the original tabular data through Embedding, and the data is converted into tensors suitable for tensorflow processing. The nonlinear transformation and attention mechanism operation are performed using entity embedding and field and event aggregation technology, and finally data classification is performed.

Benefits of technology

The end-to-end processing from the original data to the model results is realized, the representation of time series data is enriched, and the characteristics that are difficult to mine by manual are mined, which significantly improves the representation ability and model performance of the time series model, avoids manual intervention and dimensional explosion problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114996325B_ABST
    Figure CN114996325B_ABST
Patent Text Reader

Abstract

The present invention relates to an automatic data mining method based on time series data, which is used for automatic mining of user data of time series data in the financial industry. The method generates tabular data from the original event sequence data of the user, and processes the tabular data into a tensor suitable for tensorflow processing, taking the height direction as the event and the length direction as the field; based on a deep learning model, the tabular data is embedded, and an aggregation operation is performed on the field data based on the result of the embedding processing, and an event aggregation operation is performed on the result after the field data is aggregated, and finally data classification is performed on the result of the event aggregation operation. Compared with the prior art, the present invention has the advantages of no need for manual intervention and improved model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular to an automatic data mining method based on time series data. Background Art

[0002] Time series refers to a set of data arranged in chronological order, and is an important type of complex data object. As a form of data in a database, it is widely present in various large-scale commercial, social science, medical and engineering databases, such as stock prices, various exchange rates, sales volume, product production capacity and weather data. A large amount of time series type data truly records all important information of the system at each moment. Proposing an efficient data mining method will greatly improve people's understanding of such systems, and then conduct effective prediction and control.

[0003] In the financial industry, the mainstream data mining method for time series data is manual derivation, which uses manual derivation technology to derive statistical features such as summing up fields in the past month or the past three months, and then trains the model with the processed features. However, this processing solution has the following obvious problems:

[0004] 1. The efficiency of feature derivation is low, relying on the analytical ability of business personnel and limited thinking. In addition, the derived features are sparse, and the iv (Information Value) of a single feature is low, making it difficult to establish an effective model. Therefore, a large number of features need to be processed for modeling.

[0005] 2. After a large number of features are processed, dimensionality explosion may occur. The marginal value of further deriving features becomes lower and lower, and the difficulty of manually deriving features increases rapidly, resulting in a performance bottleneck in the model. Summary of the invention

[0006] The purpose of the present invention is to provide an automatic data mining method based on time series data in order to overcome the defects of the above-mentioned prior art.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] An automatic data mining method based on time series data is used for automatic mining of user data of time series data in the financial industry. The method generates tabular data from the user's original event sequence data, and processes the tabular data into a tensor suitable for tensorflow processing, with the height direction as the event and the length direction as the field; based on a deep learning model, embedding processing is performed on the tabular data, and aggregation operations are performed on the field data based on the results of the embedding processing, and event aggregation operations are performed on the results of the aggregation of the field data. Finally, data classification is performed on the results of the event aggregation operation.

[0009] Furthermore, the specific steps of Embedding the table data include:

[0010] 11) Process each cell of data in the table and use Embedding to embed the categorical variables;

[0011] 12) Perform NumEmbedding on the continuous variables, that is, perform linear transformation on each continuous field and a set of parameter matrices, cooperate with the relu activation function, perform nonlinear mapping, and obtain the result of nonlinear mapping:

[0012] 13) The embedding results of each grid of data are assembled into a three-dimensional tensor to obtain the three-dimensional tensor representation of the time series.

[0013] In step 12), the expression for performing nonlinear mapping is:

[0014] x1=NumEmbedding(x)=Relu(xD)

[0015] Where x1 is the result of performing nonlinear mapping, x is each continuous field, and D is the parameter matrix.

[0016] Furthermore, the specific steps of performing aggregation operations on field data based on the results of Embedding processing include:

[0017] 21) Perform nonlinear transformation on the three-dimensional tensor of Embedding in step 1) through multiple sets of parameter matrices and ReLU to obtain the nonlinear transformation result of the three-dimensional tensor;

[0018] 22) Perform a dimensionality reduction operation on the result obtained in step 21) through the attention mechanism and query (Q) operation to obtain a dimensionality reduction result of a two-dimensional tensor.

[0019] In step 22), the expression for performing the dimensionality reduction operation on the result obtained in step 21) is:

[0020] x3=Attention(x2)=Relu(Qx2 T )x2

[0021] Where x3 is the result of the dimensionality reduction operation, x2 is the nonlinear transformation result of the three-dimensional tensor in step 21), Relu is the activation function, and Attention(x2) is the attention mechanism operation on the result x2.

[0022] Furthermore, the specific steps of performing an event aggregation operation on the result after aggregating the field data include:

[0023] 31): Pass the result of field data aggregation into TransformerEncoder to obtain the result x4;

[0024] 32) Perform dimensionality reduction operation on the result x4 obtained in step 31) through the attention mechanism and query (Q) operation to obtain a one-dimensional result x5.

[0025] Furthermore, the TransformerEncoder includes Self-attention and FeedForward.

[0026] The expression of the one-dimensional result x5 is:

[0027] x5=Attention(x4)=Relu(Qx4 T )x4

[0028] Where Relu is the activation function, and Attention(x4) is the attention mechanism operation on the result x4.

[0029] Furthermore, the specific content of data classification of the result of event aggregation operation is as follows:

[0030] The results of the event aggregation operation are passed through a single or multiple nonlinear layers to obtain the final output of the model, that is, multiple sets of parameter matrices are nonlinearly transformed with the Relu activation function to obtain the classification results.

[0031] Furthermore, all data processing processes of the method of the present invention are implemented based on Automl.

[0032] The automatic data mining method based on time series data provided by the present invention has at least the following beneficial effects compared with the prior art:

[0033] 1) The method of the present invention is an automatic data mining method for time series data based on deep learning. It uses the embedding and field and event aggregation technology to achieve end-to-end processing from raw data to model results, which can enrich the representation of time series data and mine features that are difficult to mine manually. By adding entity embedding and field and event aggregation technology, the representation ability of the time series model can be significantly improved, thereby improving the model performance.

[0034] 2) The method of the present invention establishes a universal time series data automatic mining model framework and is implemented based on Automl. A model with better performance can be obtained without human intervention in the entire data mining process. The model results can also be used as features and input into other business models to supplement the information that is not fully mined by artificial feature derivatives. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of a model framework based on deep learning created for an automatic data mining method based on time series data in an embodiment. DETAILED DESCRIPTION

[0036] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should belong to the scope of protection of the present invention.

[0037] Example

[0038] The present invention relates to an automatic data mining method based on time series data, which is used for automatic data mining of time series type data of users of industry content in the financial industry. Compared with the prior art, this method skips the link of artificially derived features and directly models the original business data, thereby improving the model performance. The method of the present invention is mainly reflected in the following three aspects:

[0039] 1) Use deep learning modeling technology to enrich the representation of time series data and mine features that are difficult to mine manually;

[0040] 2) Add entity embedding and field and event aggregation technology to improve the representation ability of time series models and thus improve model performance;

[0041] 3) Establish a general time series data automatic mining model framework.

[0042] Specifically, the present invention uses deep learning as the basis of the model framework, such as Figure 1As shown, the model framework includes an Embedding module, a field aggregation module, an event aggregation module, and a classification module. As a preferred solution, the model framework of the present invention is implemented using the tensorflow deep learning framework, thereby establishing a set of behavior data automatic mining model framework.

[0043] In the method of the present invention, the original event sequence data for a user is a table, and the table is processed into a tensor that can be processed by tensorflow, with the height direction being events and the length direction being fields. Based on this, the specific steps of the method of the present invention are as follows:

[0044] Step 1: Embedding the table data.

[0045] Embedding is a way to convert discrete variables into continuous vector representations. In neural networks, embedding can reduce the spatial dimension of discrete variables while still representing the variable meaningfully. Embedding has the following three main purposes: 1) Finding the nearest neighbor in the embedding space can be used to make recommendations based on user interests. 2) As input for supervised learning tasks. 3) Used to visualize the relationship between different discrete variables.

[0046] In this step, each cell in the table is processed first, and the categorical variables are embedded in entities. The Embedding at this time is the regular Embedding.

[0047] Secondly, NumEmbedding is performed on the continuous variables, that is, each continuous field x is linearly transformed with a set of parameter matrices D, and the relu activation function is used to perform nonlinear mapping to obtain the result x1:

[0048] x1=NumEmbedding(x)=Relu(xD)

[0049] Then, each Embedding result is pieced together into a three-dimensional tensor to obtain the three-dimensional tensor representation of this sequence.

[0050] Step 2: Perform field aggregation operations. The specific steps include:

[0051] Step 21: Perform nonlinear transformation on the three-dimensional tensor of Embedding in step 1 through multiple sets of parameter matrices (D1, D2, ...) and Relu to obtain the nonlinear transformation result x2 of the three-dimensional tensor:

[0052] x2=output(x1)=Relu(Relu(x1D1)D2)(Twice is used as an example here)

[0053] Step 22: Perform a dimensionality reduction operation on the above result x2 through the attention mechanism (attention) and query (Q) operation to obtain the dimensionality reduction result x3:

[0054] x3=Attention(x2)=Relu(Qx2 T )x2

[0055] The operation in this step is called Field Aggregation, and the output result is a two-dimensional tensor, which is called the field aggregation tensor (a two-dimensional matrix representation of a sequence).

[0056] Step 3: Execute event aggregation operations. The specific steps include:

[0057] Step 31: The two-dimensional matrix x3 of the sequence obtained in step 2 is passed into TransformerEncoder. The TransformerEncoder here is Self-attention and FeedForward. The difference from the conventional transformerEncoder is that there is no transformerEncoder. At this time, the input and output sizes of the tensor remain consistent, and the result x4 is obtained.

[0058] Step 32: Perform a dimensionality reduction operation on the output matrix x4 obtained in step 31 through the attention mechanism (attention) and query (Q) operation to obtain the result x5:

[0059] x5=Attention(x4)=Relu(Qx4 T )x4

[0060] At this time, the operation is called item aggregation, and the output result is a one-dimensional vector, which is called the event aggregation vector (a one-dimensional vector representation of a sequence).

[0061] Step 4: Use the top-level classifier to perform classification operations.

[0062] For the result x5 obtained in step 3, the final output of the model is obtained through a single or multiple nonlinear layers, that is, multiple sets of parameter matrices (D1, D2, ...) and Relu are used for nonlinear transformation to obtain the result x6: x6 = output (x5) = Softmax (Relu (Relu (x5D1) D2)) (here two classifications are taken as an example)

[0063] In this embodiment, as a preferred solution, the above process is implemented based on Automl, using the same set of codes to adapt to multiple scenarios and multiple data sources, with high automation, almost no feature engineering required, and no manual intervention required in the training process.

[0064] In summary, the method of the present invention is an automatic data mining method for time series data based on deep learning. It utilizes Embedding and field and event aggregation technology to achieve end-to-end processing from raw data to model results. No human intervention is required in the entire data mining process to obtain a model with better performance. The model results can also be used as features and input into other business models to supplement information that is not fully mined by artificial feature derivatives.

[0065] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person familiar with the technical field can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. An automatic data mining method based on time series data, used for automatic mining of user data of time series data in the financial industry, characterized in that: Generate tabular data from the user's original event sequence data, and process the tabular data into tensors suitable for tensorflow processing, with the height direction as the event and the length direction as the field; Based on the deep learning model, embedding is performed on the table data, and aggregation is performed on the field data based on the results of the embedding processing. Event aggregation is performed on the results of the field data aggregation, and finally data classification is performed on the results of the event aggregation operation. The specific steps for Embedding tabular data include: 11) Process each cell of data in the table and use Embedding to embed the categorical variables; 12) Perform NumEmbedding on the continuous variables, that is, perform linear transformation on each continuous field and a set of parameter matrices, cooperate with the relu activation function, perform nonlinear mapping, and obtain the result of nonlinear mapping: 13) The embedding results of each grid of data are assembled into a three-dimensional tensor to obtain the three-dimensional tensor representation of the time series; The specific steps to perform aggregation operations on field data based on the results of Embedding processing include: 21) Perform nonlinear transformation on the three-dimensional tensor of Embedding in step 13) through multiple sets of parameter matrices and ReLU to obtain the nonlinear transformation result of the three-dimensional tensor; 22) Performing a dimensionality reduction operation on the result obtained in step 21) through an attention mechanism and a query operation to obtain a dimensionality reduction result of a two-dimensional tensor; The specific steps of performing event aggregation operations on the results of field data aggregation include: 31): Pass the result of field data aggregation into TransformerEncoder to obtain the result x4; 32) Perform dimensionality reduction operation on the result x4 obtained in step 31) through attention mechanism and query operation to obtain a one-dimensional result x5.

2. The automatic data mining method based on time series data according to claim 1 is characterized in that: In step 12), the expression for performing nonlinear mapping is: x1=NumEmbedding(x)=Relu(xD) Where x1 is the result of performing nonlinear mapping, x is each continuous field, and D is the parameter matrix.

3. The automatic data mining method based on time series data according to claim 1 is characterized in that: In step 22), the expression for performing the dimensionality reduction operation on the result obtained in step 21) is: x3=Attention(x2)=Relu(Qx2 T )x2 Where x3 is the result of the dimensionality reduction operation, x2 is the nonlinear transformation result of the three-dimensional tensor in step 21), Relu is the activation function, and Attention(x2) is the attention mechanism operation on the result x2.

4. The automatic data mining method based on time series data according to claim 1 is characterized in that: The TransformerEncoder includes Self-attention and FeedForward.

5. The automatic data mining method based on time series data according to claim 1 is characterized in that: The expression of the one-dimensional result x5 is: x5=Attention(x4)=Relu(Qx4 T )x4 Where Relu is the activation function, and Attention(x4) is the attention mechanism operation on the result x4.

6. The automatic data mining method based on time series data according to claim 1 is characterized in that: The specific contents of data classification for the results of event aggregation operation are as follows: The results of the event aggregation operation are passed through a single or multiple nonlinear layers to obtain the final output of the model, that is, multiple sets of parameter matrices are nonlinearly transformed with the Relu activation function to obtain the classification results.

7. The automatic data mining method based on time series data according to claim 1 is characterized in that: All data processing processes of this method are implemented based on Automl.

Citation Information

Patent Citations

  • Method and device for analyzing time series data

    CN104239477A

  • Time series data clustering model establishment method and time series data clustering method

    CN114168822A