Financial time series data prediction method, device and medium

The financial time series data prediction method addresses resource constraints and overfitting issues by incorporating data augmentation and lightweight learning, enhancing model performance and adaptability across diverse scenarios.

CN120316474AInactive Publication Date: 2025-07-15BANK OF CHANGSHA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510389635.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing financial timing data prediction models are complex in resource-constrained environments, difficult to process multivariate data, and are sensitive to small data volumes, are prone to overfitting, lack of adaptability and generalization capabilities, especially in the face of non-stationary time series or cross-scale analysis.

Method used

The data enhancement module is used to process the input data, build a core network of feature core and time core, combine the light element learning training framework and temperature parameter adjustment, and optimize the model through gating and jump connection, reducing resource consumption and improving model generalization capabilities.

Benefits of technology

The convergence speed of the model is accelerated, the influence of noise and outliers is reduced, the generalization performance of the model in small data scenarios is improved, the computing resource requirements is reduced, and the learning ability of long-term series is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316474A_ABST
    Figure CN120316474A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a financial time series data prediction method and device and a medium, and the method comprises the steps: obtaining financial data as input data, and employing a data enhancement module to enhance the input data to obtain time series data; a core network is constructed, the core network comprises a feature core and a time core, the feature core is used for learning the relation between the time sequence data and the time sequence, and the time core is used for learning the relation between the time sequence before and after; training the core network based on the financial data to obtain an initial model; parameters of the initial model are adjusted based on the input data size and temperature parameters, a financial time series data prediction model is obtained, financial data to be predicted are input into the financial time series data prediction model, and financial time series data prediction is achieved. According to the method, main features are amplified according to the probability of feature values through gating and jump connection, unimportant features are reduced, and the information extraction capability of the model is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of time series data prediction, and specifically relates to a financial time series data prediction method, device, and medium. Background Art

[0002] The time series data in the financial field is complex, uncertain, and highly volatile, making the prediction of financial time series data extremely challenging. In recent years, with the progress of large language models (LLMs) and other basic models, the use of these models in time series and spatio-temporal data mining has been increasing. Especially in the financial field, deep learning technologies such as long short-term memory networks (LSTMs), convolutional neural networks (CNNs), and Transformers have been widely applied to time series prediction, demonstrating the ability to handle non-linear data and long-term dependencies. However, these models usually require a large number of parameters and computing resources, limiting their application in resource-constrained environments.

[0003] When dealing with long time series data, time series models need to consider the stability and trend changes of the data, which increases the complexity of the models. As the amount of data increases, the models require more computing resources to process this data, including memory and processor capabilities. Without sufficient hardware support, the training and prediction processes of the models may become very slow or even infeasible. To reduce the computational burden, some lightweight time series models have been designed to handle public data sets. Although multivariate time series data is very common in the real world, many existing large time series models cannot effectively handle this multivariate data. All of the above make large time series models perform poorly when facing new and unseen data because they lack sufficient generalization ability, limiting their application in diverse and complex environments.

[0004] In addition, in most time series problems in the financial field, due to time constraints, the amount of data is small. And because of the small amount of data, it is very easy to have the problem of overfitting when using large models for training, resulting in many time series models being very sensitive to outliers and noise in the data, and the model performance deteriorates. And time series problems usually require time series models to be trained on specific time scales, which limits their adaptability to data on different time scales and further limits their generalization performance. When facing non-stationary time series or scenarios that require cross-scale analysis, the models may need to be redesigned and retrained.

[0005] In summary, there is an urgent need for a financial time series data prediction method, device, and medium to solve the problems in the prior art. Summary of the Invention

[0006] The object of the present invention is to provide a financial time series data prediction method, device, and medium, and the specific technical solutions are as follows:

[0007] A financial time series data prediction method, comprising the following steps:

[0008] S1: Data preprocessing, obtaining financial data as input data, and using a data augmentation module to augment the input data to obtain time series data;

[0009] S2: Constructing a core network, the core network including a feature core and a time core; the feature core is used to learn the relationship between time series data and time series, and obtain time series features; the time core is used to learn the relationship before and after time series, and obtain a result output;

[0010] S3: Training the core network based on financial data to obtain an initial model;

[0011] S4: Adjusting the parameters of the initial model based on the size of the input data volume and temperature parameters to obtain a financial time series data prediction model, and inputting the financial data to be predicted into the financial time series data prediction model to achieve financial time series data prediction.

[0012] Optionally, in S1, using a data augmentation module to augment the input data, including:

[0013] Random sampling: According to the input feature labels, perform stratified random sampling of time series from financial data according to the feature labels, and select k time series segments, where k is any integer between 1 and a certain maximum value K, and the length of the data slice is 5%-20% of the financial data;

[0014] Scaling: Scaling the adopted financial data to compare and combine different time series segments on the same scale;

[0015] Convex combination: The scaled time series are mixed together by convex combination to generate new augmented samples.

[0016] Optionally, in S1, the weights of the convex combination are sampled from a symmetric Dirichlet distribution, and the expression of the convex combination process is as follows:

[0017]

[0018] Among them, represents the augmented time series, is the i-th scaled time series, ω i is the weight sampled from the Dirichlet distribution.

[0019] Optionally, in S2, the feature core includes a fully connected layer, an activation function, a regularization module, and a gating function. The time series data is input into the feature core and sequentially passes through a fully connected layer, an activation function, a regularization module, and another fully connected layer, and finally outputs through the gating function to obtain the time series feature.

[0020] Optionally, in S2, the time core includes a fully connected layer, an activation function, a regularization module, and a gated attention mechanism. The time series feature is input into the time core and sequentially passes through a fully connected layer, an activation function, a regularization module, and another fully connected layer, and finally outputs through the gated attention mechanism and the activation function to obtain the result output.

[0021] Optionally, in S2, the gated attention mechanism adopts GLU gating, and the expression is as follows:

[0022] GLU(x) = x × σ(g(x));

[0023] Among them, GLU(x) represents the output of the GLU gating, x represents the input, σ represents the sigmoid function, and g(x) represents the output of the input x passing through the multi-layer perceptron.

[0024] Optionally, in S3, the core network is trained based on financial data and the light meta-learning training framework, and the process is as follows:

[0025] S3.1: Put in different types of financial data and randomly extract repeatedly to obtain training data;

[0026] S3.2: Randomly initialize the parameters of the core network;

[0027] S3.3: Randomly sample the training data to form a data batch.

[0028] S3.4: Start gradient update, and use each training data in the data batch to update the parameters of the core network respectively;

[0029] S3.5: Use the support set in a certain training data in the data batch to calculate the gradient of each parameter;

[0030] S3.6: End the gradient update;

[0031] S3.7: According to the parameters obtained from the previous gradient update, calculate the next gradient update through gradient by gradient; the gradient calculated during the next gradient update acts on the core network through stochastic gradient descent;

[0032] S3.8: Repeat steps S3.3 to S3.7, and when the parameters of the core network meet the stop condition, the training of the core network is completed.

[0033] Optionally, in S4, adjust the parameters of the initial model based on the size of the input data volume and the temperature parameter, and the expression is as follows:

[0034]

[0035] Where P(y i ) represents the output probability distribution, z i represents the score calculated by the model, and T represents the temperature parameter.

[0036] In addition, the present invention further includes a computer device, including a memory and a processor;

[0037] The memory is used to store a computer program that can run on the processor;

[0038] The processor is used to implement the steps of the financial time series data prediction method as described above when executing the computer program.

[0039] In addition, the present invention further includes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the financial time series data prediction method as described above are implemented.

[0040] Applying the technical solution of the present invention has the following beneficial effects:

[0041] In the preprocessing of the method of the present invention, data from different parts is added. The gradient vanishing problem is alleviated through the core network, the network convergence speed is accelerated, and the influence of noise and outliers is reduced. In addition, by setting two gates, the method of the present invention can learn long-term memory, and the core network structure is composed of fully connected layers, which improves the operation speed and reduces the resource consumption. The method of the present invention further controls the parameter adjustment ratio during model fine-tuning, reducing the requirement for computing resources when applied to new scenarios. In the core network of the method of the present invention, the gates and skip connections amplify the main features and shrink the unimportant features according to the probability of the eigenvalue, scale the initial input features through the weight matrix, enhance the information extraction ability of the neural network, and ensure that the model can learn long-term time series data.

[0042] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0044] Figure 1 is the flowchart of the steps of the financial time series data prediction method in the preferred embodiment of the present invention;

[0045] Figure 2 is the network schematic diagram of the feature core in the preferred embodiment of the present invention;

[0046] Figure 3 is the network schematic diagram of the time core in the preferred embodiment of the present invention. Detailed implementation manners

[0047] In order to enable those skilled in the art of this technology to better understand the solution of the present invention, the following will further elaborate on the present invention in conjunction with the drawings and specific implementation manners. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0048] As Figure 1 shown, this embodiment provides a financial time series data prediction method, including the following steps:

[0049] S1: Data preprocessing, obtaining financial data as the input data, and using a data enhancement module to enhance the input data to obtain time series data;

[0050] S2: Constructing a core network, the core network includes a feature core and a time core; the feature core is used to learn the relationship between time series data and time series, and obtain time series features; the time core is used to learn the relationship before and after time series, and obtain the result output;

[0051] S3: Training the core network based on financial data to obtain an initial model;

[0052] S4: Adjusting the parameters of the initial model based on the size of the input data volume and the temperature parameter to obtain a financial time series data prediction model, and inputting the financial data to be predicted into the financial time series data prediction model to achieve financial time series data prediction.

[0053] Optionally, in S1, using a data enhancement module to enhance the input data includes:

[0054] Random sampling: According to the input feature tags, perform stratified random sampling of time series from financial data according to the feature tags, and select k time series segments, where k is any integer between 1 and a certain maximum value K, and the length of this data slice is 5%-20% of the financial data;

[0055] Scaling: Scale the adopted financial data to compare and combine different time series segments on the same scale;

[0056] Convex combination: The scaled time series are mixed together through convex combination to generate new enhanced samples.

[0057] Optionally, in S1, the weights of the convex combination are sampled from a symmetric Dirichlet distribution, that is, in this embodiment, the importance of each time series segment in the final combination is determined by random weights obtained from the Dirichlet distribution. Specifically, the expression of the convex combination process is as follows:

[0058]

[0059] where, represents the enhanced time series, is the i-th scaled time series, ω i is the weight sampled from the Dirichlet distribution.

[0060] In this embodiment, the main role of data preprocessing is to clean and process the input data and transform it into information that the model can utilize. Compared with conventional preprocessing, this embodiment adds a data augmentation module in the current operation. Since the amount of time series-related data is usually small, it is easy to have the problem of overfitting when the large model is learning due to insufficient data volume. Therefore, this embodiment designs a data augmentation module to improve the generalization performance of the model, so that the large model can be well applied in scenarios with small data volume. It should be noted that for the data augmentation module in this embodiment, users can select the corresponding augmentation type and set the corresponding parameters according to actual needs for data augmentation. The specific augmentation types include random sampling, scaling, and convex combination, etc.

[0061] Such as Figure 2As shown, in S2, the feature core includes a fully connected layer, an activation function, a regularization module, and a gating function. The time series data is input into the feature core and sequentially passes through a fully connected layer, an activation function, a regularization module, and another fully connected layer, and finally outputs through the gating function to obtain time series features. The feature core first performs normalization and transposition operations on the data to facilitate the model's learning of the relationships between multivariate features, and then enters network training. The feature core in this embodiment consists of two fully connected layers, a RELU activation function, and a softmax gate. The RELU activation function is widely used in deep learning due to its high computational efficiency and ability to alleviate the vanishing gradient problem. At the same time, it also helps to accelerate the convergence of the neural network, and the sparsity of its output also helps to improve the generalization ability of the model. Therefore, relu is used as the activation function in the first layer of the feature core. The second layer of the feature core uses softmax as the gate, and the expression is as follows:

[0062]

[0063] where, e s is the exponential score matrix s, and the denominator belongs to the normalization operation to prevent the value from being too large after exponentiation.

[0064] The probability distribution obtained after passing through the softmax function can be regarded as the probability of each hidden unit being selected. These probabilities can be used as gating coefficients to determine which hidden unit outputs will be combined into the final output. In this embodiment, since the output value of the Softmax function is always between 0 and 1, the problem of gradient explosion or disappearance is avoided, and each unit output will have a non-zero probability value, which means that all units will be activated, that is, a non-sparse gating function is realized. Softmax gating is a form of conditional computation designed to increase the capacity and computational efficiency of the model.

[0065] Furthermore, a first skip connection is also provided in the feature core. The first skip connection is used to randomly select a part of the data that does not pass through the time core and directly combines it with the output result as the final output.

[0066] As Figure 3 shown, in S2, the time core includes a fully connected layer, an activation function, a regularization module, and a gated attention mechanism. The time series features are input into the time core and sequentially pass through a fully connected layer, an activation function, a regularization module, and another fully connected layer, and finally output through the gated attention mechanism and an activation function to obtain the result output.

[0067] Furthermore, a second skip connection is also provided in the time core. The second skip connection is used to randomly select a part of the data, which is directly transmitted to the next core without passing through the feature core, and combined with the output result of the core layer through addition as the input of the next core.

[0068] Optionally, in S2, the gated attention mechanism of the time core adopts GLU gating, and the expression is as follows:

[0069] GLU(x) = x × σ(g(x));

[0070] Among them, GLU(x) represents the output of the GLU gating, x represents the input, σ represents the sigmoid function, and g(x) represents the output of the input x passing through a multi-layer perceptron.

[0071] Optionally, in S3, the core network is trained based on financial data and the light meta-learning training framework, and the process is as follows:

[0072] S3.1: Put in different types of financial data, randomly extract repeatedly to obtain training data;

[0073] S3.2: Randomly initialize the parameters of the core network;

[0074] S3.3: Randomly sample the training data to form a data batch.

[0075] S3.4: Start gradient update, and use each training data in the data batch to update the parameters of the core network respectively;

[0076] S3.5: Use the support set in a certain training data in the data batch to calculate the gradient of each parameter;

[0077] S3.6: End the gradient update;

[0078] S3.7: According to the parameters obtained from the previous gradient update, calculate the next gradient update through gradient by gradient; the gradient calculated during the next gradient update acts on the core network through stochastic gradient descent;

[0079] S3.8: Repeat steps S3.3 to S3.7, and the training of the core network is completed when the parameters of the core network meet the stop condition.

[0080] It should be noted that the light meta-learning training framework involves sampling tasks from the task distribution, and these tasks are used to train the model. Each task can be regarded as a small dataset, including a support set and a query set. The light meta-learning training framework also includes an inner loop and an outer loop.

[0081] Inner Loop: For each task, use a small number of gradient descent steps to update the model parameters to obtain task-specific parameters. This process is carried out on the support set of each task, aiming to enable the model to quickly adapt to the current task.

[0082] Outer Loop: In the outer loop, calculate the task loss on the parameters updated in the inner loop and perform backpropagation with respect to the original parameters to update the global model parameters. This process is actually a second-order optimization problem because it is necessary to calculate the "gradient of the gradient".

[0083] Parameter Update: The parameter update process involves two learning rates: one is the learning rate for the inner loop update step, and the other is the learning rate for the outer loop update step. The learning rate of the outer loop is used to update the initial parameters of the model to make it more sensitive to new tasks.

[0084] Optionally, in S4, adjust the parameters of the initial model based on the size of the input data volume and the temperature parameter, and the expression is as follows:

[0085]

[0086] where P(y i ) represents the output probability distribution, z i represents the score calculated by the model, and T represents the temperature parameter.

[0087] When the temperature parameter T > 1, it makes the probability distribution output by softmax more uniform, increasing the randomness and innovation of the model output. This means that words with lower scores also have a reasonable probability of being selected, making the model more diverse and creative when generating text.

[0088] When the temperature parameter T < 1, it makes the probability distribution output by softmax sharper, that is, the probability of high-score words will be higher, while the probability of low-score words will be lower. This causes the model to be more inclined to select those most likely words, generating more conservative and predictable text.

[0089] Furthermore, in this embodiment, when fine-tuning, adjust the parameters of the model's core network and other parts according to the size of the input data volume and the temperature control parameter, so as to improve the running speed of the model while ensuring the performance of the model on other datasets.

[0090] Specifically, the temperature control parameter is controlled between 0 and 2. According to this parameter, the parameters of the corresponding proportion of the model except the backbone are directly adjusted. For example, when the temperature control parameter is 1.4, 70% of the parameters except the core network will be adjusted during fine-tuning, and at the same time, the probability distribution of the output will be affected. The parameters of the backbone part are 20% of the adjustable parameters. This part is comprehensively adjusted by the amount of input data and the temperature control parameter. When the amount of data is small, the temperature control parameter will be reduced, and the number of overall fine-tuning parameters will also be reduced. This selection of adjusting parameters can effectively reduce the problem of overfitting in model training when the amount of data is small. At the same time, since only some parameters are fine-tuned, the model training speed is also enhanced.

[0091] The fine-tuning mechanism in this embodiment optimizes the performance of the model by adjusting the temperature parameter, achieving an ideal balance between randomness and determinacy, enabling the model to produce diverse output results while maintaining accuracy.

[0092] In addition, this embodiment also provides a computer device, including a memory and a processor;

[0093] The memory is used to store a computer program that can run on the processor;

[0094] The processor is used to implement the steps of the above financial time series data prediction method when executing the computer program.

[0095] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and this instruction segment is used to describe the execution process of the computer program in the computer device.

[0096] The computer device can be a computing device such as a mobile phone, a desktop computer, a notebook, a handheld computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. For example, the computer device may further include input / output devices, network access devices, a bus, etc.

[0097] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines.

[0098] The memory can be used to store the computer program and / or modules. The processor realizes the computer program by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0099] Among them, if the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0100] In addition, an embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned financial time series data prediction method are implemented.

[0101] This embodiment provides a financial time series data prediction method, which alleviates the problem of gradient disappearance through a core network, accelerates the network convergence speed, and reduces the influence of noise and outliers. In addition, the method of the present invention learns long-term memory by setting two gates, and the core network structure is composed of fully connected layers, which improves the operation speed and reduces the resource consumption. The method of the present invention further controls the parameter adjustment ratio during model fine-tuning, reducing the requirements for computing resources when applied to new scenarios. In the core network of the method of the present invention, the gates and skip connections amplify the main features and shrink the unimportant features according to the probability of the eigenvalue, scale the initial input features through the weight matrix, enhance the information extraction ability of the neural network, and ensure that the model can learn long-term time series data.

[0102] It should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0103] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A financial time series data prediction method, characterized in that, It includes the following steps: S1: Data preprocessing. Obtain financial data as the input data, and use a data augmentation module to augment the input data to obtain time series data. S2: Build a core network. The core network includes a feature core and a time core. The feature core is used to learn the relationship between time series data and time series, and obtain time series features. The time core is used to learn the relationship before and after time series, and obtain the result output. S3: Train the core network based on the financial data to obtain an initial model. S4: Adjust the parameters of the initial model based on the size of the input data volume and the temperature parameter to obtain a financial time series data prediction model. Input the financial data to be predicted into the financial time series data prediction model to achieve financial time series data prediction.

2. The financial time series data prediction method according to claim 1, wherein In S1, using the data augmentation module to augment the input data includes: Random sampling: According to the input feature labels, perform stratified random sampling of time series from the financial data according to the feature labels, and select k time series segments, where k is any integer between 1 and a certain maximum value K. The length of this data slice is 5%-20% of the financial data. Scaling: Scale the adopted financial data to compare and combine different time series segments on the same scale. Convex combination: The scaled time series are mixed together by convex combination to generate new augmented samples.

3. The financial time series data prediction method according to claim 2, wherein In S1, the weights of the convex combination are sampled from a symmetric Dirichlet distribution. The expression of the convex combination process is as follows: Among them, represents the enhanced time series, is the i-th scaled time series, ω i is the weight sampled from the Dirichlet distribution.

4. The financial time series data prediction method according to claim 3, wherein In S2, the feature core includes a fully connected layer, an activation function, a regularization module, and a gating function. The time series data is input into the feature core and passes through a fully connected layer, an activation function, a regularization module, and another fully connected layer in sequence, and finally outputs through the gating function to obtain time series features.

5. The financial time series data prediction method according to claim 4, wherein In S2, the time core includes a fully connected layer, an activation function, a regularization module, and a gated attention mechanism. The time series features are input into the time core and pass through a fully connected layer, an activation function, a regularization module, and another fully connected layer in sequence, and finally output through the gated attention mechanism and the activation function to obtain the result output.

6. The financial time series data prediction method according to claim 5, wherein In S2, the gated attention mechanism adopts GLU gating, and the expression is as follows: GLU(x) = x × σ(g(x)); Among them, GLU(x) represents the output of the GLU gating, x represents the input, σ represents the sigmoid function, and g(x) represents the output of the input x passing through a multi-layer perceptron.

7. The financial time series data prediction method according to claim 6, wherein In S3, train the core network based on the financial data and the light meta-learning training framework. The process is as follows: S3.1: Put in different types of financial data, randomly extract repeatedly to obtain training data. S3.2: Randomly initialize the parameters of the core network. S3.3: Randomly sample the training data to form a data batch. S3.4: Start gradient update, and use each training data in the data batch to update the parameters of the core network respectively. S3.5: Use the support set in a certain training data in the data batch to calculate the gradient of each parameter. S3.6: End the gradient update. S3.7: Based on the parameters obtained from the previous gradient update, calculate the next gradient update through gradient by gradient; the gradient calculated during the next gradient update acts on the core network through stochastic gradient descent. S3.8: Repeat steps S3.3 to S3.

7. When the parameters of the core network meet the stopping condition, the training of the core network is completed.

8. The financial time series data prediction method according to claim 7, wherein In S4, adjust the parameters of the initial model based on the size of the input data volume and the temperature parameter, and the expression is as follows: Among them, P(y i ) represents the output probability distribution, z i represents the score calculated by the model, and T represents the temperature parameter.

9. A computer device, characterized in that, It includes a memory and a processor; The memory is used to store computer programs that can run on the processor; The processor is used to implement the steps of the financial time series data prediction method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the financial time series data prediction method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Financial time series data multi-scale feature analysis and prediction method and system

    CN121961728A