A method, system and storage medium for automatically generating stock market news

By combining preset description rules and deep neural network technology, stock market news is generated, which solves the problem of time-consuming and subjective generation of news in the existing technology, and realizes automation and rapid generation of stock market news with news value.

CN116303997BActive Publication Date: 2025-07-18JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211616030.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-10
Publication Date
2025-07-18
Estimated Expiration
2043-05-10

AI Technical Summary

Technical Problem

The existing technology is difficult to generate stock market news with good news value, and it is time-consuming and subjective by manual writing.

Method used

Preset description rules are used to generate opening and closing description texts, and early trading, afternoon and late trading description texts are generated by combining deep neural network models. The increase, trend and range increase sequences are obtained as input through data preprocessing to reduce the computational complexity.

Benefits of technology

It realizes automation and quickly generates stock market news with news value, saves manpower and is not subjective, and is suitable for news with very short timeliness to seize the initiative.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303997B_ABST
    Figure CN116303997B_ABST
Patent Text Reader

Abstract

The present invention provides a method, a system and a storage medium for automatically generating stock market news. This method combines the characteristics of preset description rules and neural network technology, gives full play to the advantages of each technology, uses the preset description rules to generate relatively fixed expressions in stock market news, and uses deep neural network technology to generate trend descriptions in stock market news. Specifically, the preset description rules are used to generate the opening description text and the closing description text of the stock market, and a deep neural network model is used to generate the morning session description text, the afternoon session description text and the closing session description text of the stock market. Before using the neural network model, the present invention fully explores the hidden information of time series data through data preprocessing to obtain the gain sequence, the trend sequence, the interval gain sequence and the auxiliary description data. By using these data to replace the original time series data as the input of the model, the computational complexity of the model is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method, a system and a storage medium for automatically generating stock market news. Background Art

[0002] With the continuous development of artificial intelligence technology, artificial intelligence technology has begun to be widely applied in various industries. In the field of news, the use of machines to write news has attracted the attention of large technology companies and research institutions. Especially for short news describing relevant data and being broadcast frequently, it has become urgently needed to use machines to replace human beings to complete such cumbersome tasks. Stock market news has attracted the attention of a large number of stockholders. However, due to its large amount of data and various types, writing by hand is not only time-consuming and laborious but also has a certain degree of subjectivity.

[0003] Currently, the research on generating text based on time series data has just started. Most of the existing technologies generate text by using rules. However, due to the characteristics that time series data is all composed of numerical data, with a large amount of data and complex feature extraction, there are thousands of possible changes in the whole sequence, and it is difficult to cover all situations with the definition of rules. Therefore, the text generated by using rules has a relatively fixed expression, covers less information, and does not have good news value.

[0004] Based on this, it is necessary to propose a method, a system and a storage medium for automatically generating stock market news to solve the above technical problems. Summary of the Invention

[0005] In view of the above situation, the main purpose of the present invention is to propose a method, a system and a storage medium for automatically generating stock market news to solve the above technical problems.

[0006] An embodiment of the present invention provides a method for automatically generating stock market news, wherein the method includes the following steps:

[0007] Step 1, generate an opening description text and a closing description text according to a preset description rule;

[0008] The preset description rule corresponding to the opening description text is:

[0009] When it is judged that the opening price change range A is in the interval (-∞, -0.7%), it is a large low opening; when it is judged that the opening price change range A is in the interval [-0.7%, -0.25%), it is a low opening; when it is judged that the opening price change range A is in the interval [-0.25%, 0), it is a slightly low opening; when it is judged that the opening price change range A is equal to 0, it is a flat opening; when it is judged that the opening price change range A is in the interval (0, 0.25%), it is a slightly high opening; when it is judged that the opening price change range A is in the interval (0.25%, 0.7%), it is a high opening; when it is judged that the opening price change range A is in the interval (0.7%, +∞], it is a large high opening;

[0010] The preset description rules corresponding to the closing description text are:

[0011] The closing description text is expressed as "closing (up / down) F%, closing at G point", where F represents the closing increase / decrease description value, and G represents the closing price description value;

[0012] Step 2: Use the deep neural network model to generate early trading description text, afternoon trading description text, and late trading description text:

[0013] Auxiliary sequence generation:

[0014] Generate auxiliary sequences, including an increase sequence, a trend sequence, and an interval increase sequence, wherein the increase sequence is used to reflect the relative increase or decrease of the current price of the index price time series and the closing price of the previous day, the trend sequence is used to reflect the urgency and duration of the fluctuation process of the time series, and the interval increase sequence is used to reflect the relative trends in the three intervals of morning, afternoon and end of the trading day;

[0015] Auxiliary description sequence generation:

[0016] The auxiliary description sequence includes a description of the increase or decrease and a description of the maximum and minimum values;

[0017] Inputting the auxiliary sequence and the auxiliary description sequence into a deep neural network model to calculate and obtain an early trading description text, an afternoon trading description text, and an end-of-day trading description text;

[0018] Step 3: Merge the opening description text, closing description text, morning description text, afternoon description text and end-of-day description text to obtain complete news.

[0019] The present invention proposes a method for automatically generating stock market news. By combining the characteristics of preset description rules and deep neural network technology, it gives full play to the advantages of each technology. The preset description rules are used to generate relatively fixed expressions, while the deep neural network technology is used to generate descriptions of the stock market trend. Specifically, the preset description rules are used to generate the opening description text and the closing description text, and the deep neural network model is used to generate the morning session description text, the afternoon session description text, and the closing session description text. Before using the neural network model, the present invention fully explores the hidden information of time series data through data preprocessing to obtain the gain sequence, the trend sequence, the interval gain sequence, and the auxiliary description data. By using these data to replace the original time series data as the input of the model, the computational complexity of the model is greatly reduced. The present invention can automatically generate market descriptions of the three major indexes, such as the opening, morning session, afternoon session, closing session, and closing trend descriptions, providing new ideas for stock market news workers, helping to save manpower, and the text generation process is automatically completed by the machine, which is fast and not subjective, seizing the opportunity for this type of news with extremely short timeliness.

[0020] For the method for automatically generating stock market news, in step one, the calculation formula corresponding to the opening gain or loss A is expressed as:

[0021] A = (C - B) / B

[0022] where B represents the closing price of the previous day, and C represents the opening price;

[0023] The calculation formula corresponding to the closing gain or loss D is expressed as:

[0024] D = (E - B) / B

[0025] where E represents the closing price;

[0026] The calculation formula corresponding to the closing gain or loss description value F is expressed as:

[0027] F = round(100 * D, 2)

[0028] The calculation formula corresponding to the closing price description value G is expressed as:

[0029] G = round(E, 0)

[0030] where round(m, n) represents rounding m to n decimal places, and rounding to 0 decimal places means taking the integer.

[0031] For the method for automatically generating stock market news, in step two, the calculation formula corresponding to the gain sequence is:

[0032]

[0033] where Denotes the value of the increase rate sequence at time t, X t Denotes the value of the original time series at time t;

[0034] The corresponding calculation formula for the trend sequence is expressed as:

[0035]

[0036] Among them, Denotes the value of the trend sequence at time t, X t-1 Denotes the value of the original time series at time t - 1;

[0037] The corresponding calculation formula for the interval increase rate sequence is expressed as:

[0038]

[0039] Among them, Denotes the value of the interval increase rate sequence at time t, and p denotes the price at the first moment of the interval data.

[0040] The method for automatically generating stock market news, wherein, in the second step, the deep neural network model includes four bidirectional LSTM encoders, and the hidden vector obtained by each bidirectional LSTM encoder is expressed as:

[0041]

[0042]

[0043]

[0044] Among them, z′ t Denotes the hidden vector obtained by each bidirectional LSTM encoder at time t, Denotes the forward hidden activation vector corresponding to time t, Denotes the backward hidden activation vector corresponding to time t, x′ t Denotes the input sequence, Denotes the forward encoder parameter, Denotes the backward encoder parameter, and LSTM(·) represents the long short-term neural network operation.

[0045] The method for automatically generating stock market news, wherein, the hidden state vector of each bidirectional LSTM encoder is expressed as:

[0046]

[0047] Among them, Z t Denotes the hidden state vector of each bidirectional LSTM encoder;

[0048] The method further includes:

[0049] Encoding the increase rate sequence, trend sequence, interval increase rate sequence, and auxiliary description sequence using a bidirectional LSTM encoder to obtain corresponding hidden state vectors;

[0050] Using a linear model to calculate the sum of the hidden state vectors of the increase rate sequence, trend sequence, and interval increase rate sequence to obtain the hidden state vector at the corresponding moment of the time series data, and the corresponding calculation formula is:

[0051]

[0052] where d t represents the hidden state vector at the t-th moment of the time series data, represents the hidden state vector of the increase rate sequence at the t-th moment, represents the hidden state vector of the trend sequence at the t-th moment, represents the hidden state vector of the interval increase rate sequence at the t-th moment, and Linear(·) represents a linear function operation;

[0053] Concatenating the hidden state vector of the time series data and the hidden state vector of the auxiliary description sequence to obtain the overall hidden activation vector sequence of the neural network model, and the corresponding formula is:

[0054]

[0055] where h represents the overall hidden activation vector sequence of the deep neural network model, represents the hidden state vector at the t-th moment corresponding to the auxiliary description sequence.

[0056] The automatic stock market news generation method, wherein, in the second step, the deep neural network model further includes a decoder, and in the decoding stage, the state representation of the decoder at the k-th step is:

[0057]

[0058] where S k represents the state corresponding to the k-th time step of the decoder, S k-1 represents the state corresponding to the (k - 1)-th time step of the decoder, γ D represents the decoder parameter, y′ k represents the embedded vector representation of the k-th time step in the auxiliary description sequence, represents the context vector of the sequence data during decoding at the (k - 1)-th time step;

[0059] The initial state s0 of the decoder is composed of the final state Z of the encoder L and the initial token y′0 and is obtained by calculating the embedding vector; among them, the final state z of the encoder L The calculation formula is expressed as:

[0060]

[0061] Among them, represents the final state of the increase rate sequence encoder, represents the final state of the trend sequence encoder, represents the final state of the interval increase rate sequence encoder, represents the final state of the auxiliary description sequence encoder, represents a vector of all zeros.

[0062] In the method for automatically generating stock market news, an attention mechanism is introduced during the decoding process to enable the deep neural network model to focus on the corresponding sequence data when decoding and generating descriptive words. The corresponding attention weights are expressed as:

[0063]

[0064]

[0065] Among them, represents a scalar, v T represents the transpose operation on the vector v, tanh(·) represents the hyperbolic tangent function operation, and the matrix W h , the matrix W s and the vector v, the vector b attn are all learnable parameters, h t represents the hidden vector of the overall hidden activation vector sequence h of the deep neural network model at time t, represents the attention weight corresponding to the hidden vector of the overall hidden activation vector sequence h of the deep neural network model at time t during decoding at the k-th time step. L represents the length of the overall hidden activation vector sequence h of the deep neural network model;

[0066] The context vector of the sequence data during decoding is calculated based on the attention weights. The corresponding formula is:

[0067]

[0068] Among them, represents the context vector of the sequence data during decoding at the k-th time step;

[0069] Connect the context vector of the sequence data during decoding with the state S corresponding to the k-th time step of the decoder k and generate the vocabulary distribution P vocab through two linear layers. The corresponding formula is expressed as:

[0070]

[0071] Among them, P vocab represents the probability distribution of all words in the vocabulary, soft max(·) represents the normalization process, and V′, V, ε, and ε′ all represent learnable parameters.

[0072] The method for automatically generating stock market news, wherein the method further includes:

[0073] Calculate the word with the highest probability according to the vocabulary distribution P vocab The corresponding calculation formula is expressed as:

[0074]

[0075] Among them, y k represents the word with the highest probability, arg max represents the operation of taking the maximum value, P vocab (w) represents the probability of generating the word w, and V vocab represents the vocabulary.

[0076] The present invention also proposes a system for automatically generating stock market news, wherein the above-mentioned method for automatically generating stock market news is applied, and the system includes:

[0077] The first generation module is used to generate the opening description text and the closing description text according to the preset description rules;

[0078] The preset description rule corresponding to the opening description text is:

[0079] When it is judged that the opening gain or loss A is in the interval (-∞, -0.7%), it is a large gap down; when it is judged that the opening gain or loss A is in the interval [-0.7%, -0.25%), it is a gap down; when it is judged that the opening gain or loss A is in the interval [-0.25%, 0), it is a slightly gap down; when it is judged that the opening gain or loss A is equal to 0, it is an even opening; when it is judged that the opening gain or loss A is in the interval (0, 0.25%], it is a slightly gap up; when it is judged that the opening gain or loss A is in the interval (0.25%, 0.7%], it is a gap up; when it is judged that the opening gain or loss A is in the interval (0.7%, +∞], it is a large gap up;

[0080] The preset description rule corresponding to the closing description text is:

[0081] The closing description text is expressed as "Closing (up / down) F%, closing at G points", where F represents the closing gain or loss description value and G represents the closing price description value;

[0082] A second generation module for generating morning session description text, afternoon session description text, and end-of-day session description text using a deep neural network model:

[0083] Auxiliary sequence generation:

[0084] Generate an auxiliary sequence, which includes a gain sequence, a trend sequence, and an interval gain sequence. The gain sequence is used to reflect the relative increase or decrease of the current price in the index price time series compared to the closing price of the previous day. The trend sequence is used to reflect the urgency and duration of the fluctuation process of the time series. The interval gain sequence is used to reflect the relative trends in the three intervals of the morning session, afternoon session, and end-of-day session;

[0085] Auxiliary description sequence generation:

[0086] The auxiliary description sequence includes gain / loss description and maximum / minimum value description;

[0087] Input the auxiliary sequence and the auxiliary description sequence into the deep neural network model for calculation to obtain the morning session description text, afternoon session description text, and end-of-day session description text;

[0088] A text combination module for combining the opening description text, closing description text, morning session description text, afternoon session description text, and end-of-day session description text to form a complete news item.

[0089] The present invention also proposes a storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the above-described method for automatically generating stock market news is implemented.

[0090] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the embodiments of the present invention. Description of the Drawings

[0091] Figure 1 Is a flowchart of a method for automatically generating stock market news proposed by the present invention;

[0092] Figure 2 Is a schematic diagram of the principle of a method for automatically generating stock market news proposed by the present invention;

[0093] Figure 3 Is a schematic diagram of the structural principle of the deep neural network text generation model in the present invention;

[0094] Figure 4 Is a schematic diagram of the structural principle of the text fusion model in the present invention;

[0095] Figure 5 An example diagram of generating the maximum / minimum value description in the present invention;

[0096] Figure 6 The structural schematic diagram of a stock market news automatic generation system proposed by the present invention. Detailed implementation manners

[0097] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0098] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed to represent some manners of implementing the principles of the embodiments of the present invention. However, it should be understood that the scope of the embodiments of the present invention is not limited thereto. On the contrary, the embodiments of the present invention include all variations, modifications and equivalents falling within the spirit and connotation of the appended claims.

[0099] Please refer to Figure 1 and Figure 2 , the present invention proposes a method for automatically generating stock market news, wherein the method includes the following steps:

[0100] Step 1: Generate an opening description text and a closing description text according to a preset description rule.

[0101] By simply counting the description texts, it is found that the opening description texts are divided into seven situations: significantly lower opening, lower opening, slightly lower opening, flat opening, slightly higher opening, higher opening, and significantly higher opening. In the present invention, the preset description rule corresponding to the opening description text is:

[0102] When it is judged that the opening rise and fall A is in the interval (-∞, -0.7%), it is a significantly lower opening; when it is judged that the opening rise and fall A is in the interval [-0.7%, -0.25%), it is a lower opening; when it is judged that the opening rise and fall A is in the interval [-0.25%, 0), it is a slightly lower opening; when it is judged that the opening rise and fall A is equal to 0, it is a flat opening; when it is judged that the opening rise and fall A is in the interval (0, 0.25%], it is a slightly higher opening; when it is judged that the opening rise and fall A is in the interval (0.25%, 0.7%], it is a higher opening; when it is judged that the opening rise and fall A is in the interval (0.7%, +∞], it is a significantly higher opening;

[0103] The preset description rule corresponding to the closing description text is:

[0104] The closing description text is expressed as "Closing (up / down) F%, closing at G points", where F represents the closing rise and fall description value and G represents the closing price description value.

[0105] In step 1, the calculation formula corresponding to the opening price increase or decrease A is expressed as:

[0106] A=(CB) / B

[0107] Among them, B represents the closing price of the previous day, and C represents the opening price;

[0108] The calculation formula corresponding to the closing price increase or decrease D is expressed as:

[0109] D=(EB) / B

[0110] Among them, E represents the closing price;

[0111] The calculation formula corresponding to the closing price change description value F is expressed as:

[0112] F = round(100*D, 2)

[0113] The calculation formula corresponding to the closing price description value G is expressed as:

[0114] G = round(E, 0)

[0115] Among them, round(m, n) means retaining n decimal places for m, and retaining 0 decimal places means rounding.

[0116] As a supplementary explanation, (up / down) is obtained through a simple judgement, E>B corresponds to "up", otherwise it corresponds to down.

[0117] Step 2: Use the deep neural network model to generate early trading description text, afternoon trading description text, and late trading description text.

[0118] Since the description generation of early, afternoon and late trading data involves a lot of data reasoning and calculation, the features are relatively abstract and hidden deep, and feature extraction is difficult. In order to reduce the calculation of the deep neural network model, the present invention performs data preprocessing on the data before using the neural network model based on empirical knowledge. Among them, data preprocessing mainly includes two parts: auxiliary sequence generation and auxiliary description generation.

[0119] (1) Auxiliary sequence generation:

[0120] Generate auxiliary sequences, which include an increase sequence, a trend sequence and an interval increase sequence, wherein the increase sequence is used to reflect the relative increase or decrease between the current price of the index price time series and the closing price of the previous day, the trend sequence is used to reflect the urgency and duration of the fluctuation process of the time series, and the interval increase sequence is used to reflect the relative trends in the three intervals of morning, afternoon and end of the day.

[0121] Among them, the calculation formula corresponding to the increase sequence is expressed as:

[0122]

[0123] Among them, P t chg represents the value of the increase rate sequence at time t, and X t represents the value of the original time series at time t;

[0124] The corresponding calculation formula for the trend sequence is expressed as:

[0125]

[0126] Among them, P t trend represents the value of the trend sequence at time t, and X t-1 represents the value of the original time series at time t - 1;

[0127] The corresponding calculation formula for the interval increase rate sequence is expressed as:

[0128]

[0129] Among them, P t schg represents the value of the interval increase rate sequence at time t, and p represents the price at the first moment of the interval data.

[0130] (2) Generation of auxiliary description sequence:

[0131] The auxiliary description sequence includes the increase / decrease description and the maximum / minimum value description;

[0132] Input the auxiliary sequence and the auxiliary description sequence into the deep neural network model for calculation to obtain the morning session description text, the afternoon session description text, and the closing session description text.

[0133] Specifically, perform the following statistical analysis on the time series data of the three time periods of the morning session, the afternoon session, and the closing session:

[0134] 1) Increase / decrease analysis: Obtain the maximum increase rate and the minimum increase rate of the time period data, and generate the corresponding auxiliary description text according to the maximum increase rate and the minimum increase rate. The generation rule of the increase / decrease description text is shown in Table 1.

[0135] Table 1 Increase rate description rule

[0136]

[0137] As can be seen from Table 1(a) and Table 1(b): Generate the corresponding description according to the interval where the maximum and minimum increase rates are located. When it is in the interval [-0.8%, 0.8%], it is an empty string.

[0138] 2) Maximum and minimum value analysis is performed to obtain the intervals where the maximum and minimum values of the overall time series data are located, and the corresponding maximum and minimum value descriptions of "maximum value" or "minimum value" are obtained for the corresponding intervals. An example of generating the maximum and minimum value descriptions is as follows Figure 5 shown: The minimum value is in the early trading session interval, and the maximum value is in the late trading session interval. Therefore, the maximum and minimum value descriptions for the early trading session are: "minimum value", and for the late trading session are: "maximum value", while for the afternoon session it is an empty string.

[0139] The increase and decrease description is concatenated with the maximum and minimum value description text to obtain the auxiliary description text. Taking Figure 5 as an example, the minimum value of the early trading session interval sequence is the minimum value of the overall sequence and is greater than 0, so no description is required. The maximum value of the early trading session interval is in the interval (0.8%, 1%], and the corresponding description is: "rose nearly 1%". Similarly, the increase and decrease descriptions for the afternoon and late trading sessions are: "rose more than 1%". Therefore, the auxiliary description texts for the three intervals are shown in Table 2.

[0140] Table 2 Examples of auxiliary description texts for each interval

[0141] Segment interval Auxiliary description text Morning session "Rose nearly 1%, minimum value" Afternoon session "Rose more than 1%" Closing session "Rose more than 1%, maximum value"

[0142] In the present invention, the deep neural network model includes four bidirectional LSTM encoders (as Figure 3 shown), and the hidden vector obtained by each bidirectional LSTM encoder is represented as:

[0143]

[0144]

[0145]

[0146] where z′ t represents the hidden vector obtained by each bidirectional LSTM encoder at time t, represents the forward hidden activation vector corresponding to time t, represents the backward hidden activation vector corresponding to time t, x′ t represents the input sequence, represents the forward encoder parameter, represents the backward encoder parameter, and LSTM(·) represents the long short-term neural network operation.

[0147] To ensure that the model pays more attention to the forward sequence changes, in the present invention, the forward activation vector is given twice the weight, and the hidden state vector of each bidirectional LSTM encoder is represented as:

[0148]

[0149] where, Z tRepresents the hidden state vector of each bidirectional LSTM encoder.

[0150] In the present invention, further, a bidirectional LSTM encoder is used to encode the increase sequence, trend sequence, interval increase sequence, and auxiliary description sequence to obtain corresponding hidden state vectors;

[0151] The sum of the hidden state vectors of the increase sequence, trend sequence, and interval increase sequence is calculated using a linear model to obtain the hidden state vector at the corresponding moment of the time series data. The corresponding calculation formula is:

[0152]

[0153] where d t represents the hidden state vector at the t-th moment of the time series data, represents the hidden state vector of the increase sequence at the t-th moment, represents the hidden state vector of the trend sequence at the t-th moment, represents the hidden state vector of the interval increase sequence at the t-th moment, and Linear(·) represents a linear function operation;

[0154] The hidden state vector of the time series data and the hidden state vector of the auxiliary description sequence are concatenated to obtain the overall hidden activation vector sequence of the deep neural network model. The corresponding formula is expressed as:

[0155]

[0156] where h represents the overall hidden activation vector sequence of the deep neural network model, represents the hidden state vector of the auxiliary description sequence at the t-th moment.

[0157] Further, the deep neural network model further includes a decoder. In the decoding stage, the state representation of the decoder at the k-th step is:

[0158] s k = LSTM(s k-1 , Lznear(y′ k , V c k-1 )); γ D )

[0159] where S k represents the state corresponding to the k-th time step of the decoder, S k-1 represents the state corresponding to the (k - 1)-th time step of the decoder, γ D represents the decoder parameter, y′ k represents the embedded vector representation of the k-th time step in the auxiliary description sequence, Represents the context vector of the sequence data during decoding at the (k-1)-th time step.

[0160] In the present invention, the initial state s0 of the decoder is determined by the final state Z of the encoder L along with the initial token y′0 and the embedding vector of. Among them, the final state z of the encoder L is calculated by the formula:

[0161]

[0162] where represents the final state of the increase rate sequence encoder, represents the final state of the trend sequence encoder, represents the final state of the interval increase rate sequence encoder, represents the final state of the auxiliary description sequence encoder, is a vector of all zeros.

[0163] During the decoding process, an attention mechanism is introduced to enable the deep neural network model to focus on the corresponding sequence data when decoding and generating descriptive words. The corresponding attention weights are expressed as:

[0164]

[0165]

[0166] where represents a scalar, v T represents the transpose operation on the vector v, tanh(·) represents the hyperbolic tangent function operation, the matrix W h and the matrix W s along with the vector v and the vector b attn are all learnable parameters, h t represents the hidden vector of the overall hidden activation vector sequence h of the deep neural network model at time t, represents the attention weight corresponding to the hidden vector of the overall hidden activation vector sequence h of the neural network model at time t during decoding at the k-th time step, and L represents the length of the overall hidden activation vector sequence h of the neural network model.

[0167] Furthermore, based on the attention weights, the context vector of the sequence data during decoding is calculated, and the corresponding formula is:

[0168]

[0169] where represents the context vector of the sequence data during decoding at the k-th time step;

[0170] Further, the context vector of the sequence data during decoding is concatenated with the state S corresponding to the k-th time step of the decoder k and passed through two linear layers to generate the vocabulary distribution P vocab , and the corresponding formula is expressed as:

[0171]

[0172] where P vocab represents the probability distribution of all words in the vocabulary, soft max(·) represents the normalization process, and V′, V, ε, and ε′ all represent learnable parameters.

[0173] Further, the word with the highest probability is calculated according to the vocabulary distribution P vocab , and the corresponding calculation formula is expressed as:

[0174]

[0175] where y k represents the word with the highest probability, arg max represents the operation of taking the maximum value, P vocab (w) represents the probability of generating the word w, and V vocab represents the vocabulary.

[0176] As a supplementary note, during the training process, the loss at the k-th time step is the negative log-likelihood value of the target word at this time step , and the corresponding calculation formula is expressed as:

[0177]

[0178] where loss k represents the loss at the k-th time step, represents the probability of generating the target word ;

[0179] and the overall loss of the entire sequence of hidden activation vectors of the deep neural network model is:

[0180]

[0181] where loss represents the overall loss of the entire sequence of hidden activation vectors of the deep neural network model, and K represents the maximum value of the time step.

[0182] Step 3: Combine the opening description text, closing description text, early morning description text, afternoon description text, and late trading description text to form a complete news item.

[0183] In this step, the generated description texts are simply connected by strings in the order of opening, morning, afternoon, end and closing. The order is made to conform to the real description order, such as the description text: "Today, the Shanghai Composite Index opened sharply higher, fluctuated slightly higher in the morning, once rose by more than 1%, then fell back and weakened, fluctuated down and turned green, fluctuated downward in the afternoon, once fell by more than 1%, then rebounded upward after stabilizing at the bottom, and the decline narrowed. After a slight fluctuation and rise in the end, it fluctuated and ran near the flat line. It closed down 0.05% and closed at 3244 points." The five parts of the description text are clear and easy to distinguish, but when the trend performance of the morning, afternoon or end is consistent, it is often described in a unified way to avoid redundant descriptions, such as the description: "Today, the Shanghai Composite Index opened slightly higher, fluctuated downward and turned green in the morning, then rebounded and strengthened, fluctuated and rose throughout the day, and fluctuated and rushed high in the end. It closed up 2.47% and closed at 3096 points." The afternoon trend is included in the morning description.

[0184] In this regard, the present invention designs a text-to-text generation model architecture based on the LSTM sequence-to-sequence model (such as Figure 4 As shown in Figure 1), the generated five-part description text is used as input to output the final description text. This sequence-to-sequence model can not only combine description texts of multiple intervals, but also achieve a correction effect by reinterpreting the text that does not conform to the language rules generated in the sequence-to-text generation model to generate a more understandable description text.

[0185] The method for automatically generating stock market news proposed in the present invention combines the characteristics of preset description rules and deep neural network technology, gives full play to the advantages of their respective technologies, generates relatively fixed expressions by using preset description rules, and uses deep neural network technology to generate time series data trend description. Specifically, the opening description text and the closing description text are generated by using the preset description rules, and the early trading description text, the afternoon description text and the late trading description text are generated by using the deep neural network model. Before using the deep neural network model, the present invention fully explores the hidden information of the time series data through data preprocessing to obtain the increase sequence, trend sequence, interval increase sequence and auxiliary description data, and uses these data to replace the original time series data as the input of the model, which greatly reduces the calculation complexity of the model. The present invention can automatically generate the market description of the three major indexes, such as the opening, early trading, afternoon, late trading and closing trend description, which provides new ideas for stock market journalists, helps save manpower, and the text generation process is automatically completed by the machine, which is fast and not subjective, and seizes the opportunity for such news with extremely short timeliness.

[0186] See also Figure 5, the present invention also provides a stock market news automatic generation system, wherein the stock market news automatic generation method as described above is applied, and the system includes:

[0187] A first generation module, configured to generate an opening description text and a closing description text according to a preset description rule;

[0188] The preset description rule corresponding to the opening description text is:

[0189] When it is determined that the opening gain or loss A is in the interval (-∞, -0.7%), it is a large gap down; when it is determined that the opening gain or loss A is in the interval [-0.7%, -0.25%), it is a gap down; when it is determined that the opening gain or loss A is in the interval [-0.25%, 0), it is a slightly gap down; when it is determined that the opening gain or loss A is equal to 0, it is an even opening; when it is determined that the opening gain or loss A is in the interval (0, 0.25%], it is a slightly gap up; when it is determined that the opening gain or loss A is in the interval (0.25%, 0.7%], it is a gap up; when it is determined that the opening gain or loss A is in the interval (0.7%, +∞], it is a large gap up;

[0190] The preset description rule corresponding to the closing description text is:

[0191] The closing description text is expressed as "closing (up / down) F%, closing at G points", where F represents the closing gain or loss description value and G represents the closing price description value;

[0192] A second generation module, configured to generate a morning session description text, an afternoon session description text, and a late session description text using a deep neural network model:

[0193] Auxiliary sequence generation:

[0194] Generate an auxiliary sequence, where the auxiliary sequence includes a gain sequence, a trend sequence, and an interval gain sequence. The gain sequence is used to reflect the relative gain or loss of the current price of the index price time series and the closing price of the previous day. The trend sequence is used to reflect the urgency and duration of the fluctuation process of the time series. The interval gain sequence is used to reflect the relative trend in the three intervals of the morning session, the afternoon session, and the late session;

[0195] Auxiliary description sequence generation:

[0196] The auxiliary description sequence includes gain or loss descriptions and maximum and minimum value descriptions;

[0197] Input the auxiliary sequence and the auxiliary description sequence into the deep neural network model for calculation to obtain the morning session description text, the afternoon session description text, and the late session description text;

[0198] A text combination module for combining the opening description text, closing description text, morning session description text, afternoon session description text, and end-of-day session description text to obtain a complete news item.

[0199] The present invention also provides a storage medium having a computer program stored thereon, wherein when the program is executed by a processor, it implements the stock market news automatic generation method described above.

[0200] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0201] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0202] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the present invention's patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention's patent should be subject to the appended claims.

Claims

1. A method for automatically generating stock market news, characterized in that, The method comprises the following steps: Step 1: Generate stock market opening description text and closing description text according to preset description rules; The preset description rules corresponding to the opening description text are: When it is judged that the opening price change range A is in the interval (-∞, -0.7%), it is a large low opening; when it is judged that the opening price change range A is in the interval [-0.7%, -0.25%), it is a low opening; when it is judged that the opening price change range A is in the interval [-0.25%, 0), it is a slight low opening; When it is judged that the opening price change range A is equal to 0, it is a flat opening; when it is judged that the opening price change range A is in the range of (0, 0.25%], it is a slightly high opening; when it is judged that the opening price change range A is in the range of (0.25%, 0.7%], it is a high opening; when it is judged that the opening price change range A is in the range of (0.7%, +∞], it is a significantly high opening; The preset description rules corresponding to the closing description text are: The closing description text is expressed as "closing (up / down) F%, closing at G point", where F represents the closing increase / decrease description value, and G represents the closing price description value; Step 2: Use the deep neural network model to generate early trading description text, afternoon trading description text, and late trading description text: Auxiliary sequence generation: Generate auxiliary sequences, including an increase sequence, a trend sequence, and an interval increase sequence, wherein the increase sequence is used to reflect the relative increase or decrease of the current price of the index price time series and the closing price of the previous day, the trend sequence is used to reflect the urgency and duration of the fluctuation process of the time series, and the interval increase sequence is used to reflect the relative trends in the three intervals of morning, afternoon and end of the trading day; Auxiliary description sequence generation: The auxiliary description sequence includes a description of the increase or decrease and a description of the maximum and minimum values; Inputting the auxiliary sequence and the auxiliary description sequence into a deep neural network model to calculate and obtain an early trading description text, an afternoon trading description text, and an end-of-day trading description text; Step 3: Merge the opening description text, closing description text, morning description text, afternoon description text and end-of-day description text to obtain complete news.

2. The automatic stock market news generation method according to claim 1, characterized in that In step 1, the calculation formula corresponding to the opening price fluctuation A is expressed as: A=(CB) / B Among them, B represents the closing price of the previous day, and C represents the opening price; The calculation formula corresponding to the closing price increase or decrease D is expressed as: D=(EB) / B Among them, E represents the closing price; The calculation formula corresponding to the closing price change description value F is expressed as: F = round(100*D, 2) The calculation formula corresponding to the closing price description value G is expressed as: G = round(E, 0) Among them, round(m, n) means retaining n decimal places for m, and retaining 0 decimal places means rounding.

3. The automatic stock market news generation method according to claim 2, characterized in that, In step 2, the calculation formula corresponding to the increase sequence is expressed as: Among them, P t chg represents the value of the increase rate sequence at time t, and X t represents the value of the original time series at time t; The calculation formula corresponding to the trend sequence is expressed as: Among them, P t trend represents the value of the trend sequence at time t, and X t-1 represents the value of the original time series at time t - 1; The calculation formula corresponding to the interval increase sequence is expressed as: Among them, P t schg represents the value of the interval increase rate sequence at time t, and p represents the price at the first moment of the interval data.

4. The automatic stock market news generation method according to claim 3, wherein In step 2, the deep neural network model includes four bidirectional LSTM encoders, and the hidden vector obtained by each bidirectional LSTM encoder is expressed as: Among them, z′ t represents the hidden vector obtained by each bidirectional LSTM encoder at time t, represents the forward hidden activation vector corresponding to time t, represents the backward hidden activation vector corresponding to time t, x′ t represents the input sequence, represents the forward encoder parameter, represents the backward encoder parameter, and LSTM(·) represents the long short-term neural network operation.

5. The automatic stock market news generation method according to claim 4, wherein The hidden state vector of each bidirectional LSTM encoder is represented as: Among them, Z t represents the hidden state vector of each bidirectional LSTM encoder; The method further comprises: Encode the increase sequence, trend sequence, interval increase sequence, and auxiliary description sequence using a bidirectional LSTM encoder to obtain corresponding hidden state vectors; Use a linear model to calculate the sum of the hidden state vectors of the increase sequence, trend sequence, and interval increase sequence to obtain the hidden state vector at the corresponding moment of the time series data. The corresponding calculation formula is: Among them, d t represents the hidden state vector of the time series data corresponding to the t-th moment, represents the hidden state vector of the increase rate series at the t-th moment, represents the hidden state vector of the trend series at the t-th moment, represents the hidden state vector of the interval increase rate series at the t-th moment, and Linear(·) represents a linear function operation; Concatenate the hidden state vector of the time series data and the hidden state vector of the auxiliary description sequence to obtain the overall hidden activation vector sequence of the deep neural network model. The corresponding formula is: Among them, h represents the sequence of hidden activation vectors of the entire neural network model, represents the hidden state vector corresponding to the t-th moment of the auxiliary description sequence.

6. The automatic stock market news generation method according to claim 5, characterized in that In the second step, the deep neural network model further includes a decoder. In the decoding stage, the state representation of the decoder at the k-th step is: Among them, S k represents the state corresponding to the k-th time step of the decoder, and S k-1 represents the state corresponding to the (k - 1)-th time step of the decoder, and γ D represents the decoder parameters, and y′ k represents the embedded vector representation of the k-th time step in the auxiliary description sequence, represents the context vector of the sequence data when decoding at the (k - 1)-th time step; The initial state s0 of the decoder is calculated from the final state z of the encoder L and the initial token y′0 and the embedding vector of; where, the final state z of the encoder L The calculation formula of is expressed as: Among them, represents the final state of the increase rate sequence encoder, represents the final state of the trend sequence encoder, represents the final state of the interval increase rate sequence encoder, represents the final state of the auxiliary description sequence encoder, represents a zero vector.

7. A method for automatically generating stock market news according to claim 6, characterized in that, Introduce an attention mechanism during the decoding process to enable the deep neural network model to focus on the corresponding sequence data when decoding and generating descriptive words. The corresponding attention weights are: Among them, represents a scalar, v T represents the transpose operation on the vector v, tanh(·) represents the hyperbolic tangent function operation, and the matrix W h , the matrix W s , the vector v, and the vector b attn are all learnable parameters, h t represents the hidden vector of the hidden activation vector sequence h of the entire neural network model at time step t, represents the attention weight corresponding to the hidden vector of the hidden activation vector sequence h of the entire neural network model at time step t during decoding at the k-th time step, and L represents the length of the hidden activation vector sequence h of the entire neural network model; Calculate the context vector of the sequence data during decoding based on the attention weights. The corresponding formula is: Among them, represents the context vector of the sequence data at the k-th time step during decoding; The context vector of the sequence data during decoding is concatenated with the state S corresponding to the k-th time step of the decoder k and passed through two linear layers to generate the vocabulary distribution P vocab , and the corresponding formula is expressed as: Among them, P vocab represents the probability distribution of all words in the vocabulary, softmax(·) represents the normalization process, and V′, V, ε, and ε′ all represent learnable parameters.

8. A method for automatically generating stock market news according to claim 7, characterized in that The method further includes: According to the vocabulary distribution P vocab Calculate the word with the highest probability, and the corresponding calculation formula is expressed as: Among them, y k represents the word with the highest probability, argmax represents the operation of taking the maximum value, and P vocab (w) represents the probability of generating the word w, and V vocab represents the vocabulary.

9. An automatic stock market news generation system, characterized in that, Apply the stock market news automatic generation method according to any one of claims 1 to 8 above. The system includes: A first generation module for generating an opening description text and a closing description text according to a preset description rule; The preset description rule corresponding to the opening description text is: When it is determined that the opening gain or loss A is in the interval (-∞, -0.7%), it is a large opening decline; when it is determined that the opening gain or loss A is in the interval [-0.7%, -0.25%), it is an opening decline; when it is determined that the opening gain or loss A is in the interval [-0.25%, 0), it is a slight opening decline; when it is determined that the opening gain or loss A is equal to 0, it is an opening flat; when it is determined that the opening gain or loss A is in the interval (0, 0.25%], it is a slight opening increase; when it is determined that the opening gain or loss A is in the interval (0.25%, 0.7%], it is an opening increase; when it is determined that the opening gain or loss A is in the interval (0.7%, +∞], it is a large opening increase; The preset description rule corresponding to the closing description text is: The closing description text is expressed as "Closing (up / down) F%, closing at G points", where F represents the closing gain or loss description value and G represents the closing price description value; A second generation module for generating an early morning description text, an afternoon description text, and a late afternoon description text using the deep neural network model: Auxiliary sequence generation: Generate an auxiliary sequence, which includes an increase sequence, a trend sequence, and an interval increase sequence. The increase sequence is used to reflect the relative gain or loss of the current price of the index price time series and the closing price of the previous day. The trend sequence is used to reflect the urgency and duration of the fluctuation process of the time series. The interval increase sequence is used to reflect the relative trend in the three intervals of early morning, afternoon, and late afternoon; Auxiliary description sequence generation: The auxiliary description sequence includes gain or loss descriptions and maximum and minimum value descriptions; Input the auxiliary sequence and the auxiliary description sequence into the deep neural network model for calculation to obtain an early morning description text, an afternoon description text, and a late afternoon description text; A text combination module is used to combine the opening description text, closing description text, morning session description text, afternoon session description text, and end-of-day session description text to obtain a complete news item.

10. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements a method for automatically generating stock market news as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A multi-feature fusion Chinese news text abstract generation method based on a neural network

    CN109344391A

  • System, method and device for realizing automatic generation of stock market closing comment information, processor and storage medium thereof

    CN112347762A