Text classification processing method, device and equipment of news corpus and storage medium

CN115774781BActive Publication Date: 2026-09-29CHINA CONSTR BANK CORP SICHUAN BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211447898.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2026-09-29
Estimated Expiration
2042-11-18

AI Technical Summary

Benefits of technology

[0046]另一方面,本公开提供一种计算机程序产品,包括计算机程序,该计算机程序被处理器执行时实现任一项上述的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115774781B_ABST
    Figure CN115774781B_ABST
Patent Text Reader

Abstract

The present disclosure provides a news corpus text classification processing method, device and equipment and storage medium. The method comprises: performing word segmentation processing on the obtained news corpus to obtain segmented word groups; performing word embedding coding processing on the segmented word groups to obtain multi-dimensional tensor representation of the segmented word groups; determining the correlation between the target index tensor and the multi-dimensional tensor representation of the segmented word groups by using the pre-trained BERT-CNN based text category classification model, so as to determine the corresponding text category of the news corpus based on the correlation, wherein the target index tensor is used to represent the tensor corresponding to at least one technical index obtained by pre-labeling. To solve the problem of low accuracy of text category classification of news corpus, resulting in low accuracy of quantitative strategy involving news corpus, thereby improving the technical effect of improving the accuracy of quantitative strategy involving news corpus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer natural language processing technology, and more particularly to a method, apparatus, device, and storage medium for text classification processing of news corpora. Background Technology

[0002] In existing technologies, numerous factors are involved in quantification strategies. Among them, besides technical indicators, news corpora also have a significant impact on the strategy. Properly handling news corpora can effectively improve the accuracy of quantification decisions. Currently, the mainstream approach to processing news corpora is based on Long Short-Term Memory (LSTM) networks or BERT networks for supervised training in sentiment classification.

[0003] These two approaches have the following drawbacks: First, many news corpora contain both optimistic and pessimistic views on the market, often using techniques like transitions where the content following "but" is the key point. Long Short-Term Memory (LSTM) networks, due to their nature, will also extract the content before "but" during feature extraction, introducing irrelevant information. Second, most news corpora contain a great deal of textual description, much of which states objective facts rather than opinions. Traditional training methods use the entire text as input, resulting in low training efficiency. Third, quantitative strategies need to be adapted to market changes promptly. The sequential nature of LTM training leads to excessively long training times, potentially affecting the timeliness of strategy implementation. Finally, news may be misleading, with viewpoints contradicting actual market performance and lacking a positive correlation. Summary of the Invention

[0004] This disclosure provides a method, apparatus, device, and storage medium for text classification processing of news corpora, in order to solve the problem that the low accuracy of text category classification of news corpora leads to low accuracy of quantification strategies involving news corpora.

[0005] On the one hand, this disclosure provides a text classification and processing method for news corpora, the method including:

[0006] The acquired news corpus is segmented into words to obtain segmented word groups;

[0007] The above word segmentation array is subjected to word embedding encoding to obtain word segmentation groups represented by multidimensional tensors;

[0008] A pre-trained BERT-CNN-based text category classification model is used to determine the correlation between the target indicator tensor and the word segmentation group represented by the above multidimensional tensor, so as to determine the text category corresponding to the above news corpus based on the above correlation. The target indicator tensor is used to represent the tensor corresponding to at least one technical indicator obtained by pre-annotation.

[0009] Furthermore, the above method also includes:

[0010] Obtain the aforementioned news data from the news corpus, as well as the publication time and language of the aforementioned news data;

[0011] The above news corpus is filtered to obtain the filtered corpus;

[0012] Based on the publication time and language of the above news corpus, the filtered corpus is classified to obtain the classified corpus;

[0013] The above-mentioned word segmentation processing of the acquired news corpus to obtain word segments includes: using a word segmenter to segment the text in the above-classified corpus to obtain the above-mentioned word segments.

[0014] Furthermore, the aforementioned word segmentation group is the word segmentation array output by the word segmenter, represented by a two-dimensional tensor. The word segmentation array is then subjected to word embedding encoding to obtain a word segmentation group represented by a multi-dimensional tensor, including:

[0015] The word segmentation array represented by the above two-dimensional tensor is subjected to word embedding encoding to obtain the word segmentation group represented by the above multi-dimensional tensor. The above two-dimensional tensor includes: publication time period and word segmentation, and the above multi-dimensional tensor includes: publication time period, word segmentation and word tensor.

[0016] Furthermore, the pre-trained BERT-CNN-based text category classification model used above determines the correlation between the target index tensor and the word segmentation groups represented by the aforementioned multidimensional tensor, including:

[0017] Based on the objective function, the covariance of the target index tensor and the word segmentation group represented by the multidimensional tensor is calculated to obtain the covariance value. The objective function is used to improve the positive correlation between the word segmentation group represented by the multidimensional tensor and the target index tensor.

[0018] Using the BERT-CNN-based text category classification model, the correlation between the target index tensor and the word segmentation group represented by the multidimensional tensor is determined based on the covariance value.

[0019] Furthermore, the above method also includes:

[0020] The initial BERT model was trained using machine learning with multiple sets of sample data to obtain the trained BERT model. Each set of sample data included: general corpus and its corresponding text categories, news corpus and its corresponding text categories, and news corpus related to at least one technical indicator and its corresponding text categories.

[0021] Before the output layer of the trained BERT model, a convolutional neural network (CNN) layer is concatenated to obtain the text category classification model described above.

[0022] Furthermore, the above method also includes:

[0023] Obtain the text category classification models running on different server nodes, as well as the BERT model parameters of the text category classification models.

[0024] The BERT model parameters corresponding to the different server nodes are stored on a network disk, so as to obtain the mean-normalized model parameters based on the BERT model parameters corresponding to the different server nodes.

[0025] The mean-normalized model parameters are sent to the different server nodes using a message queue, so that the different server nodes can use the mean-normalized model parameters to synchronize the text category classification model running locally.

[0026] Furthermore, the aforementioned technical indicators include at least one of the following: market technical stochastic indicators, daily net trading volume indicators, and market trend indicators.

[0027] On the other hand, this disclosure provides a text classification and processing apparatus for news corpora, the apparatus comprising:

[0028] The word segmentation module is used to segment the acquired news corpus into words to obtain word groups.

[0029] The encoding processing module is used to perform word embedding encoding processing on the above word segmentation array to obtain word segmentation groups represented by multidimensional tensors;

[0030] The text classification module is used to determine the correlation between the target indicator tensor and the word segmentation group represented by the multidimensional tensor using a pre-trained BERT-CNN-based text category classification model, so as to determine the text category corresponding to the news corpus based on the correlation. The target indicator tensor is used to represent the tensor corresponding to at least one pre-annotated technical indicator.

[0031] Furthermore, the above-mentioned device also includes: a first acquisition module, used to acquire the news data in the news corpus, as well as the publication time and language of the news data; a filtering module, used to filter the news data to obtain filtered data; and a classification module, used to classify the filtered data according to the publication time and language of the news data to obtain classified data.

[0032] The word segmentation module described above is also used to segment the text in the categorized corpus using a word segmenter to obtain the segmented word groups.

[0033] Furthermore, the aforementioned word segmentation group is the word segmentation array output by the word segmenter, represented by a two-dimensional tensor. The aforementioned encoding processing module is also used to perform word embedding encoding processing on the word segmentation array represented by the two-dimensional tensor to obtain the word segmentation group represented by the multi-dimensional tensor. The aforementioned two-dimensional tensor includes: publication time period and word segmentation, and the aforementioned multi-dimensional tensor includes: publication time period, word segmentation, and word tensor.

[0034] Furthermore, the aforementioned text classification module includes:

[0035] The covariance calculation module calculates the covariance of the target index tensor and the word segmentation group represented by the multidimensional tensor based on the objective function to obtain the covariance value. The objective function is used to improve the positive correlation between the word segmentation group represented by the multidimensional tensor and the target index tensor.

[0036] The correlation determination module is used to determine the correlation between the target index tensor and the word segmentation group represented by the multidimensional tensor based on the covariance value of the above-mentioned BERT-CNN-based text category classification model.

[0037] Furthermore, the aforementioned device also includes:

[0038] The training module is used to train an initial BERT model through machine learning using multiple sets of sample data to obtain a trained BERT model. Each set of data in the multiple sets of sample data includes: general corpus and its corresponding text categories, news corpus and its corresponding text categories, and news corpus related to at least one technical indicator and its corresponding text categories.

[0039] The concatenation module is used to concatenate a convolutional neural network (CNN) layer before the output layer of the trained BERT model to obtain the text category classification model.

[0040] Furthermore, the aforementioned device also includes:

[0041] The second acquisition module is used to acquire the text category classification model running on different server nodes, as well as the BERT model parameters of the text category classification model.

[0042] The storage module is used to store the BERT model parameters corresponding to the different server nodes using a network disk, so as to obtain the mean-normalized model parameters based on the BERT model parameters corresponding to the different server nodes.

[0043] In the subsequent training module, a message queue is used to send the mean-normalized model parameters to the different server nodes, so that the different server nodes can use the mean-normalized model parameters to synchronize the text category classification model running locally.

[0044] On the other hand, this disclosure provides an electronic device, including: a processor and a memory connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement any of the methods described above.

[0045] On the other hand, this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement any of the methods described above.

[0046] On the other hand, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements any of the methods described above.

[0047] The text classification method for news corpora disclosed herein involves segmenting the acquired news corpora into word groups; performing word embedding encoding on the word groups to obtain word groups represented by multidimensional tensors; and employing a pre-trained BERT-CNN-based text category classification model to determine the correlation between the target indicator tensor and the word groups represented by the multidimensional tensors, thereby determining the text category corresponding to the news corpora based on the correlation. The target indicator tensor is used to represent the tensor corresponding to at least one pre-annotated technical indicator. This method addresses the problems of low accuracy in text category classification of news corpora, low accuracy of quantification strategies involving news corpora, and inconsistencies with actual market performance, thereby improving the technical effectiveness of quantification strategies involving news corpora. Attached Figure Description

[0048] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0049] Figure 1 This is a flowchart illustrating a text classification and processing method for news corpora provided in an embodiment of this disclosure;

[0050] Figure 2 This is a schematic diagram of a news corpus preprocessing process provided in an embodiment of this disclosure;

[0051] Figure 3 This is a schematic diagram of a process for pre-training an initial BERT model provided in an embodiment of this disclosure;

[0052] Figure 4 A schematic diagram illustrating the framework of a text category classification model based on BERT-CNN provided in this disclosure embodiment;

[0053] Figure 5 This is a schematic diagram of a parallel training process on different server nodes provided in an embodiment of this disclosure;

[0054] Figure 6 A structural block diagram of a text classification and processing apparatus for news corpus provided in this disclosure embodiment;

[0055] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.

[0056] The accompanying drawings have illustrated specific embodiments of this disclosure, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this disclosure to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0057] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0058] First, let's explain the terms used in this disclosure:

[0059] Quantitative strategies refer to the general term for strategies and algorithms that use computers as tools, employ quantitative methods, and follow a fixed logic to analyze, judge, and trade in the financial markets. Quantitative analysis involves objectively analyzing massive amounts of data to make decisions, using models to capture price differences, and obtaining consistent and stable returns, thereby avoiding interference from subjective human factors.

[0060] Factors: These are elements that can explain the returns of different assets. The simplest factor is the excess return of the market portfolio in CAPM, also known as the market factor (MKT).

[0061] Convolutional Neural Networks (CNNs) are a type of feedforward neural network that includes convolutional computations and has a deep structure.

[0062] Recurrent Neural Networks (RNNs) are a type of recurrent neural network that takes sequential data as input, recursively moves along the direction of the sequence, and has all nodes connected in a chain-like manner.

[0063] Long Short-Term Memory (LSTM) Network: LSTM is a type of recurrent neural network designed specifically to address the long-term dependency problem present in recurrent neural networks.

[0064] Autonomous mechanism: It provides an effective modeling method for capturing global context information through triplet parameters.

[0065] Transformers neural network model: A parallel neural network model based on an autonomous force mechanism that can effectively alleviate the gradient vanishing problem in recurrent neural networks.

[0066] BERT neural network model: It is a novel pre-trained language model that uses the encoder layer of the transformer neural network model for feature extraction. It adopts a training mode of pre-training and fine-tuning, and learns deep word-level and sentence-level features through the Masked Language Model (LM) task and the Next Sentence Prediction (NSP) task. It is trained and tested on different downstream tasks through fine-tuning to obtain the final model and experimental results.

[0067] Masked LM task: In the pre-training task of BERT neural network models, it is used to capture word-level features.

[0068] Next Sentence Prediction Task: In the pre-training task of the BERT neural network model, it is used to capture sentence-level features.

[0069] Market technical stochastic oscillator (KD stochastic oscillator): also known as the full-range KDJ indicator or stochastic oscillator, is a type of technical analysis indicator.

[0070] On-Balance Volume (OBV) indicator: Used to represent the net trading volume, it is the trend of changes in demand and supply in the cumulative trading volume each day.

[0071] The Market Trend Change Indicator (MACD) is used to represent the moving average of convergence and divergence. It is derived from the double exponential moving average. Changes in the MACD indicate changes in market trends.

[0072] In an era of innovation and change, information technology, engineering technology, the internet, the Internet of Things, and even the Internet of Everything have transformed people's lives, from clothing and food to housing and transportation. Now, these technologies are gradually permeating the realm of intellectual games, such as quantitative strategies.

[0073] There are many factors involved in current quantitative strategies. Among them, in addition to technical indicators, news corpus also has a very important impact on quantitative strategies. Properly handling news corpus can effectively improve the accuracy of quantitative decision-making.

[0074] However, firstly, the current mainstream approach to processing news corpora is based on Long Short-Term Memory (LSTM) networks or BERT networks for supervised training in sentiment classification. The objective function is mainly calculated by cross-entropy between the output tensor of the neural network and a pre-labeled target tensor. Existing solutions are not correlated with market technical indicators. Secondly, traditional methods are mostly based on open-source traditional BERT models, which are built on everyday corpora and do not undergo additional domain-specific corpora retraining or downstream task setting for the pre-training task. Thirdly, traditional BERT models do not require high iterative capabilities, so they are typically trained sequentially. However, quantization requires rapid responses, necessitating rapid model iteration. Therefore, this disclosure also involves parallel training methods.

[0075] The text classification and processing method for news corpora disclosed herein aims to solve the above-mentioned technical problems in the prior art.

[0076] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.

[0077] Figure 1 This is a flowchart illustrating a text classification processing method for news corpora provided in this embodiment of the present disclosure, as shown below. Figure 1 As shown, the method includes:

[0078] S101, perform word segmentation on the acquired news corpus to obtain segmented word groups.

[0079] S102, perform word embedding encoding on the above word segmentation array to obtain word segmentation groups represented by multidimensional tensors.

[0080] S103. Using a pre-trained BERT-CNN-based text category classification model, the correlation between the target index tensor and the word segmentation groups represented by the above multidimensional tensor is determined, so as to determine the text category corresponding to the above news corpus based on the above correlation.

[0081] Optionally, in this embodiment of the disclosure, the target index tensor is used to characterize the tensor corresponding to at least one technical index obtained through pre-annotation.

[0082] In one example, the aforementioned technical indicators include at least one of the following: market technical stochastic oscillator, net daily trading volume indicator, and market trend indicator.

[0083] Optionally, the aforementioned news corpus can be collected from a news corpus, for example, by focusing on collecting well-known magazines or electronic newspapers such as McKinsey and Deakin, and the resulting corpus data can be used as the news corpus in this embodiment of the disclosure.

[0084] After collecting news corpora, magazines and articles with lower quality can be filtered out to improve the accuracy of determining the text category classification of the news corpora.

[0085] Since most news corpora use a paragraph format of introduction, body, and conclusion, with the first and last paragraphs usually being the most important, an extractor can be used to extract the text from the first and last paragraphs of the filtered news corpus. For the middle paragraphs, text is extracted using random sampling. Then, a word segmenter is used to segment the extractor's results. The purpose of word segmentation is to segment the extractor's results into words; for example, different word segmenters can be used for the extractor's results in different language categories (such as Chinese and English) to obtain corresponding word groups for each language.

[0086] Optionally, in this embodiment of the disclosure, the segmented word groups output by the word segmenter can be represented by a two-dimensional tensor (K,M), where K is the publication period to characterize the date interval dimension, and M represents the word segmentation.

[0087] Subsequently, for the word segments output by the word segmenter, the word segment array can be processed by the encoder to perform word embedding encoding to obtain word segments represented by a multidimensional tensor (K,M,N), where K is the publication period to represent the date interval dimension, M represents the word segmentation, and N represents the word tensor.

[0088] The pre-trained BERT-CNN-based text category classification model provided in this embodiment, compared with the traditional BERT neural network model which uses general corpus and its corresponding text categories for pre-training, also uses news corpus and its corresponding text categories, and news corpus related to at least one technical indicator and its corresponding text categories for pre-training. Furthermore, a convolutional neural network (CNN) layer is concatenated before the output layer of the above-trained BERT model.

[0089] Therefore, by using the specially pre-trained BERT-CNN-based text category classification model provided in this embodiment, the correlation between the target index tensor and the word segmentation group represented by the multidimensional tensor can be predicted more accurately, so as to determine the text category corresponding to the news corpus based on the correlation.

[0090] In this embodiment of the disclosure, by determining the text category corresponding to the news corpus based on the above-mentioned correlation, the factor of news corpus that affects the quantification strategy can be properly handled. This can solve the problem of low accuracy in the text category classification of news corpus, which is contrary to the actual market performance, thereby improving the technical effect of the accuracy of the quantification strategy involving news corpus.

[0091] In an optional embodiment, the method further includes:

[0092] S201, Obtain the aforementioned news data from the news corpus, as well as the publication time and language of the aforementioned news data.

[0093] S202, the above news corpus is filtered to obtain the filtered corpus.

[0094] S203. Based on the publication time and language of the above news corpus, the filtered corpus is classified to obtain the classified corpus.

[0095] In this embodiment of the disclosure, such as Figure 2 As shown, the news data can be collected from a news corpus, including the publication time and language of the news data. For example, the corpus can focus on collecting well-known magazines or electronic newspapers, such as McKinsey and Deakin, and the resulting data can be used as the news data in this embodiment.

[0096] After collecting news corpora, magazines and articles with lower quality can be filtered out to improve the accuracy of determining the text category classification of the news corpora.

[0097] The filtered corpus is then categorized based on the publication period and language type determined from the collected news corpus, resulting in categorized corpus. For example, categorizing by publication period is primarily for matching technical indicators corresponding to the time when the news corpus was published. For instance, a news corpus from February to April of a certain year is generally only associated with technical indicators from the same period.

[0098] In conjunction with the above optional embodiments, in step S101, the acquired news corpus is segmented to obtain segmented word groups, including:

[0099] S1010, A word segmenter is used to segment the text in the corpus after the above classification to obtain the above segmented word groups.

[0100] Optionally, since most news corpora use a general-specific-general paragraph format, with the first and last paragraphs usually being the most important, this embodiment of the disclosure will still follow the same format. Figure 2As shown, an extractor can also be used to extract text segments from the categorized corpus to extract the text from the first and last segments, and to extract text from the middle segments by random sampling to obtain the extraction results.

[0101] Next, a word segmenter is used to segment the extraction results of the extractor. The purpose of word segmentation is to segment the extraction results of the extractor. For example, different word segmenters are used for extraction results of different language categories (such as Chinese and English) to obtain word groups (word group list) corresponding to different languages.

[0102] Furthermore, in this embodiment of the disclosure, a corpus dictionary is used to index and map each word in the above-mentioned word segmentation group to obtain an integer array after index mapping.

[0103] In another example, the aforementioned word segmentation group is the word segmentation array represented by a two-dimensional tensor output by the word segmenter. In step S102, word embedding encoding is performed on the word segmentation array to obtain a word segmentation group represented by a multi-dimensional tensor, including:

[0104] S1020, the word segmentation array represented by the above two-dimensional tensor is subjected to word embedding encoding processing to obtain the word segmentation group represented by the above multi-dimensional tensor, wherein the above two-dimensional tensor includes: publication time period and word segmentation, and the above multi-dimensional tensor includes: publication time period, word segmentation and word tensor.

[0105] Optionally, in this embodiment of the disclosure, the segmented word groups output by the word segmenter can be represented by a two-dimensional tensor (K,M), where K is the publication period to characterize the date interval dimension, and M represents the word segmentation.

[0106] Subsequently, for the word segments output by the word segmenter, the word segment array can be processed by the encoder to perform word embedding encoding to obtain word segments represented by a multidimensional tensor (K,M,N), where K is the publication period to represent the date interval dimension, M represents the word segmentation, and N represents the word tensor.

[0107] For example, an encoder is used to perform word embedding encoding on the two-dimensional integer tensor output by the word segmenter. Each word is mapped to a word vector of length 768 according to the word vector dictionary. The final information dimension is a three-dimensional tensor of (K,M,N), where N is 768.

[0108] In this embodiment of the disclosure, word segments represented by multidimensional tensors are obtained to obtain input data for text category classification processing of a text category classification model based on BERT-CNN. This can more accurately predict the correlation between the target index tensor and the word segments represented by the multidimensional tensors.

[0109] In one example, the pre-trained BERT-CNN-based text category classification model used above determines the correlation between the target index tensor and the word segmentation groups represented by the multidimensional tensor, including:

[0110] S301, based on the objective function, calculate the covariance of the target index tensor and the word segmentation group represented by the multidimensional tensor to obtain the covariance value.

[0111] S302, using the BERT-CNN-based text category classification model, the correlation between the target index tensor and the word segmentation group represented by the multidimensional tensor is determined based on the covariance value.

[0112] In this embodiment of the disclosure, the objective function is used to improve the positive correlation between the word segmentation group represented by the multidimensional tensor and the target indicator tensor; the target indicator tensor is used to characterize the tensor corresponding to at least one technical indicator obtained by pre-annotation, namely the market technical indicator tensor.

[0113] To make the text category classification results output by the BERT-CNN-based text category classification model more closely resemble the performance of the target indicator tensor, such as market technical indicators, the positive correlation between the target indicator tensor and the word segmentation groups represented by the aforementioned multidimensional tensor is improved.

[0114] In this embodiment of the disclosure, three market technical indicator tensors Yi (such as KD stochastic oscillator, OBV indicator, MACD indicator) and word segmentation phrases represented by multidimensional tensors Xi can be selected for covariance calculation, wherein the mean output tensor is X, and the mean market technical indicator tensor is Y.

[0115] Since the value of covariance ranges from -1 to +1, -1 indicates that the two are most negatively correlated, meaning that the text category classification results output by the text category classification model are opposite to the situation reflected by market technical indicators, and +1 indicates that the two are most positively correlated, meaning that the text category classification results output by the text category classification model are basically consistent with the situation reflected by market technical indicators.

[0116] To make the covariance value closer to +1, the following objective function is pre-designed in this embodiment:

[0117] Objective function = maxnF;

[0118] covariance function Where k is the index number of the news corpus, and n is the number of news corpora.

[0119] In addition, there is an optional implementation: since the three market technical indicators may differ too much, resulting in poor fit, softmax normalization can be performed on the three market technical indicators before calculating the covariance to improve the fit.

[0120] In this embodiment of the disclosure, an objective function is designed to improve the positive correlation between news corpus and market performance by combining some representative technical indicators that can be used to reflect the market performance of news corpus. This solves the problems of weak positive correlation between prediction results and market performance and slow sequential training speed in traditional BERT models.

[0121] In one example, the above method further includes:

[0122] S401, using multiple sets of sample data to train an initial BERT model through machine learning to obtain a trained BERT model. Each set of data in the multiple sets of sample data includes: general corpus and its corresponding text categories, news corpus and its corresponding text categories, and news corpus related to at least one technical indicator and its corresponding text categories.

[0123] S402, before the output layer of the BERT model trained above, a convolutional neural network (CNN) layer is concatenated to obtain the text category classification model described above.

[0124] In this embodiment of the disclosure, the pre-trained BERT-CNN-based text category classification model mainly extracts semantic features. To improve the efficiency of pre-training, a parallel training method can be adopted, executing three training methods simultaneously. Alternatively, it can be as follows: Figure 3 The three-stage sequential training method shown is as follows: In addition to using general corpora and their corresponding text categories for pre-training, the traditional BERT neural network model also uses news corpora and their corresponding text categories, and news corpora related to at least one technical indicator and their corresponding text categories for pre-training, to obtain the trained BERT model.

[0125] like Figure 4 As shown, the network layer structure in the trained BERT model includes, in sequence: input layer, self-attention layer, summation and normalization layer, feedforward neural network layer, summation and normalization layer, and output layer.

[0126] It should be noted that, as Figure 4 The diagram shown is a simplified representation of the BERT-CNN-based text category classification model. In actual model construction, the trained BERT model on the right can be overlapped 12 times to achieve better data fit.

[0127] Still Figure 4As shown in this embodiment, the trained BERT model obtained through the above three-stage training can be further modified by concatenating a convolutional neural network (CNN) layer before the output layer of the trained BERT model, i.e., between the summation and normalization layer and the output layer, to obtain the final text category classification model. By concatenating a CNN layer before the output layer of the trained BERT model, the dimension of the output tensor of the BERT-CNN-based text category classification model can be made the same as the number of market technical indicators.

[0128] In another example, such as Figure 5 As shown, the above method also includes:

[0129] S501, obtain the text category classification models running on different server nodes, as well as the BERT model parameters of the text category classification models.

[0130] S502, use a network disk to store the BERT model parameters corresponding to the different server nodes, so as to obtain the mean-normalized model parameters based on the BERT model parameters corresponding to the different server nodes.

[0131] S503, a message queue is used to send the mean-normalized model parameters to the different server nodes, so that the different server nodes can use the mean-normalized model parameters to synchronize the text category classification model running locally.

[0132] In this embodiment of the disclosure, in order to meet the rapid iteration speed of the text category classification model and improve the prediction ability of the text category classification model, a cluster of multiple server nodes can be used to train the text category classification model running on different server nodes in parallel.

[0133] Optionally, the text category classification models run on different server nodes may be, but are not limited to, trained independently by different server nodes using the pre-training method provided in the embodiments of this disclosure; in addition, parallel training may also be performed using, but is not limited to, the pre-training method provided in the embodiments of this disclosure.

[0134] During parallel training, the BERT model parameters trained on different server nodes at the same time are inconsistent. Therefore, by obtaining the text category classification models running on different server nodes and their BERT model parameters, the BERT model parameters of the text category classification models running on different server nodes can be normalized by mean.

[0135] For example, network disks can be used to store the BERT model parameters corresponding to the different server nodes, so as to obtain the mean-normalized model parameters based on the BERT model parameters corresponding to the different server nodes, and generate a unified text category classification model.

[0136] Then, the parameters of the unified text category classification model BERT are passed to different server nodes for parameter synchronization. However, the efficiency of directly passing parameters between different server nodes in a point-to-point manner is low.

[0137] Furthermore, during the entire training process, different server nodes may synchronize parameters many times, so synchronization efficiency is also very important. Therefore, in this embodiment, a message queue is used to send the mean-normalized model parameters to the different server nodes, so that the different server nodes can use the mean-normalized model parameters to synchronize the locally running text category classification model. After that, the process of independent training by different server nodes and parallel training by multiple server nodes is executed in an infinite loop.

[0138] This disclosure, through reasonable fine-tuning of the pre-trained BERT model, connecting it to CNN layers, and formulating an objective function in conjunction with technical indicator factors for downstream task training, can effectively increase the positive correlation between news corpora and technical indicator factors, avoiding misleading news articles from making incorrect judgments. It also provides a method for parallel training of the BERT model, accelerating the training speed and enabling faster development of reasonable quantization strategies.

[0139] According to one or more embodiments of this disclosure, a text classification and processing apparatus for news corpora is provided. Figure 6 A structural block diagram of a text classification and processing apparatus for news corpora provided in this disclosure embodiment is shown below. Figure 6 As shown, the above-mentioned device includes:

[0140] The word segmentation module 600 is used to segment the acquired news corpus into words to obtain word groups.

[0141] The encoding processing module 601 is used to perform word embedding encoding processing on the above-mentioned word segmentation array to obtain word segmentation groups represented by multidimensional tensors;

[0142] The text classification module 602 is used to determine the correlation between the target indicator tensor and the word segmentation group represented by the multidimensional tensor using a pre-trained BERT-CNN-based text category classification model, so as to determine the text category corresponding to the news corpus based on the correlation. The target indicator tensor is used to represent the tensor corresponding to at least one technical indicator obtained by pre-annotation.

[0143] According to one or more embodiments of this disclosure, the above-mentioned apparatus further includes: a first acquisition module, configured to acquire the news corpus in the news corpus, as well as the publication time and language type of the news corpus; a filtering module, configured to filter the news corpus to obtain filtered corpus; and a classification module, configured to classify the filtered corpus according to the publication time and language type of the news corpus to obtain classified corpus.

[0144] The word segmentation module described above is also used to segment the text in the categorized corpus using a word segmenter to obtain the segmented word groups.

[0145] According to one or more embodiments of this disclosure, the above-mentioned word segmentation group is the word segmentation array output by the word segmenter and represented by a two-dimensional tensor. The above-mentioned encoding processing module is further used to perform word embedding encoding processing on the word segmentation array represented by the two-dimensional tensor to obtain the word segmentation group represented by the multi-dimensional tensor. The above-mentioned two-dimensional tensor includes: publication time period and word segmentation, and the above-mentioned multi-dimensional tensor includes: publication time period, word segmentation and word tensor.

[0146] According to one or more embodiments of this disclosure, the above-described text classification module includes:

[0147] The covariance calculation module calculates the covariance of the target index tensor and the word segmentation group represented by the multidimensional tensor based on the objective function to obtain the covariance value. The objective function is used to improve the positive correlation between the word segmentation group represented by the multidimensional tensor and the target index tensor.

[0148] The correlation determination module is used to determine the correlation between the target index tensor and the word segmentation group represented by the multidimensional tensor based on the covariance value of the above-mentioned BERT-CNN-based text category classification model.

[0149] According to one or more embodiments of this disclosure, the above-described apparatus further includes:

[0150] The training module is used to train an initial BERT model through machine learning using multiple sets of sample data to obtain a trained BERT model. Each set of data in the multiple sets of sample data includes: general corpus and its corresponding text categories, news corpus and its corresponding text categories, and news corpus related to at least one technical indicator and its corresponding text categories.

[0151] The concatenation module is used to concatenate a convolutional neural network (CNN) layer before the output layer of the trained BERT model to obtain the text category classification model.

[0152] According to one or more embodiments of this disclosure, the above-described apparatus further includes:

[0153] The second acquisition module is used to acquire the text category classification model running on different server nodes, as well as the BERT model parameters of the text category classification model.

[0154] The storage module is used to store the BERT model parameters corresponding to the different server nodes using a network disk, so as to obtain the mean-normalized model parameters based on the BERT model parameters corresponding to the different server nodes.

[0155] In the subsequent training module, a message queue is used to send the mean-normalized model parameters to the different server nodes, so that the different server nodes can use the mean-normalized model parameters to synchronize the text category classification model running locally.

[0156] In an exemplary embodiment, this disclosure also provides an electronic device, including: a processor, and a memory connected to the processor;

[0157] The aforementioned memory stores instructions executed by the computer;

[0158] The processor executes computer execution instructions stored in the memory to implement any of the methods described above.

[0159] In an exemplary embodiment, this disclosure also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in any of the embodiments.

[0160] In an exemplary embodiment, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements any of the methods described herein.

[0161] To implement the above embodiments, this disclosure also provides an electronic device.

[0162] refer to Figure 7The diagram illustrates a structural schematic of an electronic device 700 suitable for implementing embodiments of the present disclosure. The electronic device 700 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0163] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0164] Typically, the following devices can be connected to the I / O interface 705: input devices 707 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 707. The communication device 707 allows the electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although... Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0165] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 707, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.

[0166] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0167] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0168] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0169] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0171] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0172] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0173] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

Claims

1. A text classification and processing method for news corpora, characterized in that, The method includes: The acquired news corpus is segmented to obtain segmented word groups; the segmented word groups are word arrays represented by two-dimensional tensors output by the word segmenter. The word segmentation array represented by the two-dimensional tensor is subjected to word embedding encoding to obtain word segmentation groups represented by a multi-dimensional tensor. The two-dimensional tensor includes: publication time period and word segmentation, and the multi-dimensional tensor includes: publication time period, word segmentation and word tensor. The covariance of the target index tensor and the word segmentation group represented by the multidimensional tensor is calculated based on the objective function to obtain the covariance value. The objective function is used to improve the positive correlation between the word segmentation group represented by the multidimensional tensor and the target index tensor. Adopting based on The text category classification model determines the correlation between the target index tensor and the word segmentation group represented by the multidimensional tensor based on the covariance value, so as to determine the text category corresponding to the news corpus based on the correlation, wherein the target index tensor is used to characterize the tensor corresponding to at least one technical indicator obtained by pre-annotation. The method further includes: Obtain the news data from the news corpus, as well as the publication time and language of the news data; The news corpus is filtered to obtain the filtered corpus; Based on the publication time and language of the news corpus, the filtered corpus is classified to obtain the classified corpus; The step of performing word segmentation on the acquired news corpus to obtain word segments includes: using a word segmenter to segment the text in the categorized corpus to obtain the word segments.

2. The method according to claim 1, characterized in that, The method further includes: The initial BERT model was trained using machine learning with multiple sets of sample data to obtain the trained BERT model. Each set of sample data included: general corpus and its corresponding text categories, news corpus and its corresponding text categories, and news corpus related to at least one technical indicator and its corresponding text categories. A convolutional neural network (CNN) layer is concatenated before the output layer of the trained BERT model to obtain the text category classification model.

3. The method according to claim 1, characterized in that, The method further includes: Obtain the text category classification model running on different server nodes, and the BERT model parameters of the text category classification model; The BERT model parameters corresponding to the different server nodes are stored on a network disk, so as to obtain the mean-normalized model parameters based on the BERT model parameters corresponding to the different server nodes. The mean-normalized model parameters are sent to the different server nodes using a message queue, so that the different server nodes can use the mean-normalized model parameters to synchronize the text category classification model running locally.

4. The method according to claim 1, characterized in that, The technical indicators include at least one of the following: market technical stochastic oscillator, daily net trading volume indicator, and market trend indicator.

5. A text classification and processing device for news corpora, characterized in that, The device includes: The word segmentation module is used to segment the acquired news corpus into words to obtain segmented word groups; the segmented word groups are word arrays represented by two-dimensional tensors output by the word segmenter. The encoding processing module is used to perform word embedding encoding processing on the word segmentation array represented by the two-dimensional tensor to obtain word segmentation groups represented by a multi-dimensional tensor, wherein the two-dimensional tensor includes: publication time period and word segmentation, and the multi-dimensional tensor includes: publication time period, word segmentation and word tensor; The text classification module includes: The covariance calculation module is used to calculate the covariance between the target index tensor and the word segmentation group represented by the multidimensional tensor based on the objective function, and obtain the covariance value. The objective function is used to improve the positive correlation between the word segmentation group represented by the multidimensional tensor and the target index tensor. The correlation determination module is used to employ a correlation-based approach. The text category classification model determines the correlation between the target index tensor and the word segmentation group represented by the multidimensional tensor based on the covariance value, so as to determine the text category corresponding to the news corpus based on the correlation, wherein the target index tensor is used to characterize the tensor corresponding to at least one technical indicator obtained by pre-annotation. The device further includes: The first acquisition module is used to acquire the news corpus in the news corpus, as well as the publication time and language of the news corpus; The filtering module is used to filter the news corpus to obtain filtered corpus. The classification module is used to classify the filtered corpus according to the publication time and language of the news corpus to obtain the classified corpus. The word segmentation module is further used to perform word segmentation on the text in the categorized corpus using a word segmenter to obtain the segmented word groups.

6. An electronic device, characterized in that, include: A processor, and a memory connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 4.

8. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Fund product recommendation method, device and equipment

    CN112102095A

  • BERT-CNN-based financial text classification method and system

    CN114064888A