A text analysis method, system, device and medium based on text structure
By extracting text structures in long text machine reading and using Transformer and Longformer networks for embedded vector fusion, the problem of structural neglect in long text prediction is solved, achieving higher accuracy and interpretability.
Patent Information
- Application Number
- CN202210145827.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-02-17
AI Technical Summary
The prior art ignores text structure in long text machine reading, resulting in insufficient prediction accuracy and interpretability.
By extracting the abstract and paragraph titles and content in the text structure, using Transformer and Longformer networks for embedded vector fusion, combining the attention interaction between paragraph titles and paragraph content, optimizing the sliding window attention mechanism to achieve a structured understanding of long text.
It improves the accuracy and interpretability of long text predictions, reduces computing and memory consumption, and improves the adaptability of the application scenarios of the model.
Smart Images

Figure CN114611484B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data mining, specifically to the field of text analysis, and particularly to a text analysis method, system, device and medium based on text structure. Background Art
[0002] Data mining using various types of public information has always been an important direction in the research and development of the field of natural language processing. However, from the publicly written text by the author to the final prediction result, the processing complexity of long articles and the subjective randomness during the author's writing have brought huge challenges to the prediction accuracy. The preliminary investigation of the project shows that directly modeling the content of long texts without considering the organizational structure of long articles is not an ideal solution. Although some methods of this kind of thinking have achieved seemingly relatively ideal prediction accuracy, their algorithms ignore the organizational structure of long texts and only consider the text content, making it difficult to convince the public in terms of the interpretability of the results.
[0003] Compared with short text data such as user comments, the length of long texts has increased significantly, and the complexity and difficulty of processing have also increased accordingly. Models that perform well in short text modeling and processing often perform mediocrely in long text processing. Some "fail to grasp the key points", and some algorithms have too high complexity, consuming a lot of time and effort.
[0004] Currently, there are many works studying how to design better models to efficiently and appropriately process long text data, which are introduced separately below:
[0005] The improved model based on the Long Short-Term Memory (LSTM) algorithm has achieved good results in the field of machine reading, such as the Long Short-Term Memory Neural Network (Cheng et al., 2016). The Multi-Timescale Long Short-Term Memory Neural Network (Liu et al., 2015) is a pioneer in the field of long text modeling. This model not only solves the defect that the LSTM model has very low efficiency in processing long texts, but also can capture the connection between words that are far apart in the text, and is an excellent improvement of LSTM in the field of long text machine reading. However, the structure of this model is simple and it can only learn according to the order of text vocabulary, lacking the ability to understand the text structure.
[0006] Text Convolutional Neural Networks (Kim, 2014) is a modification of CNN in the text field, enabling it to process variable-length text data and achieving good results on datasets. Structurally, it can also be regarded as a simplified version of Dynamic Convolutional Neural Networks (Kalchbrenner et al., 2014). However, such a model structure still ignores the paragraph levels divided by the article author during writing and cannot understand the text well.
[0007] The attention mechanism also makes great contributions in the field of long text machine reading. Hierarchical attention networks (Yang et al., 2016) focus on the structural attributes of long texts, dividing the article into three levels: article, sentence, and word. In each sentence, it focuses on the words with the highest weights, and then at the article level, it focuses on the sentences with high weights, thus completing the machine's understanding of the article. This method is very close to the human reading habit in terms of thinking and has strong interpretability. The disadvantage of this model structure is that it does not show its superiority in the application scenario. That is to say, although it achieves an understanding based on the text structure, it has not found a good application scenario to demonstrate the advantages of this text structure-based understanding.
[0008] Since Google proposed BERT in 2018, its power in the text processing field has been obvious to all. Various modified versions in the field of long text machine reading have emerged as the times require. The Adaptive Attention Network (Sukhbaatar et al., 2019) modified the defect of the BERT model (Devlin et al., 2018) in calculating global self-attention, changing it to learning within a certain window span, which greatly saves computing power. However, this model ignores the text structure again, and the artificially specified window span still cuts the text rigidly, resulting in understanding deviations.
[0009] Longformer (Beltagy et al., 2020) continues to modify the attention mechanism in BERT, successively using a sliding window mechanism, a dilated sliding window mechanism, and a sliding window mechanism that integrates global information, significantly reducing the computational complexity and memory consumption of the self-attention mechanism and achieving good results. However, this model still does not pay enough attention to the text structure, and for highly structured articles, the performance of the model is still not satisfactory. Summary of the Invention
[0010] Aiming at the problem that traditional machine reading of long texts ignores the text structure, the purpose of the present invention is to provide a text analysis method, system, device and medium based on the text structure. By mining the information contained in the text structure and combining the text content for prediction, the prediction accuracy can be effectively improved.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] In the first aspect, the present invention provides a text analysis method based on the text structure, which includes the following steps:
[0013] Parse the obtained text to be analyzed to obtain its text structure;
[0014] Perform machine reading on each text structure of the text to be analyzed respectively to obtain the embedding vectors corresponding to the respective text structures;
[0015] Fuse the obtained embedding vectors to obtain a fused article embedding vector;
[0016] Based on the fused article embedding vector and a pre-constructed prediction network model, obtain the text analysis result.
[0017] Furthermore, the method for parsing the text to be analyzed to obtain its text structure includes:
[0018] Crawl the original html file of the text to be analyzed from the website;
[0019] Traverse each node of the original html file;
[0020] Judge whether there is <a-sum>Label, if any, will be <a-sum>The content below the label serves as the abstract of the article;
[0021] Then determine whether the original html file has <header>Label, if any, will <header>The content below the label is used as the paragraph title, otherwise the bold content on a separate line is used as the paragraph title;
[0022] Extract the text content below the paragraph title as the paragraph content;
[0023] Number the paragraph titles and the corresponding paragraph contents in sequence to obtain the paragraph titles and their matching paragraph contents.
[0024] Furthermore, the method for machine reading of each text structure of the text to be analyzed includes:
[0025] Input the obtained abstract part into a pre-constructed first Transformer network to obtain the embedding vector of the abstract;
[0026] Input each paragraph title into a pre-constructed second Transformer network to obtain the embedding vector of each paragraph title;
[0027] Input the paragraph content corresponding to each paragraph title into a pre-constructed Longformer network to obtain the embedding vector corresponding to the paragraph content.
[0028] Furthermore, the method for fusing the obtained embedding vectors to obtain the fused article embedding vector includes:
[0029] Fuse the embedding vectors of the paragraph title and the corresponding paragraph content to obtain the paragraph embedding vector;
[0030] Fuse the paragraph embedding vectors to obtain the article content embedding vector;
[0031] Fuse the abstract embedding vector and the article content embedding vector to obtain the article vector of the entire text to be analyzed.
[0032] Furthermore, the method for fusing the embedding vectors of the paragraph title and the corresponding paragraph content includes:
[0033] Fuse the embedding vector of the paragraph title and the embedding vector of the paragraph content by taking the average value in pairs to complete the fusion, and obtain the paragraph embedding vector.
[0034] Furthermore, the method for fusing the paragraph embedding vectors to obtain the article content embedding vector includes:
[0035] When fusing the paragraph embedding vectors, take the average value of the paragraph embedding vectors in pairs to complete the fusion, and obtain the article content embedding vector.
[0036] Furthermore, the method for obtaining the text analysis result based on the fused article embedding vector includes:
[0037] Batch-normalize the article embedding vectors of all texts to be analyzed;
[0038] Input the batch-normalized article embedding vectors into a single-layer neural network for dimensionality reduction, and obtain the classification result with the highest probability among the three potential classification results through the Softmax algorithm.
[0039] In a second aspect, the present invention provides a text analysis system based on text structure, the system comprising:
[0040] A text parsing module, configured to parse the obtained text to be analyzed to obtain its text structure;
[0041] A text reading module, configured to perform machine reading on each text structure of the text to be analyzed respectively to obtain the embedding vectors corresponding to the respective text structures;
[0042] A vector fusion module, configured to fuse the obtained respective embedding vectors to obtain a fused article embedding vector;
[0043] A prediction module, configured to obtain a text analysis result based on the fused article embedding vector.
[0044] In a third aspect, the present invention provides a processing device, the processing device at least comprising a processor and a memory, a computer program is stored on the memory, and when the processor runs the computer program, the steps of implementing the text analysis method based on text structure are executed.
[0045] In a fourth aspect, the present invention provides a computer storage medium, on which computer-readable instructions are stored, and the computer-readable instructions can be executed by a processor to implement the steps of the text analysis method based on text structure.
[0046] Due to the above technical solutions adopted by the present invention, it has the following advantages:
[0047] 1. Full utilization of article structure. The present invention takes into account the important significance of article structure for machine understanding. Traditional machine reading methods often treat all words equally and only strengthen the representation of a certain vocabulary through the attention mechanism, without reading from the perspective of article structure. The method proposed by the present invention fully considers the unique structure of the analyzed article and parses it according to the structure of abstract - paragraph {paragraph title - paragraph content}, enabling the model to have the ability to read by structure.
[0048] 2. Attention Interaction between Paragraph Headings and Paragraph Contents. Based on Longformer, the present invention further optimizes the sliding window attention mechanism in combination with the actual application of analyzing articles, and uses the semantics of paragraph headings to strengthen the semantic representation of paragraph contents. This attention interaction method greatly reduces the consumption of computing power and memory by the self-attention mechanism of Transformer, and at the same time takes into account the actual application scenarios, improving the machine reading ability of Longformer.
[0049] Therefore, the present invention can be widely applied to the field of text analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0051] Figure 1 is a comparison schematic diagram between the present invention and the traditional method;
[0052] Figure 2 is a flowchart of the text analysis method based on the text structure of the present invention;
[0053] Figure 3 is the overall research plan of the present invention;
[0054] Figure 4 is a flowchart of the article structure analysis of the present invention;
[0055] Figure 5 is a flowchart of the machine reading of each structure of the present invention;
[0056] Figure 6 is a flowchart of the vector fusion of each structure of the present invention;
[0057] Figure 7 is a flowchart of the stock price increase and decrease prediction in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.
[0059] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0060] Through analysis, the present invention discovers that: at the beginning of analysis articles, it is common for authors to summarize and generalize their viewpoints as a summary or abstract, aiming to generalize the central idea of the article and save readers' time. This is very important text structure information and content information. When performing machine reading on such analysis articles, it is natural to consider this part of the content. Therefore, by separately extracting this part of the content as an independent vector and then fusing it with the vectors extracted from other parts, the present invention can have higher interpretability than current prediction methods that do not extract this part of the content.
[0061] At the same time, based on the text structure, the present invention also observes that many authors have the writing habit of writing paragraph titles. Paragraph titles often reflect the framework of the author's overall planning and the thinking of writing, and are the "backbone" of the whole text. Although they have few words, they often serve as keywords and theme words, and their importance is self-evident. The paragraph content after the paragraph title is a series of arguments and elaborations centered around the title, playing the role of strengthening the attitude of the title, and is the "flesh and blood" of the whole text. Regarding the mapping relationship between the two, traditional machine reading methods regard them as equal relationships and cannot well learn the master-slave structure between the two, and thus lack insight into the structure of the article. If predictions are made based on this defective learning result, the results will naturally be unsatisfactory. Therefore, the present invention uses the unique format features of paragraph titles to independently identify them as "titles", and then binds the titles with the paragraph content after the titles to achieve a one-to-one correspondence between paragraph titles and paragraph content, thereby providing a feasible learning structure for machine reading of long texts. To solve the long-standing difficulty that machine reading cannot effectively learn long text information.
[0062] For a long text, the first two methods ensure machine reading of each local structure of the text. For the overall article, it is necessary to conduct a macroscopic and overall analysis of the connections between the structures and then further learn. For this reason, the present invention designs a fusion learning method, so that the final result of the model not only includes information of each local part but also covers information on the associations between the local parts.
[0063] Based on the above analysis, as Figure 1 As shown in the figure, in some embodiments of the present invention, a text analysis method based on text structure is provided, which mainly consists of a deep learning method for article summary or abstract content, a deep learning method for paragraph titles and paragraph content, the fusion of paragraph title and paragraph content vectors, and the fusion of abstract vectors and paragraph vectors. By extracting the article structure of the text to be analyzed, embedding vectors are respectively modeled and extracted for the content at different positions such as the abstract or summary, paragraph titles, and paragraph content, and then the extracted embedding vectors are fused for prediction. The present invention can make full use of the inherent paragraph structure characteristics within the article, while taking into account the modeling analysis of the article content, making the algorithm more interpretable than various previous methods, and also achieving a significant improvement in prediction performance.
[0064] Correspondingly, in some other embodiments of the present invention, a text analysis system, device, and medium based on text structure are provided.
[0065] Embodiment 1
[0066] As Figure 2 、 Figure 3 shown, this embodiment provides a text analysis method based on text structure, including the structural analysis process of the article and the reading process of each structure. Among them, the structural analysis process of the article includes: the recognition of the abstract and paragraph titles and the binding of paragraph titles and their affiliated content; the reading process of each structure includes: the machine reading of the abstract, the machine reading of paragraph titles, the machine reading of paragraph content, the fusion of paragraph title and paragraph content vectors, and the fusion of paragraph vectors and abstract vectors, realizing a general machine reading method for analyzing articles. Specifically, it includes the following steps:
[0067] 1) Parse the text to be analyzed obtained to obtain its text structure;
[0068] 2) Perform machine reading on each text structure of the text to be analyzed respectively to obtain the embedding vectors corresponding to each text structure;
[0069] 3) Fuse the obtained embedding vectors to obtain a fused article embedding vector;
[0070] 4) Based on the fused article embedding vector and a pre-constructed prediction model, obtain the text analysis result.
[0071] In some implementations, in the above step 1), after parsing the text to be analyzed obtained in this embodiment, the text structure mainly includes the abstract or summary, paragraph titles, and the paragraph content matching them.
[0072] Furthermore, as Figure 4 shown, the method for parsing the text to be analyzed in this embodiment includes the following steps:
[0073] 1.1) Crawl the original html file of the text to be analyzed from the website;
[0074] 1.2) Traverse each node of the original html file;
[0075] 1.3) Determine whether there is <a-sum>Label( <a-sum>The field in html that instructs the browser to display as an abstract), if any, then <a-sum>The content below the label serves as the abstract of the article;
[0076] Among them, if there is a website link in the abstract, this is an advertisement by the author in the abstract and has nothing to do with the content of the article, so it is filtered out;
[0077] 1.4) Determine whether the original html file has a normal <header>Labels (i.e., fields in the html file that instruct the browser to display as a title), if any, will <header>The content below the label serves as the paragraph title, otherwise the bold content on a separate line serves as the paragraph title;
[0078] In particular, when <header>When there are pictures in the content below the label or in the single-line bolded content, filter out the picture titles and only keep the text part as the paragraph title;
[0079] 1.5) Extract the text content below the paragraph title as the paragraph content;
[0080] Among them, when there are non-text contents such as pictures, ordered lists, unordered lists, etc. below the paragraph title, filter out the non-text contents, and at the same time filter out other meaningless non-text contents such as disclaimers;
[0081] 1.6) Number the paragraph titles and paragraph contents in sequence to obtain the paragraph titles and their matching paragraph contents.
[0082] Among them, numbering the paragraph titles and paragraph contents in sequence is to achieve the bundling and positioning of the two through the serial numbers of the titles or contents. For example: header_1, para_1, header_2, para_2, ……. If the beginning of an article is content without a matching paragraph title, skip this title, and the paragraph titles and contents of the whole article are recorded as para_1, header_2, para_2, …….
[0083] Furthermore, as Figure 2 shown, in the above step 2), the method of machine reading for each text structure of the text to be analyzed specifically includes the following steps:
[0084] 2.1) Input the obtained abstract part into a pre-constructed first Transformer network to obtain the embedding vector of the abstract.
[0085] Preferably, in this embodiment, the first Transformer network only uses the Encoder part, its attention head is set to 6, and the word vector length is set to 512. For the entire abstract part, after calculating the vector representations of all words using the first Transformer network, take the arithmetic mean of each vector representation as the embedding vector for machine reading of the abstract part.
[0086] Specifically, in this embodiment, the Encoder layer of the Transformer network is used. This Encoder layer includes a number of Block units, and each Block unit includes a self-attention module and a feed-forward network module. Among them, the self-attention module is an algorithm module that corrects the vector of a certain word through the vectors of other words in the word sequence. More specifically, the vector sequence of each word input to the Block unit will first be multiplied by the attention head (a set of parameter matrices) to obtain the corresponding Q matrix, K matrix, and V matrix of this attention head. Then, through the product and normalization of the Q matrix and the K matrix, and multiplying and summing the result with the V matrix, the word vector Z corresponding to this attention head is obtained. Repeat the above steps with different attention heads to obtain different word vectors Z, add them up and stack them after normalization, and output the result to the subsequent feed-forward network module. In the feed-forward network module, its input is the stacked Z matrix output in the previous self-attention module, and it is passed through a fully connected network with two linear layers to output a more complex fitting result.
[0087] During the calculation process of the entire Transformer network, first, what is input into the network is to convert the article composed of word sequences into a vector sequence composed of word vectors according to the word vector table. This sequence enters the self-attention module of the first Encoder layer, and after learning by multiple attention heads, a vector sequence corrected by the front and back sequences is obtained, and then through the feed-forward network module, further fitting is performed. Then, the fitting result of the feed-forward network is input into the next Block layer, and the previous steps are repeated several times to obtain the final vector representation of the text.
[0088] The advantage of this model is that it can enhance the semantic information of words through the context text of words, and this enhancement is precisely represented by the word vectors of the Encoder layer.
[0089] 2.2) Input each paragraph title into the pre-constructed second Transformer network respectively to obtain the embedding vectors of each paragraph title.
[0090] Preferably, in this embodiment, the number of attention heads of the second Transformer network is set to 2. To reduce the computing power consumption, other parameters are the same as those of the first Transformer network.
[0091] 2.3) Input the paragraph content corresponding to each paragraph title into the pre-constructed Longformer network respectively to obtain the embedding vectors corresponding to the paragraph content.
[0092] Preferably, in this embodiment, the Longformer network is used to perform machine reading on the paragraph content matched by each paragraph title. Based on the Longformer attention mechanism, the attention calculation of the paragraph title word vectors is added. In terms of specific operations, in this embodiment, the vector representation of each word in the paragraph content part is first obtained, and then the average value is taken as the embedding vector of the paragraph content part.
[0093] Specifically, the Longformer network result is an optimization of the previous Transformer network structure. Longformer modifies the Encoder layer. Each Encoder layer includes several Block units, and each Block unit includes a sliding attention module and a feed-forward network module.
[0094] Among them, the sliding attention module is a major improvement of Longformer over Transformer. It no longer considers the entire word sequence all at once, but only calculates the word vectors within a window and uses them to correct the vector of a certain word, greatly reducing the computational load. More specifically, the vector sequence of each word input to the Block will first be segmented according to the window length of the sliding attention module. The word vectors inside the window are multiplied by the attention head (a set of parameter matrices) to obtain the Q matrix, K matrix, and V matrix corresponding to this attention head. The word vectors outside the window do not participate in the calculation. Then, through the product and normalization of Q and K, the result is multiplied and summed with the V matrix to obtain the word vector Z corresponding to this attention head. Repeat the above steps with different attention heads to obtain different word vectors Z, and add and stack them with the normalized result as the result output to the subsequent feed-forward network module.
[0095] In the feed-forward network module, the input is the stacked Z matrix output from the previous self-attention module, and it passes through a fully connected network with two linear layers to output a more complex fitting result.
[0096] In the calculation process of the entire Longformerr network, first, the article composed of a word sequence is converted into a vector sequence composed of word vectors according to the word vector table and input into the network. This sequence enters the sliding attention module of the first Encoder layer, is segmented according to the window length, and then learned by multiple attention heads to obtain a vector sequence with the front and back sequences corrected. Then, it passes through the feed-forward network module for further fitting. Next, the fitting result of the feed-forward network is input into the next Block layer, and the previous steps are repeated several times to obtain the final vector representation of the text.
[0097] Furthermore, in step 3) above, it specifically includes the following steps:
[0098] 3.1) Fuse the paragraph title with the embedding vector of the corresponding paragraph content to obtain a paragraph embedding vector;
[0099] Among them, when fusing the embedding vectors of the paragraph title and the corresponding paragraph content, the embedding vectors of the paragraph title and the paragraph content are directly averaged bit by bit to complete the fusion, rather than generating multi-line paragraph content vectors for the one-dimensional embedding.
[0100] 3.2) Fuse the embedding vectors of each paragraph to obtain the article content embedding vector;
[0101] Among them, when fusing the embedding vectors of each paragraph, the embedding vectors of each paragraph are averaged bit by bit to complete the fusion, and the article content embedding vector is obtained.
[0102] 3.3) Fuse the abstract embedding vector and the article content embedding vector to obtain the article vector of the entire text to be analyzed.
[0103] Among them, when fusing the abstract embedding vector and the article content embedding vector, the article content embedding vector and the abstract embedding vector are directly stacked into two layers, and the generated two-line vectors are the article embedding vectors of the entire text to be analyzed.
[0104] Furthermore, in the above step 4), it specifically includes the following steps:
[0105] 4.1) Batch-normalize the article embedding vectors of all texts to be analyzed;
[0106] 4.2) Input the batch-normalized article embedding vectors into a single-layer neural network for dimensionality reduction, and obtain the classification result of the article embedding vectors through the Softmax algorithm.
[0107] Embodiment 2
[0108] This embodiment takes all the stock analysis articles in the analysis module of the US stock analysis website seekingalpha.com that belong to S&P 500 enterprises and have a writing date from 2018 to 2019 as an example for introduction. It should be noted that due to the uniqueness of the data source, the structural analysis details of articles from different data sources must be different. The purpose of this section is to emphasize the process of article structure analysis. The details of the following related practical operations are all set to specifically explain this process. When the data source changes, necessary modifications need to be made according to the specific situation of the source.
[0109] 1) Analysis of article structure
[0110] As Figure 4 shown, the analysis is mainly divided into three steps: analyzing the abstract, analyzing the paragraph title, and bundling the paragraph title and its subordinate paragraph content.
[0111] Parse the abstract. Through observation, the abstract part of the stock analysis articles on the website is marked with the class named "sasource". Therefore, using "sasource" as the keyword, the location of the abstract is located. Additionally, many analysts leave their Twitter links in the abstract to drive traffic to the team. Such information is naturally useless for stock analysis and must be filtered out during parsing.
[0112] Parse the paragraph titles. The website provides a standard paragraph title format for analysts writing stock analysis articles, which is represented in html as " <h2>” or "< / h2> <h3>”. If analysts write analysis articles in accordance with this format, their paragraph headings should be led by "< / h3> <h2>” or "< / h2> <h3>” when displayed on the web page. However, in fact, only a few analysts' articles are like this. In some articles, the author wrongly formats spaces, pictures, etc. as headings, resulting in the fact that the headings led by "< / h3> <h2>” or "< / h2> <h3>” are not the real paragraph headings. In such cases, the formatting of their paragraph headings needs to be removed. In some articles, the author replaces the correct paragraph heading format with another wrong format, resulting in the fact that the paragraph headings the author really wants to show are not led by "< / h3> <h2>” or "< / h2> <h3>"Guide, and in bold (in html files with " ” Guidance), emphasizing text (in html files with " <em>” Boot), italic (in html files with " ” leads) or bullet points (in html files with " "or" ”Bootstrap” indicates that for such truly but incorrectly represented paragraph titles, the code needs to check the subordinate relationship of their context to confirm their status as paragraph titles. In this embodiment, during the parsing stage of the article structure, it is expected that all articles will be parsed into the form of {Abstract, {Paragraph Title 1, Paragraph Content 1},..., {Paragraph Title n, Paragraph Content n}}. After accurately parsing the paragraph titles of the article in the previous step, it seems that the confirmation of the paragraph content only needs to divide the text between two adjacent paragraph titles into the paragraph content of the previous paragraph title. However, this is not the case. In some articles, the background is directly described after the abstract, resulting in the absence of "Paragraph Title 1". For this situation, during the article structure parsing stage of the present invention, "Paragraph Title 1" is directly skipped, so that such articles are parsed into {Abstract, {Paragraph Title 1 (empty), Paragraph Content 1},..., {Paragraph Title n, Paragraph Content n}}. In some articles, there is only one picture after the paragraph title. During parsing, a Python program that only crawls text will not crawl down the picture, resulting in the situation where two paragraph titles are adjacent. For this situation, during the article structure parsing stage of the present invention, the paragraph content where the picture is located is assigned as empty, so that such articles are parsed into {Abstract, {Paragraph Title 1, Paragraph Content 1},..., {Paragraph Title i, Paragraph Content i (empty)}, {Paragraph Title i + 1, Paragraph Content i + 1}..., {Paragraph Title n, Paragraph Content n}}.
[0113] 2) Machine reading of each structure
[0114] As Figure 5 shown, after parsing the article according to the process of the first stage, the machine reading work of each part needs to be carried out. The present invention divides the article content into three parts, and each part uses a machine reading algorithm adapted to its own situation, which is described in detail as follows:
[0115] Machine reading of the abstract part. As mentioned before, the abstract part is not long and belongs to short text reading in the field of machine reading. In this embodiment, the mainstream short text reading algorithm in the field - Transformer is correspondingly used for its machine reading. The hyperparameters of Transformer in this embodiment, such as the number of attention heads is set to 6, the length of the word vector is 512, etc., and only the Encoder part in Transformer is used to obtain the enhanced vector representation of each word. For the entire abstract part, after calculating the vector representations of all words, the arithmetic mean is simply taken to obtain the embedded vector of the machine reading of the abstract part.
[0116] Machine reading of the paragraph title part. As mentioned above, the paragraph title part usually has less than ten words and belongs to short text reading. The machine reading method for this part in the present invention is also Transformer, but the number of attention heads is set to 2 to reduce the computing power consumption, and other hyperparameters are the same as those in the first part.
[0117] Machine reading of the paragraph content part. As mentioned above, the paragraph content part belongs to long text, and the corresponding machine reading method uses Longformer. The self-attention mechanism in Transformer requires calculating the semantic association between the target word and all words in the full text, and the computational cost caused by this in the machine reading of long texts with thousands of words is unimaginable. Therefore, Longformer improves the defect of large computational cost of the self-attention mechanism and designs a variety of novel attention calculation methods (such as: attention sliding window, spaced attention sliding window, attention sliding window integrating the global, etc.), so that the computational cost of attention is reduced to an acceptable range. After pre-training another long text reading algorithm RoBERTa (Liu et al., 2019) with the attention algorithm of Longformer, RoBERTa achieved more excellent results. Thus, it can be seen that Longformer is sufficient in the field of long text machine reading. However, due to the specific application scenario of the present invention, we must consider the association between the paragraph content and the paragraph title. Therefore, on the basis of the Longformer attention mechanism, the attention calculation of the paragraph title word vectors is added. In terms of specific operations, the present invention refers to the hyperparameter settings in the original Longformer paper to obtain the vector representation of each word in the paragraph content part, and then takes the average value as the embedding vector of the paragraph content part.
[0118] 3) Fusion of each structure vector
[0119] As Figure 6 shown, the fusion of vectors referred to in the present invention includes: the fusion of the paragraph title and the corresponding paragraph content vector, and the fusion of the abstract vector and each paragraph vector.
[0120] Fusion of the paragraph title and the corresponding paragraph content vector. Since the paragraph title vector has been fully considered through the attention mechanism when generating the paragraph content vector in the steps of machine reading of each part, when generating the paragraph vector, the paragraph title vector and the paragraph content vector are directly aligned and averaged to complete the fusion, instead of generating multiple rows of paragraph content vectors in one-dimensional embedding.
[0121] Fusion of the abstract vector and each paragraph vector. This step is the last step in generating the article vector. In this step, in this embodiment, the average value of each paragraph vector is first calculated to obtain the article content vector, and then it is directly stacked with the abstract vector into two layers without any operation, and the generated two-row vector is the article embedding vector of the whole stock analysis article.
[0122] 4) Prediction of stock price increases and decreases
[0123] As Figure 7 shown, after obtaining the article embedding vectors of multiple stock analysis articles through parallel processing, batch normalization is performed, and then they are put into a single-layer neural network, and the final predicted increase and decrease classification is obtained through the Softmax algorithm.
[0124] In this embodiment, the classification of increases and decreases is a three-classification problem. In terms of the true value, if the increase and decrease do not exceed 0.05%, this embodiment regards it as neither rising nor falling (i.e., 0), if the increase rate exceeds 0.05%, it is regarded as rising (i.e., 1), and if the decrease rate exceeds 0.05%, it is regarded as falling (i.e., -1). This embodiment crawls the increase and decrease results of S&P 500 enterprises 15 days, 30 days, 90 days, and 180 days after the release of stock analysis articles as the true values, in order to examine the prediction effect of the algorithm of the present invention in different time dimensions.
[0125] Embodiment 3
[0126] The above Embodiment 1 provides a text analysis method based on text structure. Correspondingly, this embodiment provides a text analysis system based on text structure. The system provided in this embodiment can implement a text analysis method based on text structure in Embodiment 1, and this analysis system can be implemented in a software, hardware, or software-hardware combination manner. For example, the system can include integrated or separate functional modules or functional units to execute the corresponding steps in each method of Embodiment 1. Since the recognition system in this embodiment is basically similar to the method embodiment, the description process in this embodiment is relatively simple, and the relevant parts can refer to the partial description of Embodiment 1. The embodiments of the system in this embodiment are only illustrative.
[0127] A text analysis system based on text structure provided in this embodiment includes:
[0128] A text parsing module, configured to parse the obtained text to be analyzed to obtain its text structure;
[0129] A text reading module, configured to perform machine reading on each text structure of the text to be analyzed respectively to obtain the embedding vector corresponding to each text structure;
[0130] A vector fusion module, configured to fuse the obtained embedding vectors to obtain a fused article embedding vector;
[0131] A prediction module, configured to obtain a text analysis result based on the fused article embedding vector.
[0132] Embodiment 4
[0133] This embodiment provides a processing device corresponding to the text analysis method based on text structure provided in Embodiment 1. The processing device can be a processing device for a client, such as a mobile phone, a laptop, a tablet computer, a desktop computer, etc., to execute the method of Embodiment 1.
[0134] The processing device includes a processor, a memory, a communication interface, and a bus. The processor, the memory, and the communication interface are connected through the bus to complete communication with each other. A computer program that can run on the processor is stored in the memory. When the processor runs the computer program, it executes the text analysis method based on text structure provided in Embodiment 1 of this application.
[0135] In some implementations, the memory can be a high-speed random access memory (RAM: Random Access Memory), and may also include non-volatile memory, such as at least one disk memory.
[0136] In other implementations, the processor can be various general-purpose processors such as a central processing unit (CPU), a digital signal processor (DSP), etc., which are not limited here.
[0137] Embodiment 5
[0138] The text analysis method based on text structure in Embodiment 1 of this application can be specifically implemented as a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing the text analysis method described in Embodiment 1 of this application are uploaded.
[0139] The computer-readable storage medium can be a tangible device that holds and stores instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.
[0140] It should be noted that the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present application. Each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code contains one or more executable instructions for implementing the specified logical function.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent substitutions, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
[0142] The above embodiments are only used to illustrate the present invention, and the structures, connection methods, manufacturing processes, etc. of each component can all be changed. Any equivalent transformation and improvement made on the basis of the technical solutions of the present invention should not be excluded from the protection scope of the present invention. < / em> < / h3> < / header> < / header> < / header> < / header> < / header>
Claims
1. A text analysis method based on text structure, characterized in that Including the following steps: Parse the obtained text to be analyzed to obtain its text structure; Perform machine reading on each text structure of the text to be analyzed respectively to obtain the embedding vectors corresponding to each text structure; Fuse the obtained embedding vectors to obtain a fused article embedding vector; Based on the fused article embedding vector and a pre-constructed prediction network model, obtain the text analysis result; The method for parsing the text to be analyzed to obtain its text structure includes: crawling the original html file of the text to be analyzed from a website; traversing each node of the original html file; determining whether there is <a-sum>Label, if any, will <a-sum>The content below the label serves as the abstract of the article; then determine whether the original html file has <header>Label, if any, will be <header>The content below the label is used as the paragraph title, otherwise the bold content in a separate line is used as the paragraph title; extract the text content below the paragraph title as the paragraph content; number the paragraph title and the paragraph content in sequence to obtain the paragraph title and the paragraph content matching it;< / header> < / header> The method of performing machine reading on each text structure of the text to be analyzed includes: inputting the obtained abstract part into a pre-constructed first Transformer network to obtain the embedding vector of the abstract; inputting each paragraph title into a pre-constructed second Transformer network to obtain the embedding vector of each paragraph title; inputting the paragraph content corresponding to each paragraph title into a pre-constructed Longformer network to obtain the embedding vector corresponding to the paragraph content; The method of fusing the obtained embedding vectors to obtain a fused article embedding vector includes: fusing the embedding vectors of the paragraph title and the corresponding paragraph content to obtain a paragraph embedding vector; fusing each paragraph embedding vector to obtain an article content embedding vector; fusing the abstract embedding vector and the article content embedding vector to obtain the article vector of the entire text to be analyzed.
2. The text analysis method based on text structure according to claim 1, characterized in that The method of fusing the embedding vectors of the paragraph title and the corresponding paragraph content includes: Fuse the embedding vector of the paragraph title and the embedding vector of the paragraph content by taking the average value in pairs to complete the fusion and obtain the paragraph embedding vector.
3. The text analysis method based on text structure according to claim 1, wherein The method of fusing each paragraph embedding vector to obtain an article content embedding vector includes: When fusing each paragraph embedding vector, take the average value of each paragraph embedding vector in pairs to complete the fusion and obtain the article content embedding vector.
4. The text analysis method based on text structure according to claim 1, characterized in that, The method of obtaining the text analysis result based on the fused article embedding vector includes: Perform batch normalization on the article embedding vectors of all texts to be analyzed; Input the batch-normalized article embedding vectors into a single-layer neural network for dimensionality reduction, and obtain the classification result with the highest probability among the three potential classification results through the Softmax algorithm.
5. A text analysis system based on text structure, characterized in that, The system includes: A text parsing module for parsing the obtained text to be analyzed to obtain its text structure; A text reading module for performing machine reading on each text structure of the text to be analyzed respectively to obtain the embedding vectors corresponding to each text structure; A vector fusion module for fusing the obtained embedding vectors to obtain a fused article embedding vector; A prediction module for obtaining the text analysis result based on the fused article embedding vector; Parse the text to be analyzed to obtain its text structure, including: crawling the original html file of the text to be analyzed from the website; traversing each node of the original html file; determining whether there is <a-sum>Label, if any, will <a-sum>The content below the label serves as the abstract of the article; then determine whether the original html file has <header>Label, if any, will <header>The content below the label is used as the paragraph title, otherwise the bold content in a separate line is used as the paragraph title; extract the text content below the paragraph title as the paragraph content; number the paragraph title and the paragraph content in sequence to obtain the paragraph title and the paragraph content matching it;< / header> < / header> Performing machine reading on each text structure of the text to be analyzed, including: inputting the obtained abstract part into a pre-constructed first Transformer network to obtain an embedded vector of the abstract; inputting each paragraph title into a pre-constructed second Transformer network to obtain an embedded vector of each paragraph title; inputting the paragraph content corresponding to each paragraph title into a pre-constructed Longformer network to obtain an embedded vector corresponding to the paragraph content; Fusing the obtained embedded vectors to obtain a fused article embedded vector, including: fusing the embedded vectors of the paragraph title and the corresponding paragraph content to obtain a paragraph embedded vector; fusing the paragraph embedded vectors to obtain an article content embedded vector; fusing the abstract embedded vector and the article content embedded vector to obtain an article vector of the entire text to be analyzed.
6. A processing device, the processing device at least includes a processor and a memory, and a computer program is stored on the memory, characterized in that, When the processor runs the computer program, it executes the steps of the text analysis method based on text structure according to any one of claims 1 to 4.
7. A computer storage medium, characterized in that, Stored thereon are computer-readable instructions, and the computer-readable instructions can be executed by a processor to implement the steps of the text analysis method based on text structure according to any one of claims 1 to 4.
Citation Information
Patent Citations
Document processor and document processing method
JP1999250041A
System and Method for Processing Contract Documents
US20200327151A1
Cited By
Reusable mail block generation and suggestion
US20260113294A1