A method and system for generating text summary based on chart
By constructing a graph summary generation model, converting the graph into a text summary, the problem of interpreting complex graph data is solved and decision-making efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202411817772.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The existing technology is difficult to effectively interpret and analyze complex chart data, which affects decision efficiency and accuracy.
By constructing a graph summary generation model, using the variational autoencoding model VAE and Transformer deep learning model, the graph is converted into a text summary, and the model is trained using the public data set.
It realizes the rapid conversion of chart data into text summary, improving the efficiency of decision-makers' information acquisition and decision-making accuracy.
Smart Images

Figure CN119296117B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for generating a text summary from a chart. Background Art
[0002] In the actual production and operation process, enterprises will continuously generate a wide variety of data, including but not limited to financial data used to evaluate financial status, conduct financial planning and forecasting, and set reasonable financial goals; thereby helping enterprises understand market demand and trends, and providing market data for product development, market positioning and marketing. In order to more effectively understand and use this data, enterprises often use various charts such as trend charts, scatter charts, pie charts, etc., and carefully draw these data on a chart so that the company's decision-makers can clearly interpret and analyze them. However, although charts can present data intuitively, readers' ability to quickly interpret complex information is limited. Faced with these charts full of data and information, decision makers often need to spend a long time and energy to understand and analyze, which will undoubtedly affect the efficiency and accuracy of decision-making. Therefore, how to more effectively present and use this data and improve decision-making efficiency has become an important issue facing enterprises. Summary of the invention
[0003] In view of the deficiencies in the prior art, the present invention provides a method for generating a text summary based on a chart, comprising the following steps:
[0004] Obtaining a public data set on the Internet as model training data, wherein the public data set includes high-quality charts and text summary information corresponding to the charts;
[0005] A high-quality chart is used as an input sample, and text summary information corresponding to the chart is used as an output label to construct a training data set, and the training data set is used to train an original chart summary generation model, wherein the chart summary generation model is composed of an encoder of a variational autoencoder model VAE, a decoder of the variational autoencoder model VAE, a text summary generation model established by a masked decoder using a Transformer deep learning model, and a loss function layer, wherein the loss function layer includes a first loss function for data with text summary and a second loss function for data without text summary;
[0006] A PDF conversion tool is used to identify and extract each target chart in the PDF document to be analyzed, and the target chart is input into the trained chart summary generation model to output the text summary information corresponding to each target chart.
[0007] Preferably, the text summarization generation model is , where Y is the text summary information corresponding to the input chart X, and the length of the summary text is T; and Is The chain rule is used to decompose and calculate the derivatives of the composite function; ψ is the parameter of the decoder with mask in the Transformer deep learning model, and s represents the encoding of the input graph X after passing through the encoder. is the independent word or phrase in the text summary information corresponding to the graph X, y=[y1,...,y t ] is the text summary corresponding to Chart X and the length of the summary text is T. ] is the text summary corresponding to Chart X and the length of the summary text is T.
[0008] Preferably, the encoder of the variational autoencoder model VAE is
[0009] , Indicates is the mean value, is the multivariate normal distribution with a variance matrix, where , X is the input chart, which is an RGB chart and a 3D tensor, F1 and F2 are the parameters adjusted by the encoder learning, and pool represents the pool transformation in the neural network CNN. For is the real number vector output by the multi-layer perceptron MLP, For is the output positive vector of the last layer of the multi-layer perceptron MLP, is the encoder parameter;
[0010] The decoder of the variational autoencoder model VAE is ,in Indicates is the mean, is the multivariate normal distribution with the variance matrix, Adjust parameters for subsequent learning; , , S is the code generated by the encoder, D1 and D2 are the parameters learned and adjusted by the decoder, represents the convolution of two 3D tensors, unpool represents the unpool transformation in the neural network CNN, and I is the identity matrix.
[0011] Preferably, the first loss function for textual summary data is configured as:
[0012] ;
[0013] The second loss function for the text-free summarization data is configured as:
[0014] ;
[0015] in For Find the expected value of the distribution s, For Find the expected value of s and z for the distribution, A setting value between 0 and 1 that is used to stabilize the training of the variational autoencoder model VAE. is the joint distribution probability function of the three variables X, s, and z; represents the code generation probability function of the encoder, For distribution The conditional distribution function of the encoder's final encoding s obtained by integrating z, where z is a random vector that obeys a multinomial distribution.
[0016] Preferably, the loss function layer is
[0017] ,in is a training data set of graphs with corresponding text summaries, A training dataset for graphs without corresponding text summaries.
[0018] The present invention also discloses a system for generating text summaries according to charts, comprising a data acquisition module, a model training module and a summary generation module, wherein the data acquisition module is used to acquire a public data set on the Internet as model training data, wherein the public data set includes high-quality charts and text summary information corresponding to the charts; the model training module is used to take the high-quality charts as input samples and the text summary information corresponding to the charts as output labels, to construct a training data set, and to use the training data set to train an original chart summary generation model, wherein the chart summary generation model is composed of an encoder of a variational autoencoder model VAE, a decoder of the variational autoencoder model VAE, a text summary generation model established by a masked decoder using a Transformer deep learning model, and a loss function layer, wherein the loss function layer includes a first loss function for data with text summary and a second loss function for data without text summary; the summary generation module is used to use a PDF conversion tool to identify and extract each target chart in a PDF document to be analyzed, and input the target chart into the chart summary generation model that has been trained to output the text summary information corresponding to each target chart.
[0019] Preferably, the text summarization generation model is
[0020] , where Y is the text summary information corresponding to the input chart X, and the length of the summary text is T; and Is The chain rule is used to decompose and calculate the derivatives of the composite function; ψ is the parameter of the decoder with mask in the Transformer deep learning model, and s represents the encoding of the input graph X after passing through the encoder. is the independent word or phrase in the text summary information corresponding to the graph X, y=[y1,...,y t ] is the text summary corresponding to Chart X and the length of the summary text is T. ] is the text summary corresponding to Chart X and the length of the summary text is T.
[0021] Preferably, the encoder of the variational autoencoder model VAE is
[0022] , Indicates is the mean value, is the multivariate normal distribution with the variance matrix, , X is the input graph, which is an RGB graph and a 3D tensor, F1 and F2 are the parameters adjusted by the encoder learning, and pool represents the pool transformation in the neural network CNN. For is the real number vector output by the multi-layer perceptron MLP, For is the output positive vector of the last layer of the multi-layer perceptron MLP, is the encoder parameter;
[0023] The decoder of the variational autoencoder model VAE is
[0024] ,in Indicates is the mean, is the variance matrix of the multivariate normal distribution, Adjust parameters for subsequent learning; , , S is the code generated by the encoder, D1 and D2 are the parameters learned and adjusted by the decoder, represents the convolution of two 3D tensors, unpool represents the unpool transformation in the neural network CNN, and I is the identity matrix.
[0025] The present invention also discloses a device for generating a text summary based on a chart, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the aforementioned methods when executing the computer program.
[0026] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the aforementioned methods are implemented.
[0027] The method and system for generating text summaries based on charts disclosed in the present invention use public data sets on the Internet as training samples to train a chart summary generation model including an encoder of a variational autoencoder model VAE, a decoder of the variational autoencoder model VAE, a text summary generation model established with a masked decoder using a Transformer deep learning model, and a loss function layer, and use high-quality charts in the training samples as input samples X of the chart summary generation model to be trained, and the text summary corresponding to the chart as the output label Y of the chart summary generation model, and train the original chart summary generation model, so that after the training is completed, the model can convert each target chart in the input PDF document extracted by PDF conversion tool into corresponding text summary information, so as to help decision makers quickly obtain the information content in the chart, and improve the efficiency of information acquisition and decision accuracy of decision makers.
[0028] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings described herein are used to provide further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0030] Figure 1 The figure is a schematic diagram of the steps of a method for generating a text summary based on a chart disclosed in an embodiment of the present invention.
[0031] Figure 2 A schematic diagram of an input sample disclosed in an embodiment of the present invention.
[0032] Figure 3 It is a schematic diagram of the structure of a system for generating a text summary based on a chart according to another embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0034] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0035] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may include that the first and second features are in direct contact, or may include that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, a first feature being "above", "above" and "above" a second feature includes that the first feature is directly above and obliquely above the second feature, or simply indicates that the first feature is higher in level than the second feature. A first feature being "below", "below" and "below" a second feature includes that the first feature is directly below and obliquely below the second feature, or simply indicates that the first feature is lower in level than the second feature.
[0036] Unless otherwise defined, the technical or scientific terms used herein shall have the common meanings understood by persons with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the patent application specification and claims of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one" or "a" do not indicate a quantity limitation, but indicate the existence of at least one.
[0037] In the actual production and operation process, enterprises will continuously generate a wide variety of data, including but not limited to financial data used to evaluate financial status, conduct financial planning and forecasting, and set reasonable financial goals; market data that helps enterprises understand market demand and trends and provide a basis for product development, market positioning and marketing; customer data that is indispensable in enterprise management, etc. In order to more effectively understand and utilize this data, enterprises often use various chart forms such as trend charts, scatter charts, pie charts, etc., and carefully draw these data on a chart so that the company's decision-makers can clearly interpret and analyze them. However, although charts can present data intuitively, human nature is limited in its ability to quickly interpret complex information. Faced with these charts full of data and information, decision makers often need to spend a long time and energy to understand and analyze, which will undoubtedly affect the efficiency and accuracy of decision-making. Therefore, how to more effectively present and utilize this data and improve decision-making efficiency has become an important issue facing enterprises. In this embodiment, a method for generating a text summary based on a chart is disclosed, such as Figure 1 As shown, the method may specifically include the following steps.
[0038] Step S1, obtaining a public data set on the Internet as a training sample, wherein the public data set includes high-quality charts and texts corresponding to the charts.
[0039] Specifically, this patent uses a chart summary generation model to generate a text summary interpretation of the chart by inputting a chart, so as to help decision makers effectively use chart data to make decisions. In essence, this problem can be regarded as a standard image captioning problem, which embodies the idea of multimodal processing, that is, using a variety of different data modalities including text and images to train the model to achieve accurate mapping from graphic modality to text modality. The model can automatically output a text description describing the image content and scene by receiving image input, thereby realizing automatic conversion between image and text, greatly improving the understanding and processing efficiency of graphic data.
[0040] The sample data set for model training can be constructed by selecting the public data set SCICAP on the Internet as the source of training samples. This data set consists of large-scale graphics and texts intercepted from papers published on arXiv between 2010 and 2020 in the computer field. The graphics are all high-quality graphics in these papers, ensuring the professionalism and accuracy of the graphics. Figure 2 As shown in the figure, we use the high-quality charts as the input samples X of the chart summarization generation model, and use the text corresponding to each high-quality chart as the output label Y of the chart summarization generation model. The content of the output label Y of the figure is: "The average number of iterations required for decoding the PLDPC-Hadamard code when the parameters r=8 and k=204,800 are given". It is ensured that the chart summarization generation model learns the correspondence between charts and texts during the training process, so that it can more accurately understand and generate text descriptions related to charts in the future, laying a solid foundation for subsequent image and text processing and analysis tasks.
[0041] Step S2, taking a high-quality chart as an input sample and the text summary information corresponding to the chart as an output label, constructing a training data set, and using the training data set to train the original chart summary generation model, wherein the chart summary generation model is composed of an encoder of a variational autoencoder model VAE, a decoder of a variational autoencoder model VAE, a text summary generation model established using a masked decoder of a Transformer deep learning model, and a loss function layer, wherein the loss function layer includes a first loss function for data with text summary and a second loss function for data without text summary.
[0042] Specifically, Transformer is a deep learning model based on the attention mechanism, which mainly includes two parts, an encoder and a decoder. It was originally designed for natural language processing tasks and has achieved remarkable results in machine translation tasks. Compared with traditional RNN or CNN, the Transformer deep learning model relies entirely on the self-attention mechanism to capture the dependency between input and output, enabling it to efficiently process sequence data in parallel, thereby significantly improving computational efficiency. In the current field of image-to-text conversion, many models with excellent performance mostly use the Transformer framework to process image data. However, despite the excellent performance of Transformer, its huge parameter scale has brought speed and cost problems in practical applications. This embodiment proposes a method based on the variational autoencoder model VAE for this point, and still retains the Transformer deep learning model when generating the corresponding text summary. VAE is a generative model based on deep learning, which also consists of two parts, an encoder and a decoder. Its main idea is to enable the model to generate new data similar to the training data by probabilistically modeling the implicit representation of the data and maximizing the lower bound of the likelihood function. Therefore, by combining the Transformer deep learning model and the VAE model, compared to simply using the Transformer deep learning model to process the graph, the complexity of the overall model parameters is effectively reduced, and the processing speed of the model is effectively improved without significantly reducing the model performance.
[0043] Specifically, the encoder of the variational autoencoder model VAE is configured as
[0044] ,in Indicates is the mean value, is the variance matrix of the multivariate normal distribution, , X is the input graph, which is an RGB graph and a 3D tensor, F1 and F2 are the parameters adjusted by the encoder learning, and pool represents the pool transformation in the neural network CNN. For is the real number vector output by the multi-layer perceptron MLP, For is the output positive vector of the last layer of the multi-layer perceptron MLP, Table encoder parameters.
[0045] The decoder of the variational autoencoder model VAE is configured as
[0046] ,in Indicates is the mean, is the multivariate normal distribution with the variance matrix, Adjust parameters for subsequent learning; , , S is the code generated by the encoder, D1 and D2 are the parameters learned and adjusted by the decoder, represents the convolution of two 3D tensors, unpool represents the unpool transformation in the neural network CNN, and I is the identity matrix.
[0047] The text summarization generation model built using the Transformer deep learning model with masked decoder is configured as
[0048] , where Y is the text summary information corresponding to the input chart X, and the length of the summary text is T; and Is The chain rule is used to decompose and calculate the derivatives of the composite function; ψ is the parameter of the decoder with mask in the Transformer deep learning model, and s represents the encoding of the input graph X after passing through the encoder. is the independent word or phrase in the text summary information corresponding to the graph X, y=[y1,...,y t ] is the text summary corresponding to chart X and the length of the summary text is T.
[0049] In this embodiment, the graph summarization generation model further includes a loss function layer, wherein the loss function layer includes a first loss function for textual summary data and a second loss function for non-textual summary data.
[0050] Specifically, since the generative model of the assumed text summary Y depends on the encoding s, even if some graphs X do not have clear text summary information Y corresponding to them during the training stage, the model can still indirectly improve its ability to generate text summaries by learning how to encode these graphs more effectively. Therefore, our training framework combines both data with and without text summary data, and still includes graphs X that do not directly correspond to text summary Y into the training process to optimize the encoding ability of the model, so that it can learn more accurate encoding s and latent space representation z, so that the model performs better when processing data with text summary, and provides potential text generation capabilities for data that originally lack text descriptions.
[0051] Specifically, the first loss function for textual summary data is:
[0052] .
[0053] The second loss function for summarizing data without text is:
[0054] ,in For Find the expectation of the distribution with respect to s, For Find the expectation of the distribution with respect to s and z, A setting value between 0 and 1 to stabilize VAE training. is the joint distribution probability function of the three variables X, s, and z; represents the code generation probability function of the encoder, For distribution Conditional distribution function of s obtained by integrating z, where X is the original graph, s is the final encoding of the encoder, and z is a random vector following a multinomial distribution.
[0055] Then the loss function layer can be comprehensively expressed as:
[0056] , where D c is a set of chart data with corresponding text summaries, D u A collection of chart data without a corresponding text summary.
[0057] Furthermore, for the gradient solution of the loss function, the first loss function is used as an example to illustrate the other loss functions ( ) The gradient solution is similar. After a series of mathematical transformations, equivalent conversions are performed, and each term is solved separately through direct solution and reparameterization techniques:
[0058] ,in To transform the first loss function, The second loss function is, is a collection of chart data with corresponding text summaries, A collection of chart data without corresponding textual summary. is the KL distance between two distributions.
[0059] Specifically, the first KL divergence term can be solved in an explicit form, and its gradient is also an explicit result. The gradient of the second term can be solved using the reparameterization technique commonly used in VAE. The specific mathematical expression is long and is omitted here. After solving all the expressions in the first loss function, this expression can be used as the loss function. Since some terms in the loss function (such as the reconstruction error term) involve neural networks, these parameters are learned through backpropagation. Once the model converges, we can then use the model to generate new text summaries. The overall chart summary generation model can be expressed as:
[0060] ,in .
[0061] Step S3, using a PDF conversion tool to identify and extract each target chart in the PDF document to be analyzed, inputting the target chart into the trained chart summary generation model to output text summary information corresponding to each target chart.
[0062] Specifically, after the model training is completed, the corresponding text summary can be obtained by inputting the chart data. Usually, the model can be combined with open source tools such as pdfplumber and PDF-Extract-Kit to extract charts from PDF to help decision makers quickly understand the meaning of the chart. Among them, since the text in the chart X in the data set selected in this embodiment is English, the corresponding text summary Y can be selected in English or translated Chinese as needed, so the model can be applied to both Chinese and English text summaries.
[0063] The method for generating a text summary based on a chart disclosed in this embodiment uses a public data set on the Internet as a training sample to train a chart summary generation model including an encoder of a variational autoencoder model VAE, a decoder of a variational autoencoder model VAE, a text summary generation model established with a masked decoder using a Transformer deep learning model, and a loss function layer, and uses high-quality charts in the training samples as input samples X of the chart summary generation model to be trained, and the text summary corresponding to the chart is used as the output label Y of the chart summary generation model, and the original chart summary generation model is trained, so that after the training is completed, the model can convert each target chart in the input PDF document extracted by the PDF conversion tool into corresponding text summary information, so as to help decision makers quickly obtain the information content in the chart, and improve the efficiency of decision makers in information acquisition and the accuracy of decision making.
[0064] In another embodiment, if Figure 3As shown, a system for generating text summaries based on charts is also disclosed, including a data acquisition module 1, a model training module 2 and a summary generation module 3. The data acquisition module 1 is used to obtain a public data set on the Internet as model training data, wherein the public data set includes high-quality charts and text summary information corresponding to the charts. The model training module 2 is used to take the high-quality charts as input samples and the text summary information corresponding to the charts as output labels, to construct a training data set, and to use the training data set to train the original chart summary generation model, wherein the chart summary generation model is composed of an encoder of a variational autoencoder model VAE, a decoder of a variational autoencoder model VAE, a text summary generation model established using a masked decoder of a Transformer deep learning model, and a loss function layer, wherein the loss function layer includes a first loss function for data with text summary and a second loss function for data without text summary. The summary generation module 3 is used to use a PDF conversion tool to identify and extract each target chart in the PDF document to be analyzed, and input the target chart into the chart summary generation model that has been trained to output the text summary information corresponding to each target chart.
[0065] In this embodiment, the text summary generation model is , where Y is the text summary information corresponding to the input chart X, and the length of the summary text is T; and Is The chain rule is used to decompose and calculate the derivatives of the composite function; ψ is the parameter of the decoder with mask in the Transformer deep learning model, and s represents the encoding of the input graph X after passing through the encoder. is the independent word or phrase in the text summary information corresponding to the graph X, y=[y1,...,y t ] is the text summary corresponding to Chart X and the length of the summary text is T. ] is the text summary corresponding to Chart X and the length of the summary text is T.
[0066] In this embodiment, the encoder of the variational autoencoder model VAE is , Indicates is the mean value, is the multivariate normal distribution with a variance matrix, where , X is the input chart, which is an RGB chart and a 3D tensor, F1 and F2 are the parameters adjusted by the encoder learning, and pool represents the pool transformation in the neural network CNN. For is the real number vector output by the multi-layer perceptron MLP, For is the output positive vector of the last layer of the multi-layer perceptron MLP, is the encoder parameter;
[0067] The decoder of the variational autoencoder model VAE is ,in Indicates is the mean, is the variance matrix of the multivariate normal distribution, Adjust parameters for subsequent learning; , , S is the code generated by the encoder, D1 and D2 are the parameters learned and adjusted by the decoder, represents the convolution of two 3D tensors, unpool represents the unpool transformation in the neural network CNN, and I is the identity matrix.
[0068] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and similar parts between the various embodiments can be referred to each other. For the system for generating a text summary based on a chart disclosed in the embodiment, since it corresponds to the method for generating a text summary based on a chart disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the aforementioned method part.
[0069] In other embodiments, a device for generating a text summary based on a chart is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the various steps of the method for generating a text summary based on a chart as described in the above embodiments when executing the computer program. The server may include, but is not limited to, a processor and a memory. It can be understood by those skilled in the art that the schematic diagram is only an example of a server and does not constitute a limitation on the server, and may include more or fewer components than shown in the diagram, or combine certain components, or different components.
[0070] If the device for generating a text summary according to a chart is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiments for generating a text summary according to a chart can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc.
[0071] In conclusion, the above is only a preferred embodiment of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the patent of the present invention.
Claims
1. A method for generating a text summary based on a chart, characterized in that: The steps include: Obtaining a public data set on the Internet as model training data, wherein the public data set includes high-quality charts and text summary information corresponding to the charts; A high-quality chart is used as an input sample, and the text summary information corresponding to the chart is used as an output label to construct a training data set, and the training data set is used to train the original chart summary generation model, wherein the chart summary generation model is composed of an encoder of a variational autoencoder model VAE, a decoder of the variational autoencoder model VAE, a text summary generation model established using a masked decoder of a Transformer deep learning model, and a loss function layer, wherein the loss function layer includes a first loss function for data with text summary and a second loss function for data without text summary; the text summary generation model is , where Y is the text summary information corresponding to the input chart X, and the length of the summary text is T; and Is The chain rule is used to decompose and calculate the derivatives of the composite function; ψ is the parameter of the decoder with mask in the Transformer deep learning model, and s represents the encoding of the input graph X after passing through the encoder. is the independent word or phrase in the text summary information corresponding to the graph X, y=[y1,...,y t ] is the text summary corresponding to the graph X and the length of the summary text is T. is the text summary corresponding to the graph X and the length of the summary text is T. The encoder of the variational autoencoder model VAE is , Indicates is the mean value, is the multivariate normal distribution with the variance matrix, , X is the input graph, which is an RGB graph and a 3D tensor, F1 and F2 are the parameters adjusted by the encoder learning, and pool represents the pool transformation in the neural network CNN. For is the real number vector output by the multi-layer perceptron MLP, For is the output positive vector of the last layer of the multi-layer perceptron MLP, is the encoder parameter; the decoder of the variational autoencoder model VAE is ,in Indicates is the mean, is the variance matrix of the multivariate normal distribution, Adjust parameters for subsequent learning; , , S is the code generated by the encoder, D1 and D2 are the parameters learned and adjusted by the decoder, represents the convolution of two 3D tensors, unpool represents the unpool transformation in the neural network CNN, and I is the unit matrix; A PDF conversion tool is used to identify and extract each target chart in the PDF document to be analyzed, and the target chart is input into the trained chart summary generation model to output the text summary information corresponding to each target chart.
2. The method for generating a text summary based on a chart according to claim 1, characterized in that: The first loss function for data with text summarization is configured as: ; The second loss function for the text-free summarization data is configured as: ; in For Find the expected value of the distribution s, For Find the expected value of s and z for the distribution, A setting value between 0 and 1 that is used to stabilize the training of the variational autoencoder model VAE. is the joint distribution probability function of the three variables X, s, and z; represents the code generation probability function of the encoder, For distribution The conditional distribution function of the encoder's final encoding s obtained by integrating z, where z is a random vector that obeys a multinomial distribution.
3. The method for generating a text summary based on a chart according to claim 2, characterized in that: The loss function layer is ,in is a training data set of graphs with corresponding text summaries, A training dataset for graphs without corresponding text summaries.
4. A system for generating a text summary based on a chart, characterized in that: include: A data acquisition module, used to acquire a public data set on the Internet as model training data, wherein the public data set includes high-quality charts and text summary information corresponding to the charts; The model training module is used to take a high-quality chart as an input sample and the text summary information corresponding to the chart as an output label, construct a training data set, and use the training data set to train the original chart summary generation model, wherein the chart summary generation model is composed of an encoder of a variational autoencoder model VAE, a decoder of the variational autoencoder model VAE, a text summary generation model established using a masked decoder of a Transformer deep learning model, and a loss function layer, wherein the loss function layer includes a first loss function for data with text summary and a second loss function for data without text summary; the text summary generation model is , where Y is the text summary information corresponding to the input chart X, and the length of the summary text is T; and Is The chain rule is used to decompose and calculate the derivatives of the composite function; ψ is the parameter of the decoder with mask in the Transformer deep learning model, and s represents the encoding of the input graph X after passing through the encoder. is the independent word or phrase in the text summary information corresponding to the graph X, y=[y1,...,y t ] is the text summary corresponding to the graph X, and the length of the summary text is T. is the text summary corresponding to the graph X, and the length of the summary text is T. The encoder of the variational autoencoder model VAE is , Indicates is the mean value, is the multivariate normal distribution with the variance matrix, , X is the input graph, which is an RGB graph and a 3D tensor, F1 and F2 are the parameters adjusted by the encoder learning, and pool represents the pool transformation in the neural network CNN. For is the real number vector output by the multi-layer perceptron MLP, For is the output positive vector of the last layer of the multi-layer perceptron MLP, is the encoder parameter; The decoder of the variational autoencoder model VAE is ,in Indicates is the mean, is the variance matrix of the multivariate normal distribution, Adjust parameters for subsequent learning; , , S is the code generated by the encoder, D1 and D2 are the parameters learned and adjusted by the decoder, represents the convolution of two 3D tensors, unpool represents the unpool transformation in the neural network CNN, and I is the unit matrix; The summary generation module is used to use a PDF conversion tool to identify and extract each target chart in the PDF document to be analyzed, input the target chart into the trained chart summary generation model, and output the text summary information corresponding to each target chart.
5. A device for generating a text summary based on a chart, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.
6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Chart abstract generation method and generation model training method and device
CN115309888A