Method and system for realizing diagram visualization based on natural language processing

By loading the initial text paragraphs into the adjusted natural language processing network, generating recognition results and performing chart visualization, the problem of identifying text visual elements and utilizing text paragraph relationships in the prior art is solved, and efficient and accurate chart visualization is achieved.

CN119293230BActive Publication Date: 2025-05-16JIERUAN TECH (GRP) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411141798.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2025-05-16
Estimated Expiration
2044-08-20

AI Technical Summary

Technical Problem

When the prior art applies natural language processing technology to the field of chart visualization, it is difficult to accurately identify visual elements in text, ignore the timing or logical relationship between text paragraphs, and has high computational complexity and low processing efficiency.

Method used

By obtaining the initial text paragraph collection, loading it on a pre-tuned natural language processing network, generating a collection of recognition results, determining the recognition results of each text paragraph, and visually generating the chart based on these results. The network is tuned through text learning examples, using template implicit representation and mapping data types to improve the accuracy of visual type recognition.

Benefits of technology

It improves the accuracy of visual type recognition of target visual text, enhances the correlation between text paragraphs, reduces the computational complexity of the model, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293230B_ABST
    Figure CN119293230B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for realizing chart visualization based on natural language processing, which analyzes input text based on a natural language processing model to obtain key visualization types, and generates charts matching the visualization types. The target natural language processing network is a natural language processing network obtained by loading text learning samples into an initial natural language processing network to be calibrated, and the text learning samples include multiple template implicit representations extracted from multiple text paragraph training templates and mapping data types and feedback signals corresponding to each template implicit representation. The implicit representation of the preceding text in the text paragraph set is used, and the network is calibrated by fusion and feedback of the implicit representation of the preceding text and the implicit representation of the current text, thereby improving the accuracy of recognition of the visualization type of the target visual text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing, and more specifically, to a method and system for realizing chart visualization based on natural language processing. Background Art

[0002] With the rapid development of information technology and the advent of the big data era, the processing and analysis of text information plays an increasingly important role in all walks of life. Especially in the fields of finance, scientific research, and news, how to quickly extract key information from massive text data and present it to users in an intuitive and easy-to-understand way has become a major challenge for current information processing technology. Traditional text analysis methods often rely on manual feature extraction and rule matching, which is not only inefficient but also difficult to cope with complex and changeable text data. In recent years, the rapid development of natural language processing (NLP) technology has provided new ideas for the automated processing of text information. By training deep learning models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), or transformers, computer systems can automatically learn the inherent laws of text data and realize the conversion from text to structured information. However, when applying NLP technology to the field of chart visualization, there are still many challenges. For example, how to accurately identify the visualization elements in the text and map them to the appropriate chart type; how to effectively use the temporal or logical relationship between text paragraphs to improve the accuracy and relevance of chart visualization; and how to reduce the computational complexity of the model and improve processing efficiency while ensuring accuracy. Although there are some NLP-based chart visualization solutions on the market, most of these solutions focus on simple text-to-chart mapping, ignoring the intrinsic connections between text paragraphs and the importance of contextual information. Summary of the invention

[0003] The purpose of the present invention is to provide a method and system for realizing chart visualization based on natural language processing. The embodiment of the present application is implemented as follows:

[0004] In a first aspect, the present application provides a method for realizing chart visualization based on natural language processing, comprising: obtaining an initial text paragraph set to be processed, wherein the initial text paragraph set includes target visual text to be recognized; loading the initial text paragraph set into a target natural language processing network that has been calibrated in advance to obtain a recognition result set, wherein the target natural language processing network is a natural language processing network obtained by calibrating a text learning sample loaded into an initial natural language processing network to be calibrated, and the text learning sample includes a plurality of template implicit representations extracted from a plurality of text paragraph training templates, and a mapping data type and a feedback signal corresponding to each template implicit representation. The mapping data type is used to indicate a visualization type implicitly represented by a template, the feedback signal is used to indicate a matching degree corresponding to the visualization type, the multiple text paragraph training templates are multiple continuous text paragraphs, and the implicit representation of each template includes an implicit representation obtained by taking a text paragraph training template and a previous text paragraph training template of the text paragraph training template; determining a recognition result corresponding to each text paragraph in the initial text paragraph set according to the recognition result set, wherein the recognition result represents the visualization type of the target visual text; and generating a chart visualization according to the recognition result corresponding to each text paragraph to obtain a text visualization chart.

[0005] In a second aspect, the present application provides a computer system comprising: one or more processors; a memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method as described above is implemented.

[0006] Beneficial effects of the present application: The present application obtains an initial text paragraph set to be processed, wherein the initial text paragraph set includes target visual text to be recognized, and loads the initial text paragraph set into a target natural language processing network that has been calibrated in advance to obtain a recognition result set. The target natural language processing network is a natural language processing network obtained by loading text learning samples into the initial natural language processing network to be calibrated. The text learning samples include multiple template implicit representations extracted from multiple text paragraph training templates and a mapping data type and a feedback signal corresponding to each template implicit representation. The mapping data type is used to indicate the visualization type of a template implicit representation, and the feedback signal is used to indicate the degree of matching corresponding to the visualization type. The multiple text paragraph training templates are multiple continuous text paragraphs. Each template implicit representation includes an implicit representation obtained by taking a text paragraph training template and a previous text paragraph training template of a text paragraph training template. The recognition result corresponding to each text paragraph in the initial text paragraph set is determined based on the recognition result set, wherein the recognition result represents the visualization type of the target visual text, and the implicit representation of the preceding text in the text paragraph set is adopted. The network is calibrated by fusing and feeding back the implicit representation of the preceding text and the implicit representation of the current text, thereby improving the accuracy of the recognition of the visualization type of the target visual text. Furthermore, classification is performed based on the aforementioned implicit representation of the text and the implicit representation of the current text, and the correlation between text paragraphs helps improve the accuracy of the network without the need for additional data and links, and requires low computing power consumption.

[0007] Other features will be described in part in the following description. Those skilled in the art will discover these features in part upon inspection of the following content and drawings, or may learn these features through production or use. The features of the present application may be implemented and obtained by practicing or using various aspects of the methods, tools, and combinations listed in the detailed examples described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.

[0009] Figure 1 It is a flowchart of a method for realizing chart visualization based on natural language processing provided in an embodiment of the present application.

[0010] Figure 2 It is a schematic diagram of the functional module architecture of the visualization device provided in an embodiment of the present application.

[0011] Figure 3 It is a schematic diagram of the composition of a computer system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method part of the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0013] The execution subject of the method for realizing the visualization of a chart based on natural language processing in the embodiment of the present application is a computer system, including but not limited to a server, a personal computer, a laptop, a tablet computer, a smart phone, etc. The computer system includes user devices and network devices. Among them, user devices include but are not limited to computers, smart phones, PADs, etc.; network devices include but are not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers in cloud computing, wherein cloud computing is a kind of distributed computing, a super virtual computer composed of a group of loosely coupled computer sets. Among them, the computer system can be run alone to implement this application, or it can be connected to the network and implement this application through interactive operations with other computer systems in the network. Among them, the network where the computer system is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, VPN network, etc.

[0014] The embodiment of the present application provides a method for realizing chart visualization based on natural language processing, which is applied to a server, such as Figure 1 As shown, the method includes:

[0015] Step S100: obtaining an initial text paragraph set to be processed, wherein the initial text paragraph set includes target visible text to be recognized.

[0016] In step S100, the computer system retrieves and collects a series of intrinsically related text paragraphs from a predetermined data source. These text paragraphs can be in various forms, but there must be some kind of logical or temporal continuity between them so that useful information that can be used for chart visualization can be extracted from them. For example, suppose a financial news website aims to automatically extract key financial indicators (such as revenue, profit, assets, etc.) from financial reports about a listed company's quarterly financial report, and generate corresponding charts so that readers can more intuitively understand the changing trends of these data. In this scenario, the execution process of step S100 can be as follows: the computer system is connected to a news database that stores a large number of financial report articles. Then, all financial report articles related to the listed company are retrieved according to specific query conditions (such as the company name and the keyword "financial report" contained in the article title). Since these articles are published in chronological order, each article contains the company's financial report data for different quarters, so these articles naturally form an ordered set of text paragraphs. Next, the computer system preprocesses these articles and splits them into separate text paragraphs. In this process, the computer system can identify structural elements such as subheadings and paragraph marks in the article to ensure the accuracy of the splitting. For those paragraphs containing key financial indicators (i.e., target visual text), the computer system specially marks them for focused processing in subsequent steps.

[0017] For example, a story might contain the following paragraphs (simplified):

[0018] "In the first quarter, the company's total revenue reached 1 billion yuan, a year-on-year increase of 5%."

[0019] "In the second quarter, despite the challenges of intensified market competition, the company's total revenue still grew to 1.2 billion yuan, with the year-on-year growth rate slowing to 3%."

[0020] These two paragraphs describe the company's financial performance for two consecutive quarters. There is an obvious chronological order and logical relationship between them, which is part of the ideal initial text paragraph set. The system will collect such paragraphs to form a serialized set for subsequent natural language processing and recognition of target visual text. Through this step, the computer system not only obtains the necessary text data, but also ensures the inherent relevance and orderliness between these data, laying a solid foundation for the subsequent chart visualization.

[0021] Step S200: Load the initial text paragraph set into a target natural language processing network that has been calibrated in advance to obtain a recognition result set, wherein the target natural language processing network is a natural language processing network obtained by loading text learning samples into an initial natural language processing network to be calibrated, and the text learning samples include multiple template implicit representations extracted from multiple text paragraph training templates, and a mapping data type and a feedback signal corresponding to each template implicit representation, the mapping data type is used to indicate a visualization type of a template implicit representation, and the feedback signal is used to indicate a matching degree corresponding to the visualization type, the multiple text paragraph training templates are multiple continuous text paragraphs, and each template implicit representation includes an implicit representation obtained by taking a text paragraph training template and a previous text paragraph training template of the text paragraph training template.

[0022] In step S200, the target natural language processing network is, for example, a complex neural network model, such as a recurrent neural network (RNN), a long short-term memory network (LSTM) or a Transformer, which are good at processing sequence data, such as text paragraphs. During the construction process, developers will use a large number of text learning samples to train the network. The text learning samples consist of multiple text paragraph training templates, which are continuous text paragraphs that simulate text sequences that may appear in practical applications. For each template, the computer system extracts its implicit representation, which not only contains the information of the current text paragraph, but also integrates the information of the previous text paragraph that is closely related to it. For example, if the text paragraph is a description of a company's quarterly financial report, the implicit representation will take into account the financial report data of the current quarter and the previous quarter, as well as the relationship between them.

[0023] Specific methods for extracting implicit representations include, for example, word embedding, sentence embedding, or more advanced text representation learning techniques. Subsequently, the computer system assigns a mapping data type to each implicit representation, which is actually a classification label indicating the most suitable visualization type (such as a bar chart, line chart, etc.) for the implicit representation. At the same time, the computer system also generates a feedback signal based on the degree of match between the network's prediction results and the actual mapping data type. This signal is used to guide the optimization direction of the network during the training process.

[0024] In step S200, the computer system inputs the initial text paragraph set obtained in step S100 into the calibrated target natural language processing network. This process is similar to inputting an unlabeled data set into a trained model to obtain prediction results.

[0025] Suppose an application scenario is still the visualization of financial report data in financial news. The initial text paragraph set may contain multiple financial report reports about different companies and different time points. When these text paragraphs are input into the target network, the network will use the knowledge it has learned during the training process to parse each paragraph of text and generate the corresponding implicit representation. Based on the implicit representation of each paragraph of text, the target natural language processing network will predict the most suitable visualization type for the text paragraph, that is, the mapping data type. This prediction process may involve multiple steps, including feature extraction, application of attention mechanism, and final classification decision. For example, for a text paragraph describing the changes in a company's quarterly revenue, the network can identify the key financial indicators (such as revenue, growth rate, etc.) and predict that the most suitable visualization type is a line chart based on the changing trends of these indicators. The prediction result will be collected into the recognition result set together with the corresponding text paragraph as a recognition result.

[0026] Step S300: determining a recognition result corresponding to each text paragraph in the initial text paragraph set according to the recognition result set, wherein the recognition result represents a visualization type of the target visual text.

[0027] In step S300, the computer system matches the recognition result set generated in step S200 with the initial text paragraph set, and determines the visualization type corresponding to each text paragraph. In step S200, the computer system has processed the initial text paragraph set through the target natural language processing network and generated a recognition result set. Each item in this set corresponds to a text paragraph in the initial text paragraph set, and contains the prediction results of the most suitable visualization type (such as a bar chart, a line chart, etc.) for the text paragraph. In step S300, the computer system traverses the recognition result set. For each item in the set, it extracts the predicted visualization type and associates it with the corresponding text paragraph in the initial text paragraph set. This association is achieved, for example, through an index or identifier to ensure that each text paragraph can accurately match its corresponding visualization type.

[0028] Assume that the initial text paragraph set contains the following two paragraphs (for simplicity, only key financial data is shown here):

[0029] "In the first quarter, XYZ Company's total revenue reached $100 million, a 10% year-over-year increase."

[0030] "By the second quarter, XYZ Company's total revenue climbed to $120 million, with a slightly lower year-over-year growth rate of 8%."

[0031] In step S200, the target natural language processing network may have generated the following set of recognition results for the two paragraphs:

[0032] Paragraph 1-> Line chart (because the data changes over time, line charts are suitable for showing trends);

[0033] Paragraph 2-> Line chart (also shows the trend and is continuous with Paragraph 1);

[0034] In step S300, the computer system associates the two recognition results with corresponding paragraphs in the initial text paragraph set, which means that the computer system already knows that paragraph 1 and paragraph 2 are both suitable for visualizing the financial data therein using a line graph.

[0035] Once each text paragraph is assigned a corresponding visualization type, the computer system can proceed to step S400, that is, generate a chart visualization based on these recognition results. In this process, the computer system can further parse the specific financial data (such as income, growth rate, etc.) in the text paragraph and use this data as the data source of the chart. Then, it selects a suitable chart library (such as D3.js, ECharts, etc.) to generate the chart based on the assigned visualization type (such as a line chart). Although the target natural language processing network can accurately predict the visualization type of the text paragraph in most cases, the prediction results can be biased in some edge cases. In order to improve the robustness and accuracy of the system, some post-processing mechanisms can be integrated into the system, such as manual review, user feedback loops, or automatic adjustment strategies. In addition, as new text types and visualization requirements emerge, the computer system can regularly update its training data and model parameters to adapt to changes.

[0036] Step S300 plays a bridging role in the implementation of chart visualization based on natural language processing. It connects complex natural language processing results with specific chart visualization requirements and provides a clear guiding direction for subsequent chart generation.

[0037] Step S400: generating a visualization chart based on the recognition result corresponding to each text paragraph to obtain a text visualization chart.

[0038] In step S400, the computer system converts the recognition result corresponding to each text paragraph determined in step S300 into an actual chart visualization form, thereby generating a text visualization chart.

[0039] The execution process of step S400 may include the following exemplary steps:

[0040] 1. Data extraction and preprocessing. Before step S400 begins, the computer system has completed the recognition and classification of text paragraphs (step S300), and knows the data type (such as income, profit) and the appropriate visualization type (such as bar chart, line chart) described by each text paragraph. In this step, the computer system further extracts specific data values ​​from the text paragraph, such as "revenue in the first quarter is 10 billion yuan". These data values ​​will be converted into numerical data for subsequent visualization processing.

[0041] 2. Chart template selection. According to the recognition result determined in step S300, the computer system selects a corresponding chart template. For example, if the recognition result is a line chart, the system will select a line chart template suitable for displaying time series changes. These templates usually contain style information such as the layout, color, font, etc. of the chart, as well as a data interface for drawing the chart.

[0042] 3. Data mapping and chart drawing. Next, the computer system maps the extracted data values ​​to the selected chart template. This process involves converting text data (such as specific values ​​of revenue and profit) into points, lines or areas on the chart. For example, for a line chart, each data point will correspond to a coordinate point on the chart, and the lines connecting the data points will form the line of change trend. When drawing a chart, the computer system will also add auxiliary information such as titles, legends, and axis labels to the chart based on the description of the text paragraph to improve the readability of the chart. For example, the title may be "Revenue change trend of XYZ Company in the past three years", and the legend indicates which year's data the lines of different colors represent.

[0043] 4. Chart optimization and adjustment. After the initial drawing of the chart is completed, the computer system can perform a series of optimization and adjustment work to ensure the aesthetics of the chart and the accuracy of information transmission. This includes adjusting the size, color matching, font style and other visual elements of the chart, as well as adjusting the scale and range of the coordinate axis according to the characteristics of the data. In addition, the computer system may also use some intelligent algorithms to automatically adjust the layout and element position of the chart to avoid information overlap or omission.

[0044] 5. Chart output and display. Finally, the computer system outputs the optimized charts in the form of pictures or interactive web pages. These charts can be directly embedded in the report pages of financial news for readers to view and analyze. At the same time, the computer system can also provide downloading, sharing and other functions, so that readers can save or share the charts with others.

[0045] For example, suppose a text paragraph describes the profit changes of a company in the past five years: "From 2017 to 2021, the annual profits of ABC Company were 500 million yuan, 600 million yuan, 700 million yuan, 800 million yuan and 900 million yuan respectively." In step S400, the computer system will first recognize that this text describes time series data (annual profit) and determine that the appropriate visualization type is a line chart. Then, the computer system extracts the profit values ​​for each year (500 million yuan, 600 million yuan...) and maps these values ​​to a line chart template. The resulting line chart will clearly show the growth trend of ABC Company's profits in the past five years, allowing readers to see it at a glance. Step S400 is one of the key links in the implementation of chart visualization based on natural language processing. It uses the recognition results obtained in step S300 and the specific data values ​​in the text paragraph to generate an intuitive chart visualization form, thereby improving the readability and analysis efficiency of the data.

[0046] In one implementation, the method further includes a process of extracting the implicit representations of the multiple templates, which may specifically include the following steps:

[0047] Step S21: performing a first implicit representation mining process on the plurality of text paragraph training templates to obtain a first implicit representation set, wherein an implicit representation in the first implicit representation set corresponds to a text paragraph training template in the plurality of text paragraph training templates;

[0048] Step S22: performing a second implicit representation mining process on the multiple text paragraph training templates to obtain a second implicit representation set, wherein a first implicit representation in the second implicit representation set corresponds to a first text paragraph training template in the text paragraph training templates, a previous paragraph text of the first text paragraph training template is a second text paragraph training template, and the first implicit representation is used to characterize the correlation between the first text paragraph training template and the second text paragraph training template;

[0049] Step S23: performing feature integration on the first implicit representation set and the second implicit representation set respectively to obtain the multiple template implicit representations.

[0050] In step S21, the first implicit representation mining process is involved. Specifically, first, the computer system can convert the text paragraph into a numerical vector. This can be achieved by a pre-trained word embedding model (such as the embedding layer of Word2Vec, GloVe or BERT), in which each word or phrase is mapped to a point (vector) in a high-dimensional space. For example, phrases such as "Apple" and "second quarter revenue" are converted into vectors with similar semantics. Next, the computer system aggregates all word vectors in the paragraph to form a single vector representing the entire paragraph. This can be achieved by simple averaging, weighted averaging (according to TF-IDF weights) or using more complex sequence models (such as LSTM, GRU). For example, if the paragraph is "Apple's second quarter revenue reached 10 billion US dollars", the aggregated vector will contain the semantic information of this complete sentence. Finally, the computer system outputs a first implicit representation set, in which each element is an implicit representation vector corresponding to the text paragraph training template. For example, the set may contain two different vectors representing "Apple's Q2 2022 revenue" and "Microsoft's Q2 2022 profit".

[0051] For example, consider two text paragraph training templates (understandably, this is a minimal example):

[0052] Template 1: "Apple's revenue in the second quarter of 2022 reached $10 billion."

[0053] Template 2: "Microsoft's profit for the first quarter of 2022 was $3 billion."

[0054] After processing in step S21, the computer system generates two implicit representation vectors for the two templates, denoted as v1 and v2, which capture the key financial information in template 1 and template 2 respectively.

[0055] In step S22, the computer system further mines the correlation between the training templates of the text paragraphs, especially the relationship between adjacent templates (i.e., temporally continuous or logically related paragraphs). This correlation is crucial for subsequent visualization type predictions because it can help the system understand how the data changes over time or conditions. In order to capture the correlation between templates, the computer system can adopt a variety of methods. A common method is to use an attention mechanism, which allows the model to focus on and fuse the implicit representation of the previous template when generating the implicit representation of the current template. Another method is to build a sequence model (such as LSTM, Transformer), which is naturally suitable for processing sequence data and can implicitly include information from the previous step in the output of each step. In addition to directly fusing the implicit representations of the previous and next templates, the computer system can also calculate the difference between them (such as the change in vector difference or cosine similarity) to generate a second implicit representation that specifically represents the amount of change. This representation is particularly useful for predicting trend charts (such as line charts). Finally, the computer system outputs a second implicit representation set, each element of which is an implicit representation corresponding to a specific text paragraph training template (and its previous template), which emphasizes the correlation between the templates.

[0056] For example, continuing with the previous example, suppose template 1 is "Apple's revenue in the second quarter of 2022 reached 10 billion U.S. dollars", and template 3 follows template 1, which is "Apple's revenue in the third quarter of 2022 increased to 11 billion U.S. dollars". In step S22, the computer system generates a second implicit representation that not only contains the information of template 3 itself, but also incorporates the information of template 1 to highlight the growth of revenue. This second implicit representation may be a composite vector combining the implicit representations of v1 and template 3, or a measure of the difference between them.

[0057] In step S23, the computer system integrates the first implicit representation and the second implicit representation to generate a final template implicit representation. This step is key to extracting comprehensive and representative text features, which ensures that the generated graph reflects both the content of a single text paragraph and the relationships and trends between paragraphs.

[0058] Specifically, splicing and fusion can be used for integration. A simple method is to directly splice the first implicit representation and the second implicit representation into a longer vector. This method may increase the dimension of the vector, making subsequent processing complicated. In other embodiments, a fusion layer (such as a fully connected layer, a gating mechanism, or an attention layer) can be used to map the two representations into a common space and generate a new vector that combines the information of both. During the fusion process, the computer system can assign different weights to different features. This can be obtained through learning, for example, using a back-propagation algorithm to adjust the weight parameters during training. The size of the weight reflects the importance of different features to the final visualization type prediction.

[0059] In order to maintain numerical stability and avoid gradient vanishing or exploding problems, the computer system can normalize the fused vector (such as L2 normalization). For example, returning to the previous example, assuming that after steps S21 and S22, the first implicit representation v1 of template 1 and the second implicit representation r of the relationship between template 1 and template 3 are obtained. 1,3 In step S23, the computer system may combine the two representations using a fusion layer. This fusion layer may be a simple linear combination (weighted sum) or a more complex neural network layer that learns how to combine v1 and r 1,3 Mapped to a new vector f of possibly different dimensions 1,3 This new vector f 1,3 It is the final template implicit representation corresponding to template 1 and template 3, which will be used for subsequent visualization type prediction.

[0060] Through the above steps S21, S22 and S23, the computer system can effectively extract text paragraphs containing key financial indicators and their changing trends from financial reports and convert them into machine-understandable implicit representations. These implicit representations not only capture the content of a single text paragraph, but also enhance the coherence of the context by modeling the relationship between paragraphs. These representations will become the basis for subsequent visualization type prediction and chart generation, helping the financial news analysis platform to provide more intuitive and useful data visualization services.

[0061] In one implementation, the step S22, performing a second implicit representation mining process on the plurality of text paragraph training templates to obtain a second implicit representation set, may specifically include:

[0062] Step S221: performing a second implicit representation mining process on the plurality of text paragraph training templates based on the following operations to obtain a second implicit representation set, wherein each time a text paragraph training template subjected to the second implicit representation mining process is treated as a current text paragraph training template, the obtained implicit representation is treated as a current implicit representation, and the second implicit representation set includes the current implicit representation;

[0063] Step S222: performing text detection processing on the current text paragraph training template to determine a first core text character set, wherein the first core text character set is used to indicate the distribution coordinates of the target visible text in the current text paragraph training template;

[0064] Step S223: performing text detection processing on a previous text paragraph training template of the current text paragraph training template to determine a second core text character set, wherein the second core text character set is used to indicate the distribution coordinates of the target visible text in the previous text paragraph training template;

[0065] Step S224: Determine the current implicit representation based on the first core text character set and the second core text character set, wherein the current implicit representation represents changes between corresponding core text characters in the first core text character set and the second core text character set.

[0066] In step S221, the computer system begins to iteratively process multiple text paragraph training templates to generate a second implicit representation set. These templates are continuous text paragraphs, usually arranged in chronological order or logical order, and are used to simulate the financial data update process in financial reports in actual applications. In each iteration, the template currently being processed is called the "current text paragraph training template", and the generated implicit representation is called the "current implicit representation" and is collected in the second implicit representation set. Specifically, the computer system initializes an empty set for storing the second implicit representation (i.e., the second implicit representation set). Then, it traverses all text paragraph training templates, and for each template, executes steps S222 to S224 to generate its corresponding current implicit representation, and adds the representation to the set. This process is automatic and iterative, ensuring that each template can be processed fairly and consistently.

[0067] In step S222, the computer system performs text detection processing on the current text paragraph training template, with the purpose of identifying character sets in the template that are directly related to the target visual text (i.e., key financial indicators). These character sets, referred to as "first core text character sets", usually indicate the numerical value, unit, and possible trend description of the financial indicator, and are an indispensable source of information for generating charts.

[0068] Specifically, text detection processing can involve multiple steps, including but not limited to part-of-speech tagging, named entity recognition (NER), and regular expression matching. For example, for the text "Apple's revenue in the second quarter of 2022 reached 10 billion US dollars", the computer system can extract "10 billion US dollars" as a key financial indicator through regular expression matching of numbers, currency units, and time words. At the same time, part-of-speech tagging and NER technology can help the system identify the financial category word "revenue" to more accurately locate the core text characters. Suppose the current text paragraph training template is "Google's advertising revenue in the first quarter of 2023 increased by 15% year-on-year to 5 billion US dollars." Through text detection processing, the computer system may identify "5 billion US dollars" as a specific financial indicator value, "advertising revenue" as a financial indicator type, and "year-on-year growth of 15%" as a trend description, which together constitute the first core text character set.

[0069] In step S223, the computer system turns its attention to the previous template of the current text paragraph training template (i.e., the previous text paragraph in time or logic), and performs the same text detection processing as step S222 on the template. The purpose of this processing is to determine the "second core text character set", which also contains relevant information on key financial indicators, but reflects the data status at the previous time point or condition. By comparing the two core text character sets, the computer system can capture the changes in financial indicators.

[0070] Specifically, similar to step S222, the computer system uses the same text detection processing technology (such as part-of-speech tagging, NER and regular expression matching) to extract the key financial indicator information in the previous text paragraph training template. However, unlike step S222, the focus here is on the previous point in the time series or the previous link in the logical sequence.

[0071] Continuing with the previous example, if the current text paragraph training template is about "Google's first quarter of 2023", then the previous text paragraph training template may be about "Google's fourth quarter of 2022". The system will perform text detection processing on the template "Google's advertising revenue reached 4.5 billion US dollars in the fourth quarter of 2022", extract "4.5 billion US dollars" as the specific financial indicator value, and together with "advertising revenue" constitute the second core text character set.

[0072] In step S224, the computer system generates a current implicit representation using the first core text character set and the second core text character set. This representation not only includes the key financial indicator information in the current text paragraph training template, but also captures the change of the indicator by comparing the two core text character sets. This change is very important information in chart visualization because it can intuitively show the change trend of data over time or conditions.

[0073] Specifically, in order to generate the current implicit representation, the computer system may adopt a variety of methods. A simple method is to directly encode the difference (such as numerical difference) between the two core text character sets as part of the vector. However, a more effective method may be to use a machine learning model (such as a neural network) to automatically learn the representation of this difference. For example, a computer system can train a sequence model (such as LSTM or Transformer) that can receive two core text character sets as input and output an implicit representation vector that combines the information of both and can characterize the changes.

[0074] Another possible approach is to use feature engineering to manually construct a set of features that can reflect the changes, and then input these features into a classifier or regressor to predict the most suitable visualization type or the parameters required to generate a chart. However, in practical applications, due to the complexity and diversity of financial data, this approach is often not as flexible and effective as end-to-end machine learning methods.

[0075] Assume that we have obtained the first core text character set {100, USD, revenue} and the second core text character set {90, USD, revenue} (here, for simplicity, financial indicators are directly represented by numerical values, which may be more complex text or structured data in practice). In order to generate the current implicit representation, we can calculate the difference between the two sets and encode this difference as part of the vector. However, a more common and effective method is to use a machine learning model.

[0076] If you choose to use a neural network model (such as LSTM), the processing flow of the model is as follows: First, convert the two core text character sets into a format that the model can handle. This usually involves converting the text into a sequence of word embedding vectors and possibly adding some additional features (such as timestamps, company IDs, etc.). Input the processed input sequence into the LSTM model. LSTM units are able to capture long-term dependencies in the sequence and generate a hidden state vector that contains the key information in the sequence. The current implicit representation vector is extracted from the output of the last time step of the LSTM model. This vector not only contains the information in the current text paragraph training template, but also fuses the information of the previous template through the internal state mechanism of LSTM, so that it can characterize the changes in financial indicators.

[0077] In practical applications, this implicit representation vector can be used as one of the input features for subsequent steps (such as visualization type prediction or chart generation). By combining other types of features (such as the overall semantic representation of text, timestamp features, etc.), the computer system can more accurately predict the most suitable visualization type and generate high-quality charts to meet user needs.

[0078] In summary, step S22 effectively captures the key financial indicators and their changes in the text data by iteratively processing the text paragraph training template, determining the core text character set, and generating implicit representations based on these sets. These implicit representations not only provide a rich source of information for subsequent chart visualization, but also realize the automatic learning and representation of complex data relationships through the powerful capabilities of machine learning models.

[0079] In one implementation, the method may further include:

[0080] Determine the mapping data type and the feedback signal corresponding to each template implicit representation in the multiple template implicit representations based on the following operations, wherein the template implicit representation for determining the mapping data type and the feedback signal each time is regarded as the current template implicit representation, and the mapping data type corresponding to the current template implicit representation is regarded as the current mapping data type and the current feedback signal:

[0081] Step S31: loading the current template implicit representation into the initial natural language processing network to obtain the current mapping data type, wherein the initial natural language processing network is a natural language processing network obtained by performing parameter pre-configuration processing in advance;

[0082] Step S32: Determine the current feedback signal based on the current mapping data type and the template prior tag corresponding to the current template implicit representation, wherein the template prior tag carries the actual type in the text paragraph training template corresponding to the current template implicit representation, the actual type includes the target visualization type of the target visual text, and the feedback signal represents whether the current mapping data type is consistent with the target visualization type.

[0083] In step S31, the main task of the computer system is to input the implicit representation of each template obtained after complex processing (these representations contain the key financial information in the text paragraph and their interrelationships) into an initialized natural language processing network, in order to obtain the corresponding mapping data type from the network. In this scenario, the mapping data type refers to the most suitable visualization chart type for the text paragraph, such as bar charts, line charts, pie charts, etc. These chart types can intuitively display the characteristics and changing trends of financial data.

[0084] Specifically, first, this initial natural language processing network is obtained by adjusting parameters through a large amount of pre-training data. It has the ability to extract key information from text representation and predict its corresponding visualization type. The initialization parameters of the network are obtained, for example, by unsupervised or supervised pre-training on a large-scale corpus to ensure that the network has a certain generalization ability and robustness. When the computer system obtains a new template implicit representation (that is, the current template implicit representation), it passes this representation vector as input to the initial natural language processing network. This vector may contain semantic information of the text paragraph, specific values ​​of financial indicators, and the changing relationship between these values. After receiving the input vector, the network extracts key features through a series of nonlinear transformations (such as convolution, pooling, full connection, etc., depending on the network structure) and finally outputs a prediction result. In this scenario, the prediction result is a classification label, that is, a mapping data type, which indicates the type of visualization chart that is most suitable for the current template implicit representation.

[0085] For example, suppose there is a template implicitly representing a vector v, which represents the text paragraph of "Apple's revenue changes from 2020 to 2022". When this vector is input into the initial natural language processing network, the network can output a label such as "line chart", indicating that the revenue changes during this period are best presented by a line chart. After determining the mapping data type, the task of step S32 is to generate a feedback signal based on this prediction result and the actual visualization type of the template (i.e., the template prior label). This feedback signal is crucial for evaluating the prediction performance of the natural language processing network. It directly reflects the accuracy of the network's prediction and provides an important basis for subsequent network adjustment and optimization.

[0086] Specifically, obtain the template prior mark, which is determined during the construction of the text paragraph training template, and represents the actual visualization type corresponding to each template. These marks are, for example, manually annotated to ensure accuracy and reliability. During the training process, they are used as supervisory signals to guide the learning direction of the network. For each current template implicit representation, the computer system compares its corresponding predicted mapping data type (the output of step S31) with the template prior mark. This comparison process may be a simple string match (if the mapping data type is given in text form) or a more complex similarity calculation (if the mapping data type is given in the form of a vector or probability distribution), and is not specifically limited.

[0087] Based on the comparison results, the computer system generates a feedback signal. If the predicted mapping data type is consistent with the template prior label, the feedback signal is positive (indicating that the prediction is correct); if not, the feedback signal is negative (indicating that the prediction is wrong). In some cases, the feedback signal may also contain information about the degree of prediction error in order to more finely guide the subsequent learning of the network.

[0088] Continuing with the previous example, suppose that the visualization type predicted by the network for the template implicit representation vector v is "line graph", and the template prior label is also "line graph". In this case, the feedback signal generated by the computer system is positive, indicating that the network prediction is correct. However, if the visualization type predicted by the network is "bar graph" and the template prior label is "line graph", the feedback signal is negative. In order to more finely assess the degree of prediction error, the computer system can calculate the similarity or distance between the predicted type and the prior label. For example, if the mapping data type is given in the form of a probability distribution (such as a softmax output), the cross entropy loss function can be used to calculate the difference between the predicted distribution and the true distribution.

[0089] In practical applications, since only a binary feedback signal (positive or negative) is required, as an implementation method, Boolean comparison can also be used for determination.

[0090] Through the execution of the above steps S31 and S32, the computer system can generate corresponding mapping data types and feedback signals for each template implicit representation. These outputs not only provide necessary guidance information for subsequent chart visualization (such as which chart type to choose to display data), but also provide important feedback basis for the continuous optimization and adjustment of the natural language processing network. In practical applications, the accuracy and efficiency of these steps directly affect the performance and user experience of the entire chart visualization system. Therefore, continuously optimizing the implementation of these steps is the key to improving the overall performance of the system.

[0091] In one implementation, the step S31, loading the current template implicit representation into the initial natural language processing network to obtain the current mapping data type, may specifically include:

[0092] Step S311: loading each visualization type in the preset visualization type library together with the current template implicit representation into the type mapping function one by one to obtain a first mapping value set, wherein the first mapping value set includes a mapping value corresponding to each visualization type;

[0093] Step S312: Determine the candidate mapping data type corresponding to the largest first mapping value in the first mapping value set as the current mapping data type.

[0094] In step S311, the computer system combines each visualization type (such as a bar chart, a line chart, a pie chart, etc.) in the preset visualization type library with the current template implicit representation and inputs it into a specific type mapping function. The function of this function is to calculate a score or mapping value based on the input implicit representation and visualization type information, which reflects the matching degree between the current template implicit representation and the given visualization type.

[0095] Specifically, the computer system maintains a preset visualization type library, which contains all possible types for chart visualization. These types are usually stored in the library in some form (such as string, enumeration or index) for easy reference in the algorithm. For example, the visualization type library may contain ["bar chart", "line chart", "pie chart", "scatter chart"], etc. The computer system selects a suitable type mapping function to process the implicit representation and visualization type of the input. For example, it can be assumed that the softmax function is used as an example of a mapping function, although other types of functions (such as linear functions, support vector machines, or neural network classifiers) may be selected in actual applications. The softmax function is particularly suitable for handling multi-classification problems because it can map the input to a probability distribution, where the probability value of each category is between 0 and 1, and the sum of the probabilities of all categories is 1. For each visualization type in the preset visualization type library, the computer system passes it and the implicit representation of the current template as input to the softmax function (or other type mapping function). In this process, the implicit representation may first be converted into a fixed-length vector (if it is not already in vector form), and then combined with some representation of the visualization type (such as a one-hot encoding or embedding vector). The softmax function then outputs a set of mapping values, each element of which corresponds to the score or probability of a visualization type. For example, suppose the implicit representation of the current template is a vector v of length 100, which captures the key information about the changes in company revenue in a financial news paragraph. The preset visualization type library contains three types: bar chart, line chart, and pie chart. For each type, it is converted into a one-hot encoding vector (for example, [1,0,0] for a bar chart, [0,1,0] for a line chart, and [0,0,1] for a pie chart). Then, the implicit representation vector v is concatenated with the one-hot encoding vector of each type to form three new input vectors (each vector is 103 in length because 3 one-hot encoding bits are added). These vectors are input into the softmax function respectively, and the function calculates three mapping values ​​(probabilities) according to the internal weight and bias parameters, such as [0.1, 0.8, 0.1], which means that the implicit representation matches the line graph best.

[0096] In other implementations, the type mapping function may also be selected such as a Sigmoid function, a linear function (such as y=argimax(z i )), you can make adaptive choices based on actual needs.

[0097] In step S312, the computer system determines the most suitable visualization type for the implicit representation of the current template based on the first mapping value set obtained in step S311. This usually means selecting the visualization type corresponding to the largest mapping value in the mapping value set as the prediction result. The system traverses the first mapping value set and finds the largest mapping value therein. This maximum value represents the visualization type that best matches the implicit representation of the current template. Once the largest mapping value is found, the computer system determines the corresponding visualization type as the current mapping data type. This type will be used to guide the subsequent chart generation process. For example, continuing with the previous example, if the mapping value set output by the softmax function is [0.1, 0.8, 0.1], the largest mapping value is 0.8, which corresponds to the line chart type. Therefore, the computer system determines the line chart as the most suitable visualization type for the implicit representation of the current template, and passes this information to the subsequent chart generation module.

[0098] Through the implementation of step S31, the computer system can use the preset visualization type library and type mapping function to predict the most suitable visualization type for the given template implicit representation. This process not only relies on high-quality template implicit representation and effective type mapping function, but also relies on a large amount of training data to optimize the parameters in the type mapping function. In practical applications, the computer system may also need to tune these parameters through methods such as cross-validation, grid search or random search to ensure the accuracy and robustness of the prediction results.

[0099] In one implementation, in step S311, before each visualization type in the preset visualization type library is loaded into the type mapping function together with the current template implicit representation one by one to obtain the first mapping value set, the method further includes:

[0100] Step S301: Obtain target exploration factor;

[0101] Step S302: determining whether the target exploration factor meets a pre-determined confidence requirement;

[0102] Step S303: when the target exploration factor meets the predetermined confidence requirement, arbitrarily determine a candidate mapping data type in the preset visualization type library as the current mapping data type;

[0103] Step S304: When the target exploration factor does not meet the predetermined confidence requirement, the current template implicit representation is loaded into the initial natural language processing network to obtain the current mapping data type.

[0104] In step S301, the target exploration factor is obtained. When training the initial natural language processing network to predict the corresponding mapping data type implicitly represented by the template, the computer system not only needs to use the known optimal solution to optimize the network parameters, but also needs to continuously explore new possible solutions in order to find a better mapping relationship. In order to balance this exploration and utilization process, the computer system introduces an additional variable-the target exploration factor. This factor determines whether the computer system is more inclined to use known information or explore unknown areas in the current iteration. For example, random generation can be performed. A simple and common method is to randomly generate the target exploration factor. At the beginning of each training iteration, the computer system can randomly select a value from a preset range as the target exploration factor of the current iteration. For example, a random number between 0 and 1 can be set, where 0 represents full utilization (no exploration) and 1 represents full exploration (no utilization). Another method is to let the target exploration factor gradually decrease as the training time goes by. This method assumes that in the early stage of training, the model's understanding of the data is not deep enough, so it should be explored more; and as the training deepens, the model gradually converges to a better solution, and at this time, more known information should be used to fine-tune the model parameters. For example, an exploration factor with an initial value of 1 can be set, and decayed at a fixed decay rate (such as 0.99) after each iteration. Alternatively, other more advanced methods are to dynamically adjust the target exploration factor based on the performance of the model on the validation set. If the performance of the model does not improve significantly in multiple consecutive iterations, it may mean that the current strategy is trapped in a local optimal solution. At this time, the value of the exploration factor should be increased to encourage exploration; conversely, if the model performance continues to improve, the value of the exploration factor can be appropriately reduced to make more use of the current strategy. For example, suppose that during the training process, the computer system uses a random generation method to obtain the target exploration factor. At the beginning of an iteration, the computer system randomly generates a number between 0 and 1, such as 0.6, as the target exploration factor for the current iteration. This value indicates that in the current iteration, the computer system has a 60% probability of choosing to explore unknown areas and a 40% probability of choosing to use known information.

[0105] In step S302, it is determined whether the target exploration factor meets a predetermined confidence requirement. After obtaining the target exploration factor, the computer system determines whether the factor meets a predetermined confidence requirement. This confidence requirement is usually set based on the training strategy, data characteristics, and the current performance of the model. By comparing the target exploration factor with the confidence threshold, the computer system can decide whether to explore or exploit in the current iteration.

[0106] Specifically, before training begins, the computer system sets one or more confidence thresholds based on actual conditions. These thresholds may be obtained based on experience, cross-validation results, or theoretical analysis. For example, a number between 0 and 1 (such as 0.5) can be set as the dividing point between exploration and utilization. In each iteration, the computer system compares the obtained target exploration factor with the confidence threshold set in advance. If the exploration factor is greater than the threshold, the system tends to explore unknown areas; if it is less than or equal to the threshold, it tends to utilize known information. For example, continuing with the previous example, assume that the confidence threshold set by the system is 0.5. In the current iteration, the target exploration factor obtained is 0.6, which is greater than the threshold of 0.5, so the system decides to explore more unknown areas in the current iteration.

[0107] In step S303, when the target exploration factor meets the confidence requirement, the candidate mapping data type in the preset visualization type library is arbitrarily determined. When the target exploration factor meets the predetermined confidence requirement, the computer system tends to explore unknown areas. In this step, the computer system does not rely on the prediction results of the initial natural language processing network to determine the mapping data type, but randomly or according to a certain strategy, selects a candidate mapping data type from the preset visualization type library as the result of the current iteration.

[0108] Specifically, a type can be randomly selected directly from the preset visualization type library as the candidate mapping data type. This method is simple and fast, but may lack specificity. Alternatively, if the system has prior knowledge about the data or task (for example, knowing that certain types of data are more suitable for representation with line charts), the selection process can be guided by this prior knowledge. For example, if the system knows that the implicit representation of the current template involves time series data, it may be more inclined to select a line chart as the candidate mapping data type. In some cases, the computer system may use heuristic search algorithms to find the optimal or suboptimal solution in the preset visualization type library. These algorithms are usually based on some kind of evaluation function to evaluate the quality of different solutions, and gradually approach the optimal solution through iterative optimization. For example, suppose the system uses a random selection method to determine the candidate mapping data type. In the current iteration, the preset visualization type library contains bar charts, line charts, and pie charts. Figure 3 The system randomly selects a line chart as the candidate mapping data type for the current iteration.

[0109] In step S304, when the target exploration factor does not meet the confidence requirement, the current template implicit representation is loaded into the initial natural language processing network to obtain the current mapping data type.

[0110] When the target exploration factor does not meet the confidence requirement, the computer system tends to use known information to optimize the model parameters. In this step, the computer system loads the current template implicit representation into the initial natural language processing network and obtains the mapping data type predicted by the network as the result of the current iteration.

[0111] Specifically, the system passes the current template implicit representation (which may be a high-dimensional vector) as input to the initial natural language processing network. This network has been pre-trained with a large amount of training data and has a certain classification ability. After receiving the input, the network will perform forward propagation calculations. This process involves the sequential activation and transformation of multiple network layers, and finally outputs a probability distribution or score vector to represent the possibility of different visualization types. The system determines the current mapping data type based on the network output (such as the category corresponding to the maximum value in the probability distribution). This type will be used as the prediction result of the current iteration for subsequent processing or evaluation. For example, continuing the previous example, assume that in the current iteration, the target exploration factor is 0.4 (less than the confidence threshold of 0.5), so the system decides to use known information. It loads the current template implicit representation into the initial natural language processing network, and after forward propagation calculations, it obtains a probability distribution [0.2, 0.7, 0.1], which corresponds to the probabilities of a bar chart, a line chart, and a pie chart, respectively. The system selects the category with the largest probability (line chart) as the current mapping data type.

[0112] Through the implementation of the above steps S301 to S304, the computer system effectively balances the relationship between exploration and utilization during the training process. When the target exploration factor meets the confidence requirement, the computer system tends to explore unknown areas to find better solutions; when it does not meet the confidence requirement, it tends to use known information to optimize model parameters. This approach not only helps to avoid the model from falling into a local optimal solution, but also improves the generalization ability and final performance of the model.

[0113] In one implementation, the step S312 of determining the candidate mapping data type corresponding to the largest first mapping value in the first mapping value set as the current mapping data type includes:

[0114] Step S3121: when the to-be-selected mapping data type corresponding to the first mapping value is consistent with the target visualization type, determining the current feedback signal as a first feedback signal, wherein the first feedback signal is used to indicate that the initial natural language processing network has successfully identified the data;

[0115] Step S3122: When the to-be-selected mapping data type corresponding to the first mapping value is inconsistent with the target visualization type, determining the current feedback signal as a second feedback signal, wherein the second feedback signal is used to indicate that the initial natural language processing network recognition fails.

[0116] In an embodiment of the present application, step S312 selects the most suitable mapping data type from the first mapping value set output by the type mapping function, and generates a corresponding feedback signal according to the degree of match between the prediction result and the actual target. In step S312, the computer system finds the largest first mapping value from the first mapping value set, and the selected mapping data type corresponding to the value is preliminarily determined as the current mapping data type. Subsequently, the computer system generates different feedback signals based on the match between this prediction result and the actual target visualization type. This process embodies the core idea of ​​supervised learning in machine learning, that is, guiding the training and improvement of the model through known labels (i.e., target visualization types).

[0117] Step S3121 involves the generation of a feedback signal when the recognition is successful. When the type of candidate mapping data corresponding to the first mapping value is completely consistent with the target visualization type, it means that the initial natural language processing network has successfully identified the most suitable visualization type. In this case, the computer system generates a positive feedback signal (i.e., the first feedback signal) to encourage this correct recognition behavior (i.e., the reward idea in reinforcement learning, Reward). Specifically, the first feedback signal is, for example, a specific value or mark, which is used to clearly indicate the situation of successful recognition during the training process of the model. For example, the first feedback signal can be defined as 1 (indicating success) or 'success' (a text mark indicating successful recognition).

[0118] Once the system determines that the predicted visualization type matches the target type, it will automatically generate a first feedback signal and store it together with the current template implicit representation, the predicted visualization type and other information for subsequent training or evaluation of the model. For example, assume that the preset visualization type library contains three types: 'bar chart', 'line chart' and 'pie chart'. In a certain iteration, for a certain template implicit representation, the first mapping value set output by the type mapping function is [0.1, 0.85, 0.05], and the corresponding candidate mapping data types are 'bar chart', 'line chart' and 'pie chart'. Obviously, the mapping value corresponding to 'line chart' is the largest (0.85) and is consistent with the target visualization type. Therefore, the computer system will generate a first feedback signal (such as 1 or 'success'), indicating that the recognition is successful.

[0119] In contrast to step S3121, in step S3122, when the candidate mapping data type corresponding to the first mapping value is inconsistent with the target visualization type, it means that the initial natural language processing network fails to correctly identify the most suitable visualization type. In this case, the computer system generates a negative feedback signal (i.e., the second feedback signal) to indicate the recognition failure and prompt the model to improve in future training.

[0120] Specifically, corresponding to the first feedback signal, the second feedback signal is also a specific value or mark, which is used to indicate the situation of recognition failure. For example, the second feedback signal can be defined as 0 (indicating failure) or 'failure' (a text mark indicating recognition failure). As long as the predicted visualization type does not match the target type, the computer system will generate a second feedback signal. The generation process is similar to the first feedback signal, but the meaning of the signal is completely opposite. In some cases, in order to further guide the improvement direction of the model, the computer system may also calculate the difference or similarity between the predicted type and the target type while generating the second feedback signal, and provide this information as additional feedback to the model. For example, continuing the previous example, if the target visualization type is 'pie chart', but the predicted result is 'line chart', then although the mapping value corresponding to 'line chart' is the largest, it is not equal to the target type, so the system needs to generate a second feedback signal (such as 0 or 'failure'). In addition, the computer system may also calculate the difference between 'line chart' and 'pie chart' (such as by calculating the Euclidean distance between the two types in the visualization feature space), and record this difference as additional feedback information.

[0121] In the process of model training, feedback signals, as part of the supervisory information, play a vital role in adjusting model parameters and optimizing model performance. Specifically, when the system receives the first feedback signal, it tends to maintain or strengthen the current network parameter settings and feature extraction methods; when it receives the second feedback signal, it will try to adjust the network parameters or introduce new feature extraction methods to improve the recognition ability of the model. In addition, the feedback signal can also be used to evaluate the overall performance of the model. By counting the ratio of the first feedback signal to the second feedback signal in different batches of data, the computer system can roughly understand the recognition accuracy and error rate of the model on the current data set, and then make a preliminary judgment on the generalization ability and robustness of the model.

[0122] Through the implementation of step S312, the computer system can generate corresponding feedback signals according to the matching between the prediction results of the initial natural language processing network and the actual target. These feedback signals not only provide clear direction and guidance for further optimization of the model, but also lay a solid foundation for subsequent chart visualization generation by evaluating the recognition performance of the model.

[0123] In one implementation, the method further includes the following steps:

[0124] Step S101: obtaining a plurality of learning sample tuples from the text learning sample, wherein an x-th learning sample tuple in the plurality of learning sample tuples includes an x-th template implicit representation, an x-th mapping data type corresponding to the x-th template implicit representation, an x-th feedback signal, and an y-th template implicit representation, where x≥1 and y=x+1;

[0125] Step S102: Calibrate the initial natural language processing network to be calibrated based on the multiple learning sample tuples to obtain the target natural language processing network, wherein when the number of iterations for calibrating the initial natural language processing network reaches a set maximum number of iterations, the initial natural language processing network is used as the target natural language processing network, and when the number of iterations for calibrating the initial natural language processing network does not reach the set maximum number of iterations, the network configuration variables of the initial natural language processing network are updated based on a preset error determination function, and the input of each iterative calibration is a learning sample tuple from the multiple learning sample tuples.

[0126] In step S101, a learning example tuple is obtained. Before training or tuning a natural language processing model, it is first necessary to prepare a set of rich and representative learning examples. These examples should be able to comprehensively cover various situations that may arise in practical applications to ensure that the model can accurately respond to various inputs after deployment. In the context of chart visualization, learning examples usually include text paragraphs (or sequences of text paragraphs) and their corresponding visualization type labels (mapping data types) and feedback signals of the model's prediction performance. In addition, since the model needs to capture the temporal or logical relationships between text paragraphs, each learning example may also contain implicit representations of multiple consecutive text paragraphs.

[0127] Specifically, in step S101, the computer system extracts multiple learning example tuples from pre-constructed text learning examples. Each tuple contains the following four key elements:

[0128] The xth template implicit representation: This is the implicit representation vector obtained after feature extraction and representation learning of the current text paragraph (xth template). It captures key information in the text paragraph, such as financial indicators, timestamps, company names, etc., and exists in the form of a numerical vector. For example, for a text paragraph describing "Apple's revenue reached 10 billion US dollars in the first quarter of 2022", its template implicit representation may be a vector containing multiple features, such as revenue values, timestamps (quarter identifiers), company names, etc. The corresponding embedding vectors are spliced ​​together.

[0129] The xth mapped data type: This is the most suitable visualization type label for the current text paragraph, such as "line chart", "bar chart", etc. This label is manually annotated to indicate the target predicted by the model.

[0130] The xth feedback signal: This is the feedback that the model gets after predicting the xth template implicit representation during training. It is, for example, a numerical value or a tag that indicates how well the model's prediction matches the actual label. For example, if the model correctly predicts the mapping data type, the feedback signal is positive (such as 1 or 'success'); if the prediction is wrong, the feedback signal is negative (such as 0 or 'failure').

[0131] Implicit representation of the yth template: This is the implicit representation vector of the next text paragraph (the yth template, where y=x+1) of the current text paragraph. It is used to provide contextual information to help the model understand the temporal or logical relationship between text paragraphs. For example, if the xth template describes "Apple's first quarter 2022 revenue", the yth template may describe "Apple's second quarter 2022 revenue", and its implicit representation will contain relevant financial information for the second quarter.

[0132] For example, suppose there are the following two consecutive text paragraphs as training data:

[0133] Paragraph 1: "Apple's revenue in the first quarter of 2022 reached $10 billion."

[0134] Paragraph 2: "By the second quarter, Apple's revenue had grown to $11 billion."

[0135] For these two paragraphs, the following learning example tuples can be constructed:

[0136] The first template implicitly represents: the vector v1 obtained after feature extraction and representation learning of paragraph 1.

[0137] The first mapping data type: "Line chart" (assuming that the change in income over time is best presented using a line chart).

[0138] The first feedback signal: Initially undefined, it will be generated after the model predicts based on the degree of match between the prediction result and the actual label.

[0139] The second template implicitly represents: vector v2 obtained by the same processing of paragraph 2.

[0140] In step S102, the initial natural language processing network is calibrated based on the learning sample tuple. For example, after obtaining a sufficient number of learning sample tuples, it is then necessary to use these samples to calibrate the initial natural language processing network. The calibration process usually involves multiple iterations, each of which uses a learning sample tuple as input, calculates the prediction result through forward propagation, and updates the network parameters based on the difference between the prediction result and the actual label. As the iterations proceed, the prediction performance of the network will gradually improve until the preset convergence conditions are reached (such as the upper limit of the number of iterations, the performance of the validation set is no longer improved, etc.).

[0141] Specifically, in step S102, the computer system performs the following operations to adjust the initial natural language processing network: before starting the adjustment, a set of initial parameter values ​​need to be set for the initial natural language processing network. These parameters are obtained, for example, by random initialization, or can be fine-tuned based on the pre-trained model.

[0142] For each learning example tuple, the computer system performs the following steps for iterative tuning:

[0143] Forward propagation: Input the current learning sample tuple (including the implicit representation of the x-th template and the implicit representation of the y-th template) into the network, and obtain the prediction result (i.e. the predicted probability distribution of the mapped data type) through the forward propagation calculation of the network.

[0144] Calculate loss: Calculate the loss value based on the difference between the predicted result and the actual mapped data type (i.e., the xth mapped data type). This usually involves a loss function (such as the cross entropy loss function) to quantify the prediction error.

[0145] Backpropagation: Update the network's parameters based on the loss value through the backpropagation algorithm. This process involves calculating the gradient of the loss with respect to the network parameters and using the gradient descent (or its variant) algorithm to update the parameter values.

[0146] Update network configuration variables: After each iteration, adjust the network configuration variables (i.e., weights and biases) according to the parameter update rules to better approximate the true mapping relationship in the next iteration.

[0147] Convergence judgment: After each iteration, the computer system checks whether the convergence conditions are met (such as the number of iterations reaches the upper limit, the performance of the validation set is no longer significantly improved, etc.). If the conditions are met, the iteration is stopped and the current network parameters are used as the final parameters of the target natural language processing network; if the conditions are not met, one of the remaining learning sample tuples is selected for the next iteration.

[0148] For example, continuing with the previous example, assume that a natural language processing network has been initialized and is ready to be calibrated using a tuple of learning examples containing paragraphs 1 and 2. In the first iteration, the computer system passes v1 and v2 as input to the network and obtains the predicted probability distribution of the mapped data type. The computer system then calculates the loss value based on the difference between the predicted result and the actual label "line chart" and updates the network parameters through the back-propagation algorithm. This process will be repeated many times until the network's performance on the validation set stabilizes or reaches the preset upper limit of the number of iterations. Ultimately, the fully calibrated network will be able to accurately predict the visualization type of a given text paragraph, providing strong support for subsequent chart visualization generation.

[0149] In one implementation, the method further includes the following steps:

[0150] The target natural language processing network is obtained by tuning the initial natural language processing network to be tuned through multiple groups of learning sample tuples based on the following operations, wherein the learning sample tuple loaded into the initial natural language processing network is the xth learning sample tuple:

[0151] Step SA1: Determine whether the text paragraph training template implicitly represented by the yth template is the last text paragraph training template in the sample text paragraph set.

[0152] During the training process, the computer system iteratively processes these text paragraphs in order to gradually optimize the performance of the natural language processing network. Step SA1 is performed during this iterative process, and its purpose is to determine whether the currently processed text paragraph (identified by its corresponding template implicit representation) is the last one in the sequence. Specifically, when the computer system loads the x-th learning sample multi-tuple for training, it checks the y-th template implicit representation in the group. This template implicit representation is extracted from a text paragraph training template, and it contains the semantic information and key financial indicators of the text paragraph. The system needs to know whether this text paragraph is the last one in the sample text paragraph set (i.e., the entire training data set), because this affects the subsequent training steps and error calculation methods. To determine this, the computer system can maintain an internal counter or index to track the position of the currently processed text paragraph in the sequence. At the beginning of each iteration, the computer system checks whether this index is equal to the total length of the sequence minus one (because the index, for example, starts at 0, and the sequence length includes all elements). If they are equal, it means that the currently processed text paragraph is the last one in the sequence; if they are not equal, it means that there are subsequent text paragraphs to be processed.

[0153] For example, suppose the sample text paragraph collection contains the following three consecutive text paragraphs:

[0154] "Apple's revenue reaches $10 billion in the first quarter of 2022."

[0155] "Apple's revenue grew to $11 billion in the second quarter."

[0156] "In the third quarter, Apple's revenue continued to grow to $12 billion."

[0157] For these three paragraphs, the computer system generates template implicit representations for them respectively, and combines these representations with the corresponding mapping data types (i.e., chart types) and feedback signals to form learning sample tuples. During the training process, when the system processes the third paragraph (i.e., the paragraph of "third quarter"), it checks whether the template implicit representation corresponding to the paragraph (assuming it is the yth template implicit representation) is the last one in the sequence. Since this is the last paragraph in the sequence, the system adjusts the subsequent training steps based on this information (as described in steps SA2 and SA3).

[0158] Step SA2: When the yth template implicitly indicates that the corresponding text paragraph training template is the last text paragraph training template, determine the intermediate variable output by the initial natural language processing network based on the xth feedback signal.

[0159] In step SA2, the computer system determines whether the text paragraph training template corresponding to the implicit representation of the currently processed yth template is the last text paragraph in the sequence. This judgment is based on the text paragraph index or counter maintained internally by the system. When the index value is equal to the sequence length minus one, it can be confirmed that the current template is the last template. Once the yth template is confirmed to be the last template, the computer system will then determine the intermediate variable output by the initial natural language processing network based on the xth feedback signal corresponding to the template. The xth feedback signal is obtained after the model predicts the implicit representation of the xth template in a previous iteration. It reflects the degree of match between the visualization type predicted by the model and the actual label.

[0160] For example, suppose the platform is processing a series of financial report text paragraphs about a company's revenue for four consecutive quarters. During the training process, when the system processes the text paragraph of the fourth quarter (i.e., the last text paragraph), it will execute step SA2.

[0161] For the text paragraph of the fourth quarter, the computer system first converts it into a template implicit representation, which may be a high-dimensional vector containing the embedded representation of the key financial data in the paragraph (such as revenue value, timestamp, etc.). Prior to this, the initial natural language processing network has predicted the text paragraph of the fourth quarter and generated a feedback signal. Assuming that the predicted visualization type is a line chart and the prediction is correct (i.e., consistent with the actual label), the feedback signal will be a positive signal, indicating that the model successfully identified the chart type of the paragraph. In step SA2, the computer system generates an intermediate variable based on this positive feedback signal. This intermediate variable may be a simple value (such as 1, indicating a successful prediction) or a more complex structure (such as a vector containing information such as prediction probability and confidence). However, in practical applications, in order to simplify calculations and improve efficiency, intermediate variables are often designed as values ​​or simple vectors that can be directly used for subsequent error calculations and parameter adjustments. In this example, it can be assumed that the intermediate variable is a simple value 1, indicating that the model's prediction of the last text paragraph is successful.

[0162] In the implementation of chart visualization based on natural language processing, step SA2 ensures that the model can get correct feedback when processing the end paragraph of the sequence text data, and generates appropriate intermediate variables accordingly. This step is crucial for optimizing model parameters and improving model performance, especially when processing text data with time series or logical relationships. By reasonably designing the generation method of intermediate variables, the computer system can more effectively use feedback signals to guide the training process of the model, thereby generating more accurate and useful chart visualization results.

[0163] Step SA3: When the yth template implicitly indicates that the corresponding text paragraph training template is not the end text paragraph training template, the intermediate variable is determined based on the xth feedback signal and the output of the reference natural language processing network, wherein the reference natural language processing network is a natural language processing network obtained by performing parameter pre-configuration processing in advance, and the network configuration variables of the reference natural language processing network are inconsistent with those of the initial natural language processing network.

[0164] In step SA3, the computer system confirms that the currently processed text paragraph training template (identified by its corresponding y-th template implicit representation) is not the last one in the sequence. This means that there are subsequent text paragraphs to be processed, so the system needs to more carefully evaluate the accuracy of the current prediction and consider how to adjust the model parameters to optimize future predictions. To achieve this goal, the computer system not only looks at the x-th feedback signal (which reflects the prediction accuracy of the initial network for the x-th template implicit representation), but also uses the control natural language processing network to predict the same template implicit representation. The control network is a natural language processing network with pre-configured parameters in advance, and its network configuration variables (such as weights and biases) are different from those of the initial network, so it can provide an independent perspective to evaluate the quality of the current prediction. Next, the computer system determines an intermediate variable based on the x-th feedback signal and the output of the control network. This intermediate variable may be a composite value that combines the prediction results of the two, which is used for subsequent error calculation and network parameter adjustment. Specifically, the computer system can use weighted averaging, voting mechanism or other fusion strategies to integrate the two prediction results and generate an intermediate variable accordingly.

[0165] For example, suppose the system is processing a sequence of text paragraphs containing multiple consecutive quarterly financial data, and the current processing is the text paragraph of the second quarter. For this paragraph, the initial natural language processing network predicts that the most suitable visualization type is a line chart, while the control network predicts a bar chart. The system also obtains the xth feedback signal, which indicates that there is a certain difference between the prediction of the initial network and the actual label (but not necessarily wrong, because the prediction of the control network may also be inaccurate). In step SA3, the computer system comprehensively considers these two prediction results and the feedback signal. For example, it can calculate the weighted average of the two prediction results, where the weight may be determined based on the performance of each network on the validation set. Assuming that the weight of the initial network is 0.6 and the weight of the control network is 0.4 (this is just an example, and the actual weight may be dynamically adjusted according to the specific situation), the weighted average prediction result may be more inclined to a line chart. Then, the computer system generates an intermediate variable based on this weighted average prediction result and the feedback signal for subsequent error calculation and network parameter adjustment.

[0166] The specific form of the intermediate variable depends on the design and implementation details of the system. In most cases, it may be a numerical value or vector that represents some kind of comprehensive evaluation of the current prediction result. However, in the dual network architecture, the intermediate variable may also contain the prediction information of the control network or the difference measurement of the prediction results of the two. This information will help the system to more comprehensively evaluate the performance of the initial network and make more refined parameter adjustments accordingly. Step SA3 enhances the training effect of the initial network by introducing a control natural language processing network. When processing non-end text paragraphs, the computer system not only considers the prediction results and feedback signals of the initial network, but also combines the output of the control network to determine the intermediate variables. This dual network architecture helps the system to more comprehensively evaluate the prediction quality and find potential room for improvement during the training process. In this way, the computer system is able to gradually optimize the performance of the initial network to generate more accurate and useful chart visualization results.

[0167] Step SA4: Based on the intermediate variables and the output of the initial natural language processing network, determine the error value of the error determination function, and modify the network configuration variables of the initial natural language processing network based on the error value of the error determination function.

[0168] In step SA4, the computer system calculates the error value based on the intermediate variables determined in the previous step (such as SA2 or SA3) and the output of the initial natural language processing network. This error value reflects the degree of difference between the network prediction result and the actual label, and is an important indicator for measuring model performance. In order to calculate the error value, the computer system uses a predefined error determination function. This function is, for example, a loss function, such as cross-entropy loss (Cross-Entropy Loss), mean squared error (Mean Squared Error, MSE), etc. The specific choice depends on the requirements of the task and the characteristics of the data.

[0169] After calculating the error value, the computer system modifies the network configuration variables of the initial natural language processing network, namely weights and biases, according to the error value. This process can be achieved through the backpropagation algorithm, which uses the chain rule to calculate the gradient of the error with respect to each network parameter and updates the parameter value based on these gradients.

[0170] Specifically, the computer system updates the parameters according to the learning rate and gradient descent (or more complex optimization algorithms such as Adam, RMSprop, etc.). The learning rate is a hyperparameter that controls the step size of the parameter update; gradient descent is an iterative method used to gradually minimize the error function. For example, suppose that when the initial natural language processing network processes a text paragraph, it predicts that the most suitable visualization type is a line chart, but the actual label is a bar chart. In step SA3, the computer system determines an intermediate variable based on the feedback signal and the output of the control network, which reflects the difference between the predicted result and the actual label. In step SA4, the computer system uses the cross entropy loss function to calculate the quantitative value of this difference (i.e., the error value). Assume that the predicted probability distribution is [0.1, 0.8, 0.1] (corresponding to the probabilities of a bar chart, a line chart, and a pie chart, respectively), and the true label is a bar chart (i.e., [1, 0, 0]). Substituting these values ​​into the cross entropy loss formula, the computer system will obtain a positive error value, indicating that there is a difference between the predicted result and the actual label. Next, the computer system uses the back-propagation algorithm to calculate the gradient of the error with respect to each network parameter, and updates the network parameters based on these gradients and the preset learning rate. This process is repeated multiple times (i.e., multiple iterations) until the error value is reduced to an acceptable level or the preset upper limit of the number of iterations is reached.

[0171] Step SA4 is a key step in the model tuning process, which optimizes model performance by calculating the error value and adjusting the network parameters according to the error value. In the chart visualization implementation scenario based on natural language processing, this step ensures that the model can gradually learn the accurate mapping relationship from text paragraphs to chart visualization types, thereby generating more accurate and useful chart visualization results.

[0172] Step SA5: When the number of iterations of the above steps reaches the set maximum number of iterations, the initial natural language processing network is used as the target natural language processing network.

[0173] During the calibration process, the computer system performs multiple iterations according to the preset upper limit of the number of iterations. Each iteration will perform a series of operations from steps SA1 to SA4, including determining the location of the text paragraph, calculating intermediate variables, evaluating prediction errors, and adjusting network parameters. These steps together constitute the core process of model optimization. When the number of iterations reaches the set maximum number of iterations, step SA5 is triggered. At this point, the computer system believes that the initial natural language processing network has been fully trained, and its performance has reached a relatively stable state under the current training data and calibration strategy. Therefore, the computer system saves and uses the current initial natural language processing network as the target natural language processing network.

[0174] The target natural language processing network is the final product of the tuning process. It has the ability to map text paragraphs to corresponding chart visualization types. When saving, the computer system records all the configuration variables of the network (including weights and biases, etc.) so that they can be reloaded and used in subsequent chart visualization tasks. For example, assume that the maximum number of iterations set by the platform is 1000. During the tuning process, the computer system continuously iterates according to the preset steps, and each iteration adjusts the network parameters according to the prediction results and feedback signals of the model. As the iterations proceed, the performance of the model on the validation set gradually improves, and the prediction error gradually decreases. When the number of iterations reaches 1000, the computer system triggers step SA5 and saves the current initial natural language processing network as the target natural language processing network. At this point, the network can already accurately predict the visualization type of key financial indicators in financial reports, such as mapping text paragraphs describing income changes to line graphs, and mapping text paragraphs describing market share distribution to pie charts, etc.

[0175] Step SA5 is the end of the model tuning phase and the starting point of model deployment. By setting the maximum number of iterations and strictly executing the tuning process, the computer system can ensure that the target natural language processing network achieves optimal performance under the training data and tuning strategy. The implementation of this step not only depends on the smooth execution and effective coordination of the previous steps, but also requires reasonable setting of the number of iterations and selection of tuning strategies. In actual applications, platform developers need to flexibly adjust these parameters and strategies according to specific tasks and data characteristics to obtain optimal model performance.

[0176] In one implementation, in step SA3, determining the intermediate variable based on the x-th feedback signal and the output of the reference natural language processing network may specifically include:

[0177] Step SA31: loading each visualization type in the preset visualization type library and the y-th template implicit representation into the linear function one by one to obtain a second mapping value set, wherein the second mapping value set includes a mapping value corresponding to each visualization type;

[0178] Step SA32: taking the sum of the largest second mapping value in the second mapping value set and the xth feedback signal as the intermediate variable.

[0179] In step SA31, the computer system processes the case of a non-end text paragraph. At this point, the computer system not only focuses on the prediction results of the initial network, but also considers the output of the control network in order to obtain a more comprehensive evaluation. Specifically, the computer system matches the currently processed text paragraph (represented by the implicit representation of the yth template) with each visualization type in the preset visualization type library to evaluate the correlation between them. Specifically, first, the computer system maintains a preset visualization type library that contains all possible chart types, such as line charts, bar charts, pie charts, etc. Each type has a unique identifier and is represented in some form (such as enumeration, index, or one-hot encoding) within the system.

[0180] For each type in the preset visualization type library, the computer system loads it together with the yth template implicit representation (representing the text paragraph currently being processed) into a linear function. A linear function is a simple mathematical model that can calculate a scalar value (i.e., a mapping value) based on the input feature vector (here, the template implicit representation and the visualization type identifier), which reflects the linear relationship between the input and the output. In practical applications, the specific form of the linear function may vary depending on the system design and task requirements. But the basic idea is to use the template implicit representation and the visualization type identifier as input features, and obtain an output value through a linear transformation (i.e., weighted summation plus bias). For example, if the template implicit representation is a d-dimensional vector v, and the visualization type identifier is a k-dimensional one-hot encoding vector o (in which only one element is 1 and the rest are 0), then the linear function can be expressed as:

[0181] f(v,o)=w T ·(v⊕o)+b;

[0182] Among them, w is the weight vector, b is the bias term, and ⊕ represents the vector concatenation operation. Note that the weight vector w and bias term b here are pre-set by the system according to task requirements, or obtained through some form of learning (such as pre-training).

[0183] For each type in the preset visualization type library, the computer system performs the above linear function calculation to obtain a mapping value. These mapping values ​​together constitute a second mapping value set, and each element in the set corresponds to an evaluation result of a visualization type.

[0184] For example, suppose the preset visualization type library contains three types: line chart (type1), bar chart (type2) and pie chart (type3). For the yth template implicit representation vy, the computer system loads it and the identifier of each type (such as the identifier of type1 is otype1) into the linear function and calculates three mapping values: mtype1, mtype2 and mtype3. These mapping values ​​reflect the linear correlation between the current text paragraph and different chart types.

[0185] In step SA32, the computer system first finds the maximum value in the second mapping value set, and the visualization type corresponding to the value can be regarded as the optimal prediction of the control network for the current text paragraph. However, since the prediction result of the initial network (represented by the xth feedback signal) needs to be considered, the system combines this maximum value with the xth feedback signal to generate a comprehensive intermediate variable.

[0186] The system traverses the second mapping value set and finds the maximum value m max and its corresponding visualization type. This maximum value reflects the chart type that the control network thinks is the best match for the current text paragraph. The xth feedback signal is obtained after the initial network predicts the xth template implicit representation in a previous iteration. It reflects the degree of match between the prediction result and the actual label. The system will compare this feedback signal with the maximum mapping value m max Combine to generate intermediate variables. The combination can be a simple numerical operation, such as weighted sum, product, or some custom function. But here, in order to keep the explanation simple, it is assumed that the system uses a weighted sum method:

[0187] Va=α·m max +β·Fs;

[0188] Among them, α and β are predefined weight coefficients used to balance the contribution of the maximum mapping value and the feedback signal. These coefficients can be adjusted according to task requirements to achieve the best performance; Va is an intermediate variable and Fs is the feedback signal.

[0189] According to the above formula, the computer system calculates the value of the intermediate variable. This value not only contains the optimal prediction information of the control network for the current text paragraph, but also integrates the prediction result feedback of the initial network, so it can more comprehensively reflect the status and performance of the current network.

[0190] Step SA3 determines the intermediate variables by combining the output of the control natural language processing network and the xth feedback signal, providing a more comprehensive evaluation basis for the adjustment of the initial network. In SA31, the computer system uses a linear function to calculate the correlation between each visualization type and the current text paragraph; in SA32, the computer system combines the maximum mapping value and the feedback signal in a weighted sum to generate intermediate variables. This process not only takes into account the independent perspective of the control network, but also integrates the prediction result feedback of the initial network, which helps to improve the overall performance and generalization ability of the model. In practical applications, the form of the linear function, the value of the weight coefficient and other parameters can be adjusted according to the specific tasks and data characteristics to obtain the optimal model adjustment effect.

[0191] Based on Figure 1 Based on the same principle as the method shown in , the present application also provides a visualization device 10, such as Figure 2 As shown, the device 10 comprises:

[0192] A text acquisition module 11, used to acquire an initial text paragraph set to be processed, wherein the initial text paragraph set includes a target visual text to be recognized;

[0193] The network calling module 12 is used to load the initial text paragraph set into a target natural language processing network that has been calibrated in advance to obtain a recognition result set, wherein the target natural language processing network is a natural language processing network obtained by loading a text learning sample into an initial natural language processing network to be calibrated, the text learning sample includes a plurality of template implicit representations extracted from a plurality of text paragraph training templates, and a mapping data type and a feedback signal corresponding to each template implicit representation, the mapping data type is used to indicate a visualization type of a template implicit representation, the feedback signal is used to indicate a matching degree corresponding to the visualization type, the plurality of text paragraph training templates are a plurality of continuous text paragraphs, and each template implicit representation includes an implicit representation obtained by taking a text paragraph training template and a previous text paragraph training template of the text paragraph training template;

[0194] A type determination module 13, configured to determine a recognition result corresponding to each text paragraph in the initial text paragraph set according to the recognition result set, wherein the recognition result represents a visualization type of the target visual text;

[0195] The visualization module 14 is used to generate a visualization chart according to the recognition result corresponding to each text paragraph to obtain a text visualization chart.

[0196] The above embodiment introduces the visualization device 10 from the perspective of a virtual module. The following describes a computer system from the perspective of a physical module, as shown below:

[0197] The present application embodiment provides a computer system, such as Figure 3 As shown, the computer system 100 includes: a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the computer system 100 may also include a transceiver 104. It should be noted that in actual applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation on the embodiments of the present application.

[0198] In other words, an embodiment of the present application provides a computer system, and the computer system in the embodiment of the present application includes: one or more processors; a memory; one or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by one or more processors, and when the one or more programs are executed by the processor, the above method is implemented.

[0199] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a processor, the processor can execute the corresponding content in the aforementioned method embodiment.

[0200] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for realizing chart visualization based on natural language processing, characterized in that: include: Acquire an initial text paragraph set to be processed, wherein the initial text paragraph set includes target visible text to be recognized; The initial text paragraph set is loaded into a target natural language processing network that has been calibrated in advance to obtain a recognition result set, wherein the target natural language processing network is a natural language processing network obtained by loading text learning samples into an initial natural language processing network to be calibrated, the text learning samples include a plurality of template implicit representations extracted from a plurality of text paragraph training templates, and a mapping data type and a feedback signal corresponding to each template implicit representation, the mapping data type is used to indicate a visualization type of a template implicit representation, the feedback signal is used to indicate a matching degree corresponding to the visualization type, the plurality of text paragraph training templates are a plurality of continuous text paragraphs, and each template implicit representation includes an implicit representation obtained by taking a text paragraph training template and a previous text paragraph training template of the text paragraph training template; Determining a recognition result corresponding to each text paragraph in the initial text paragraph set according to the recognition result set, wherein the recognition result represents a visualization type of the target visual text; Generate a visualization chart based on the recognition result corresponding to each text paragraph to obtain a text visualization chart; Performing a first implicit representation mining process on the multiple text paragraph training templates to obtain a first implicit representation set, wherein an implicit representation in the first implicit representation set corresponds to a text paragraph training template in the multiple text paragraph training templates; Performing a second implicit representation mining process on the multiple text paragraph training templates to obtain a second implicit representation set, wherein a first implicit representation in the second implicit representation set corresponds to a first text paragraph training template in the text paragraph training templates, a previous paragraph text of the first text paragraph training template is a second text paragraph training template, and the first implicit representation is used to characterize the correlation between the first text paragraph training template and the second text paragraph training template; Feature integration is performed on the first implicit representation set and the second implicit representation set respectively to obtain the multiple template implicit representations.

2. The method according to claim 1, characterized in that The performing second implicit representation mining processing on the plurality of text paragraph training templates to obtain a second implicit representation set includes: Based on the following operations, the second implicit representation mining process is performed on the plurality of text paragraph training templates to obtain a second implicit representation set, wherein each text paragraph training template subjected to the second implicit representation mining process is regarded as a current text paragraph training template, the obtained implicit representation is regarded as a current implicit representation, and the second implicit representation set includes the current implicit representation: Performing text detection processing on the current text paragraph training template to determine a first core text character set, wherein the first core text character set is used to indicate the distribution coordinates of the target visible text in the current text paragraph training template; Performing text detection processing on a previous text paragraph training template of the current text paragraph training template to determine a second core text character set, wherein the second core text character set is used to indicate the distribution coordinates of the target visible text in the previous text paragraph training template; The current implicit representation is determined based on the first core text character set and the second core text character set, wherein the current implicit representation represents changes between corresponding core text characters in the first core text character set and the second core text character set.

3. The method according to claim 1, characterized in that The method further comprises: Determine the mapping data type and the feedback signal corresponding to each template implicit representation in the multiple template implicit representations based on the following operations, wherein the template implicit representation for determining the mapping data type and the feedback signal each time is regarded as the current template implicit representation, and the mapping data type corresponding to the current template implicit representation is regarded as the current mapping data type and the current feedback signal: Loading the current template implicit representation into the initial natural language processing network to obtain the current mapping data type, wherein the initial natural language processing network is a natural language processing network obtained by performing parameter pre-configuration processing in advance; The current feedback signal is determined based on the current mapping data type and a template priori tag corresponding to the current template implicit representation, wherein the template prior tag carries the actual type in the text paragraph training template corresponding to the current template implicit representation, the actual type includes the target visualization type of the target visual text, and the feedback signal represents whether the current mapping data type is consistent with the target visualization type.

4. The method according to claim 3, characterized in that The step of loading the current template implicit representation into the initial natural language processing network to obtain the current mapping data type includes: Loading each visualization type in the preset visualization type library together with the current template implicit representation into the type mapping function one by one to obtain a first mapping value set, wherein the first mapping value set includes a mapping value corresponding to each visualization type; Determine the candidate mapping data type corresponding to the largest first mapping value in the first mapping value set as the current mapping data type; Before loading each visualization type in the preset visualization type library together with the current template implicit representation into the type mapping function one by one to obtain the first mapping value set, the method further includes: Get the target exploration factor; Determining whether the target exploration factor meets a predetermined confidence requirement; When the target exploration factor meets the predetermined confidence requirement, arbitrarily determining a candidate mapping data type in the preset visualization type library as the current mapping data type; When the target exploration factor does not meet the predetermined confidence requirement, the current template implicit representation is loaded into the initial natural language processing network to obtain the current mapping data type.

5. The method according to claim 4, characterized in that The step of determining the candidate mapping data type corresponding to the largest first mapping value in the first mapping value set as the current mapping data type includes: When the to-be-selected mapping data type corresponding to the first mapping value is consistent with the target visualization type, determining the current feedback signal as a first feedback signal, wherein the first feedback signal is used to indicate that the initial natural language processing network has successfully identified the data; When the to-be-selected mapping data type corresponding to the first mapping value is inconsistent with the target visualization type, the current feedback signal is determined as a second feedback signal, wherein the second feedback signal is used to indicate that the initial natural language processing network recognition fails.

6. The method according to claim 1, characterized in that The method further comprises: Acquire multiple learning sample tuples from the text learning sample, wherein an x-th learning sample tuple in the multiple learning sample tuples includes an x-th template implicit representation, an x-th mapping data type corresponding to the x-th template implicit representation, an x-th feedback signal, and a y-th template implicit representation, x≥1, y=x+1; The initial natural language processing network to be calibrated is calibrated based on the multiple learning sample tuples to obtain the target natural language processing network, wherein when the number of iterations for calibrating the initial natural language processing network reaches a set maximum number of iterations, the initial natural language processing network is used as the target natural language processing network, and when the number of iterations for calibrating the initial natural language processing network does not reach the set maximum number of iterations, the network configuration variables of the initial natural language processing network are updated based on a preset error determination function, and the input of each iterative calibration is a learning sample tuple from the multiple learning sample tuples.

7. The method according to claim 6, characterized in that The method further comprises: The target natural language processing network is obtained by tuning the initial natural language processing network to be tuned through multiple groups of learning sample tuples based on the following operations, wherein the learning sample tuple loaded into the initial natural language processing network is the xth learning sample tuple: Determine whether the text paragraph training template implicitly represented by the yth template is the last text paragraph training template in the sample text paragraph set; When the yth template implicitly indicates that the corresponding text paragraph training template is the last text paragraph training template, determining the intermediate variable output by the initial natural language processing network based on the xth feedback signal; When the yth template implicitly indicates that the corresponding text paragraph training template is not the last text paragraph training template, the intermediate variable is determined based on the xth feedback signal and the output of the reference natural language processing network, wherein the reference natural language processing network is a natural language processing network obtained by performing parameter pre-configuration processing in advance, and the network configuration variables of the reference natural language processing network are inconsistent with those of the initial natural language processing network; Determine an error value of the error determination function based on the intermediate variable and the output of the initial natural language processing network, and modify the network configuration variables of the initial natural language processing network based on the error value of the error determination function; When the number of iterations of the above steps reaches the set maximum number of iterations, the initial natural language processing network is used as the target natural language processing network.

8. The method according to claim 7, characterized in that The determining the intermediate variable based on the x-th feedback signal and the output of the reference natural language processing network comprises: Loading each visualization type in the preset visualization type library and the y-th template implicit representation into the linear function one by one to obtain a second mapping value set, wherein the second mapping value set includes a mapping value corresponding to each visualization type; The sum of the largest second mapping value in the second mapping value set and the xth feedback signal is used as the intermediate variable.

9. A computer system, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-dimensional statistical graph drawing system under statistical scene based on NPL semantic analysis

    CN118296033A