Chart de-rendering system, method and program for extracting meta information and data information from chart by using artificial intelligence

The AI-based chart de-rendering system efficiently extracts meta and data information from charts by tokenizing data into single tokens and using MLPs, addressing scalability and formatting issues in conventional systems.

WO2025159586A1PCT designated stage Publication Date: 2025-07-31LG MANAGEMENT DEV INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/001512
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-25
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Conventional chart de-rendering systems face scalability issues due to the need for separate models for each chart type, and generative models inefficiently represent numbers using multiple tokens, leading to formatting errors and slow inference speeds.

Method used

A system utilizing a generative AI model that separates meta and data information, tokenizes data into single tokens, and uses multilayer perceptrons (MLPs) to extract information efficiently, minimizing token usage and formatting errors.

Benefits of technology

The system improves scalability, reduces formatting errors, and enhances inference speed by separately extracting meta and data information, optimizing token usage and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025001512_31072025_PF_FP_ABST
    Figure KR2025001512_31072025_PF_FP_ABST
Patent Text Reader

Abstract

A chart de-rendering system, method and program for extracting meta information and data information from a chart by using artificial intelligence are disclosed. The system comprises: a memory including an AI model that includes an image encoder, which inputs a chart, converts same into a first embed that can be processed by an AI model and outputs same, a meta decoder, which outputs a second embed including meta information from the first embed, and a data decoder, which outputs a fourth embed including data information from a third embed that includes information about an entity included in the second embed; and a processor for executing or training the AI model. The data information included in the fourth embed can be tokenized into single tokens for each piece of data information, and a predefined repetition template can be used when extracting the data information from the single tokens.
Need to check novelty before this filing date? Find Prior Art

Description

Chart de-rendering system, method, and program for extracting meta information and data information from a chart using artificial intelligence

[0001] The present invention relates to a chart de-rendering system, method, and program for extracting meta information and data information included in a chart, and more particularly, to a chart de-rendering system, method, and program capable of extracting meta information and data information included in a chart using artificial intelligence.

[0002] Charts included in papers, reports, and textbooks are typically created through a process that sends data tables, typically numbers and groups, and code defining the overall layout (e.g., type, orientation, color / shape composition) to a rendering engine.

[0003] Chart de-rendering is the opposite process of chart rendering, which involves analyzing and grouping visual patterns or information in a chart to extract key information, and then extracting information about the data (e.g., numbers, groups, etc.) and information about the chart layout.

[0004] The initial chart de-rendering process used a rule-based model. A rule-based model extracts chart information using predefined functions (e.g., color detection, chart axis value extraction, etc.) and combines or analyzes each piece of information using predefined rules. While rule-based models offer superior accuracy, they lack scalability because they require separate models for each chart type.

[0005] To address these shortcomings, a generative model approach utilizing a trained artificial intelligence (AI) has been introduced (DePlot: One-shot visual language reasoning by plot-to-table translation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 10381-10399). Generative models offer improved scalability compared to rule-based models, as a single model can be easily applied to all types of charts.

[0006] However, the conventional generative model used a method of including meta information (e.g., names of the X-axis and Y-axis, names of entity (single unit of data) groups recorded in the legend, etc.) and data information (e.g., numeric values ​​of the X-axis and / or Y-axis of each entity) in a single data format without distinguishing them. For example, as shown in Fig. 6, the chart title (Meta Chart), X-axis name (Epoch), Y-axis name (Experiment result), X-axis (Epoch) values ​​[1, 2, 3, 4, 5], and all Y values ​​of each corresponding entity group (Model 1 and Model 2) (i.e., [10, nan, 30, nan, 50] for Model 1, and [1, 4, 9, 16, 25] for Model 2) are recorded in a single data format. These data formats have problems such as complex data representation due to the large amount of information expressed, constraints that all Y values ​​dependent on each X value must be expressed, and an increased possibility of formatting errors due to the inability to identify entity values ​​(e.g., "nan" values).

[0007] Furthermore, conventional generative models typically treated numbers as text. This led to the representation of numbers by dividing a single number into irregular units, each of which was then input as a token. This approach unnecessarily uses a large number of tokens to represent a single number (for example, the number 13.1 requires two tokens, as it is represented by inputting "1" and "3.1" as separate tokens). Consequently, conventional language models ended up using an unnecessary number of tokens, resulting in token waste and slow inference speed.

[0008] Prior art literature Non-patent literature DePlot: One-shot visual language reasoning by plot-to-table translation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 10381-10399

[0009] The problem to be solved by the present invention is to provide a system, method and program that can efficiently and accurately extract numeric information while providing a data format that simply expresses the contents of data included in a chart.

[0010] The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0011] A system for implementing an AI model for extracting meta information and data information included in a chart of the present invention comprises at least one processor, and at least one memory for storing instructions or information for causing the at least one processor to perform an operation, wherein the instructions perform an operation including: inputting the chart into an image encoder to convert it into a first embed that can be processed by the AI ​​model; inputting the first embed into the AI ​​model to output a second embed including the meta information from the first embed, and outputting a fourth embed including the data information from a third embed including information about an entity included in the second embed; and outputting a first data format in which the meta information included in the second embed is recorded and a second data format in which the data information included in the fourth embed is recorded, respectively.

[0012] In the above system, the data information included in the fourth embed can be distinguished for each group of entities.

[0013] In the above system, the meta information may include the title of the chart, the name of the axis, and the name of the entity group included in the legend.

[0014] In the above system, the data information may include numerical information included in a chart.

[0015] In the above system, the data information is tokenized into a single token for each piece of data information and included in the fourth embedding, and when recording the data information included in the fourth embedding in a second data format, the tokenized data information can be extracted from the single tokens and recorded in the second data format.

[0016] In the above system, the data information can be extracted from the single tokens using a multilayer perceptron (MLP).

[0017] In the above system, when extracting the data information from the single tokens, a plurality of pieces of data information can be extracted simultaneously by inputting a predefined repetition template into the single tokens.

[0018] A method for extracting meta information and data information included in a chart according to another aspect of the present invention comprises: inputting the chart into an image encoder to convert it into a first embed that can be processed by the AI ​​model; inputting the first embed into the AI ​​model to output a second embed including the meta information from the first embed; outputting a fourth embed including data information from a third embed including information about an entity included in the second embed; and outputting a first data format in which the meta information included in the second embed is recorded and a second data format in which the data information included in the fourth embed is recorded, respectively.

[0019] In the above method, the data information included in the fourth embedding can be distinguished for each group of entities.

[0020] In the above method, the meta information may include the title of the chart, the name of the axis, and the name of the entity group included in the legend.

[0021] In the above method, the data information may include numerical information included in the chart.

[0022] In the above method, the data information is tokenized into a single token for each piece of data information and included in the fourth embedding, and when recording the data information included in the fourth embedding in a second data format, the tokenized data information can be extracted from the single tokens and recorded in the second data format.

[0023] In the above method, the data information can be extracted from the single tokens using a multilayer perceptron (MLP).

[0024] In the above method, when extracting the data information from the single tokens, a plurality of pieces of data information can be extracted simultaneously by inputting a predefined repetition template into the single tokens.

[0025] According to another aspect of the present invention, a program may be stored in a computer-readable recording medium to implement an AI model that extracts meta information and data information included in a chart through an AI model according to embodiments of the present invention, in combination with a computer.

[0026] According to the present invention, by extracting metadata separately, the length of each data format is shortened, allowing the content of the data to be expressed simply, thereby reducing the possibility of formatting errors occurring.

[0027] In addition, according to the present invention, since raw data is expressed independently by distinguishing it for each entity group, there is no need to input "nan" values ​​that are difficult to process for empty data values, thereby minimizing formatting errors that occur due to failure to identify "nan" values ​​or optimizing token use.

[0028] Also, according to the present invention, each number included in the entity group of the chart is a single token ( <num>) to facilitate separation of tasks between text understanding and data extraction, and to reduce tokens representing numbers, enabling efficient learning and prediction within the model.

[0029] Also, according to the present invention, a sufficient length of ' <num>By pre-supplying tokens, the present invention minimizes the need for automatic regression propagation and further significantly improves inference speed.

[0030] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0031] FIG. 1 is a schematic diagram of a system for implementing an artificial intelligence-based chart de-rendering method according to one embodiment of the present disclosure.

[0032] FIG. 2 is a block diagram illustrating a configuration of a device that performs an artificial intelligence-based chart de-rendering method according to one embodiment of the present disclosure.

[0033] FIG. 3 is a block diagram illustrating a method for extracting meta information and data information from a chart according to embodiments of the present invention.

[0034] FIG. 4 is a conceptual diagram illustrating a method of expressing the number “-0412920” in an embedding through SNE according to embodiments of the present invention.

[0035] Figure 5a is a conceptual diagram showing a method for recognizing and expressing numbers in an automatic regression manner in a conventional language model.

[0036] Figure 5b is a method of recognizing and expressing numbers in an automatic regression manner using SNE according to embodiments of the present invention.

[0037] Figure 5c is a method of recognizing and expressing numbers in a non-automatic regression manner through SNE using a repetitive template according to embodiments of the present invention.

[0038] Figure 6 is an example of a data format used in a conventional generative model.

[0039] The following examples are provided as examples to ensure that those skilled in the art can fully grasp the spirit of the present invention. Therefore, the present invention is not limited to the embodiments described below and may be embodied in other forms.

[0040] Throughout the present invention, the same reference numerals denote the same components. The present invention does not describe all elements of the embodiments, and any content that is general in the technical field to which the present invention pertains or that overlaps between the embodiments is omitted. The terms 'part, module, element, block' used in the specification may be implemented in software or hardware, and depending on the embodiments, multiple 'parts, modules, elements, blocks' may be implemented as a single component, or a single 'part, module, element, block' may include multiple components.

[0041] Throughout the specification, when a part is said to be "connected" to another part, this includes not only direct connection but also indirect connection, and indirect connection includes connection via a wireless communication network.

[0042] Additionally, when a part is said to "include" a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise specifically stated.

[0043] Throughout the specification, when we say that an element is "on" another element, this includes not only cases where the element is in contact with the other element, but also cases where another element exists between the two elements.

[0044] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.

[0045] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0046] The identification codes for each step are used for convenience of explanation and do not describe the order of each step. Each step may be performed in a different order than specified unless the context clearly indicates a specific order.

[0047] A system for de-rendering a chart according to the present invention may include a device, which may include any of a variety of devices capable of performing computational processing and providing results to a user. For example, the system for de-rendering a chart according to the present invention may include at least one of a computer, a server device, and a portable terminal, or may be any form of a device having the same or similar functions as these. However, the present invention is not limited thereto.

[0048] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.

[0049] The above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.

[0050] The above portable terminal may include, for example, a wireless communication device that ensures portability and mobility, and may include all kinds of handheld-based wireless communication devices such as a PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, a smart phone, and a wearable device such as a watch, a ring, a bracelet, an anklet, a necklace, glasses, contact lenses, or a head-mounted device (HMD).

[0051] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.

[0052] The present invention relates to a chart de-rendering system, method, and program for extracting meta information and data information included in a chart, and more particularly, to a chart de-rendering system, method, and program capable of extracting meta information and data information included in a chart using artificial intelligence.

[0053] FIG. 1 is a schematic diagram of a system that can implement a method for de-rendering a chart according to one embodiment of the present invention.

[0054] As illustrated in FIG. 1, the system (1000) may include a device (100), a database (200), and an AI model (300).

[0055] The device (100), database (200), and AI model (300) included in the system (1000) can communicate via a network (W). Here, the network (W) may include a wired network and a wireless network. For example, the network may include various networks such as a local area network (LAN), a metropolitan area network (MAN), and a wide area network (WAN).

[0056] Additionally, the network (W) may include the well-known World Wide Web (WWW). However, the network (W) according to an embodiment of the present invention is not limited to the networks listed above, and may include at least part of a well-known wireless data network, a well-known telephone network, or a well-known wired / wireless television network.

[0057] The device (100) can input a chart and output information about the chart based on the AI ​​model (300).

[0058] The chart input into the AI ​​model (300) may be in the form of a vertical / horizontal bar chart, a line chart, a pie chart, an area chart, a scatter chart, a radial chart, a histogram, and / or a waterfall chart. The chart may be in a single form or may be a form in which multiple forms are combined. However, the chart used in the present invention is not limited thereto, and the chart may include any form of chart. In addition to information that is an image of a number, etc. (e.g., a line, a circle, etc.), the chart may include text information, and specifically, may include annotation information of the chart, a legend or title of the chart, names of each axis (e.g., X-axis, Y-axis, Z-axis, etc.), raw data values ​​of points included in the chart (e.g., numerical values ​​of the X-axis and Y-axis of points included in the chart), etc.

[0059] Information about a chart output from the AI ​​model (300) may include meta information, data information, etc., as information indicating the characteristics of the chart. Meta information may be the names of the X-axis and Y-axis, entity group names recorded in a legend, etc. Data information may be the numerical values ​​of the X-axis and / or Y-axis of each entity, etc.

[0060] According to the present invention, when the device (100) outputs information about a chart based on an acquired chart (image, etc.), it can output meta information and data information included in the information about the chart separately.

[0061] The database (200) can store various types of learning data for training the AI ​​model (300). Furthermore, the database (200) may store chart images, information about charts, and the like, and in various embodiments, may store output data output by the AI ​​model (300). However, the system (1000) may not include the database (200) if the training of the AI ​​model (300) is complete.

[0062] FIG. 1 illustrates a case where a database (200) is implemented outside of a device (100). In this case, the database (200) may be connected to the device (100) via wired or wireless means. However, this is merely an example, and the database (200) may also be implemented as a component of the device (100).

[0063] FIG. 1 illustrates a case where the AI ​​model (300) is implemented outside the device (100) (e.g., cloud-based), but is not limited thereto, and may be implemented as a component in the device (100).

[0064] FIG. 2 is a block diagram illustrating the configuration of a device that performs a method of extracting meta information and data information of a chart using artificial intelligence according to one embodiment of the present invention.

[0065] As illustrated in FIG. 2, the device (100) may include a memory (110), a communication module (120), a display (130), an input module (140), and a processor (150). However, the present invention is not limited thereto, and the device (100) may have its software and hardware configurations modified / added / omitted within a range apparent from a perspective of ordinary skill in the art, depending on the required operation. In addition, the device (100) may be replaced with a system, and the device (100) may include a plurality of devices, in which case each component included in the device (100) may be included in at least one of the plurality of devices.

[0066] The memory (110) can store data supporting various functions of the device (100), programs for the operation of the processor (150), input / output data, and a plurality of application programs (or applications) run on the device, data for the operation of the device (100), commands, and AI models. At least some of these application programs can be downloaded from an external server via wireless communication.

[0067] The memory (110) may include at least one type of storage medium among a flash memory type, a hard disk type, an SSD (Solid State Disk type), an SDD (Silicon Disk Drive type), a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.

[0068] Additionally, the memory (110) may be separate from the device and may include a database connected wired or wirelessly. The database (200) illustrated in FIG. 1 may be implemented as a component of the memory (110).

[0069] The communication module (120) may include one or more components that enable communication with an external device, and may include, for example, at least one of a broadcast reception module, a wired communication module, a wireless communication module, a short-range communication module, and a location information module.

[0070] The wired communication module may include various wired communication modules such as a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, as well as various cable communication modules such as a Universal Serial Bus (USB), a High Definition Multimedia Interface (HDMI), a Digital Visual Interface (DVI), RS-232 (recommended standard 232), power line communication, or plain old telephone service (POTS).

[0071] The wireless communication module may include a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), 4G, 5G, and 6G, in addition to a WiFi module and a Wireless Broadband module.

[0072] The display (130) displays (outputs) information or data processed in the device (100), data input or output through the AI ​​model (300), etc. In addition, the display (130) can display execution screen information of an application program (e.g., an application) running in the device (100), or UI (User Interface) or GUI (Graphical User Interface) information according to such execution screen information.

[0073] The input module (140) is for receiving information from a user. When information is input through the user input unit, the processor (150) can control the operation of the device (100) to correspond to the input information.

[0074] The input module (140) may include hardware physical keys (e.g., buttons, dome switches, jog wheels, jog switches, etc. located on at least one of the front, rear, and side of the device) and software touch keys. For example, the touch keys may be formed of virtual keys, soft keys, or visual keys displayed on a touchscreen type display (130) through software processing, or may be formed of touch keys placed on a part other than the touchscreen. Meanwhile, the virtual keys or visual keys may be displayed on the touchscreen in various forms, and may be formed of, for example, graphics, text, icons, videos, or a combination thereof.

[0075] The processor (150) may be implemented as a memory that stores data on an algorithm for controlling the operation of components within the device (100) (including learning or executing an AI model) or a program that reproduces the algorithm, and at least one processor (not shown) that performs the aforementioned operation using the data stored in the memory. In this case, the memory and the processor may be implemented as separate chips, or may be implemented as a single chip.

[0076] In one embodiment, the system (1000) or device (100) according to the present invention may include at least one processor, and when including multiple processors, the multiple processors may be included in different devices (100).

[0077] In addition, the processor (150) can control any one or a combination of the components described above to implement various embodiments according to the present disclosure described below on the device (100).

[0078] FIG. 3 is a block diagram illustrating a method for extracting meta information and data information from a chart according to an embodiment of the present invention.

[0079] Referring to FIG. 3, a chart (410) including a chart title, axis names, legend information, numbers, etc. is input into an image encoder (420). As described above, the chart (410) may have a vertical / horizontal bar chart, a line chart, a pie chart, an area chart, a scatter chart, a radar chart, a histogram, and / or a waterfall chart, but is not limited thereto. In addition, the chart may have a single shape or may be a combination of multiple shapes. For example, the chart (410) may be a chart (410a) included in FIG. 3.

[0080] The image encoder (420) is an encoder that follows a commonly used encoder architecture and can be implemented as an AI model, specifically, ViT or ResNet, but is not limited thereto. The image encoder (420) converts the input chart (410a) into a first embedding (421) that can be processed by the AI ​​model (300).

[0081] Next, the first embed (421) output from the image encoder (420) is input to an AI model (400) including a meta decoder (430) that performs a meta decode task and a data decoder (450) that performs data decode.

[0082] First, the first embed (421) is first decoded through a meta decoder (430), which may operate in an automatic regression manner. The meta decoder (430) may recognize meta information from the embed (421), such as the names of the X-axis and Y-axis, the entity group names recorded in the legend, etc., but the recognized meta information is not limited thereto. The meta decoder (430) outputs a second embed (431) including the recognized meta information. For example, in the case of using a chart (410a), the second embed (431) may include a chart title (as in the embed (431a) <title>Meta Chart < / title> ), X-axis name ( <xtitle> Epoch < / xtitle> ), Y-axis name ( <ytitle> Experiment result< / ytitle> ), may contain a list of entity groups (Epoch|Model1|Model2) recorded in the legend.

[0083] The present invention has the effect of reducing the length of each data format, enabling the content of the data to be expressed simply by extracting meta information separately from the data information described below, and thus reducing the possibility of formatting errors.

[0084] Next, among the information included in the second embed (431), the information about entities (431b) serves as a prompt that guides the data decoder (450) that decodes data in the AI ​​model (400) to output an embed containing raw data. When the information about entities (431b) is input to the data encoder as the third embed (441), the fourth embed (451) containing raw data about each entity is output. At this time, the fourth embed (451) generates the raw data in the format of [x1|y1 &; x2|y2 &; ... xn|yn &; ], but independently expresses the raw data for each entity group by distinguishing them. For example, when using the chart (410a) shown in FIG. 3, raw data for Model 1 ([1|10 &; 3|30 &; 5|50 & ]) is expressed as in Embed (451a), and raw data for Model 2 ([1|1 &; 2|4 &; 3|9 & 4|16 &; 5|25 &]) is expressed independently.

[0085] In the present invention, when a chart (410a) is input, the output data format is compared with the conventional generation model, and there are the following differences. In the conventional generation model, when the Epoch value is 2 in the chart (410a), Model 2 indicates a value of "4", but Model 1 indicates no value. In order to indicate that Model 2 inputs a value of "4", but inputs a value of "nan" for the value of Model 1, thereby expressing that it is empty (see FIG. 6). In contrast, in the present invention, since the output embedding (451a) is expressed independently for each entity group, the case in which the Epoch value is 2 in the data regarding Model 1 is expressed in a format that does not define it at all. In other words, the present invention may not output a "nan" value, which frequently causes errors during the processing of the system.

[0086] The present invention uses a data format that is independently expressed for each entity group, thereby eliminating the need to input "nan" values ​​that are difficult to process for empty data values, thereby minimizing formatting errors that occur when "nan" values ​​cannot be identified, and optimizing token use by eliminating the need to assign tokens to "nan" values.

[0087] Here, the present invention can use the singularized number embedding (SNE) method as a method for expressing numbers in the fourth embedding (451) (this is in contrast to the conventional generative model that treats numbers as text). SNE is a method that converts each number included in an entity into a single token (in the data decoder (450) that decodes data in the AI ​​model (400). <num>) and outputs a number to the fourth embedding (451) using the token.

[0088] Tokenized single token ( <num>) can be used as a method of extracting numbers. Here, the MLP is composed of an input layer, one or more hidden layers, and an output layer, and each layer includes a weight and an activation function. The weights used in the MLP used in the present invention can be optimized in advance through a learning algorithm so as to be able to predict nonlinear patterns and relationships.

[0089] For example, if we look at the case where the number "-0412920" is expressed in the embedding through SNE as in Fig. 4, each number included in the entity is converted into a single token ( <num>)(451b) is tokenized and the token (451b) is input into the MLP to be expressed as a number on the fourth embedding (451).

[0090] The present invention can be expected to further improve performance by facilitating separation of tasks between text understanding and data extraction, and enables efficient learning and prediction within the model by reducing the number of tokens representing numbers. <num>Inference speed can be greatly improved by repeatedly using tokens.

[0091] Meanwhile, the present invention can further improve the inference speed by inputting a predefined repetition template for tokens used in SNE and outputting numbers.

[0092] Specifically, referring to Figure 5a, conventional language models treat numbers as text and use an autoregressive approach to sequentially predict each token. This approach infers numbers by sequentially feeding tokens at time step t to the model, generating output until reaching the End of Sentence (EOS) token (). As previously described, this approach suffers from token waste and slows inference speed.

[0093] This problem is solved by single token through SNE as shown in Fig. 5b. <num>Some improvements can be made by representing numbers in an autoregressive way using .

[0094] Here, the token is shown in Figure 5c <num>When inferring numbers through predefined repetition templates, multiple tokens <num>If we additionally have a non-automatic regression method that infers multiple numbers at once by inputting them all at once, we can further improve the performance (in this case, tokens are generated through a repeat template). <num>After the EOS token () is first output, all output is discarded.

[0095] The present invention provides a sufficient length of ' <num>By pre-provisioning tokens, the need for automatic regression propagation can be minimized, and further, inference speed can be greatly improved.

[0096] In this way, through the processing process using the AI ​​model, a data format (460a) in which meta information is recorded and a data format (460b) in which data information included in the fourth embedding is recorded can be obtained, respectively.

[0097] Meanwhile, a method for extracting meta information and data information from a chart according to embodiments of the present invention can be implemented by the system described with reference to FIG. 3.

[0098] AI models according to embodiments of the present invention can be controlled, executed, trained, driven, etc. by a processor, and thus, at least one of the tasks of executing, training, and driving the AI ​​models can be performed by at least one processor. Furthermore, the AI ​​models can be stored in memory, and feature data according to the present invention can also be stored in memory.

[0099] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0100] Computer-readable storage media include all types of storage media that store instructions that can be deciphered by a computer. Examples include read-only memory (ROM), random access memory (RAM), magnetic tape, magnetic disks, flash memory, and optical data storage devices.

[0101] The disclosed embodiments have been described with reference to the attached drawings as described above. Those skilled in the art will understand that the present disclosure can be implemented in forms other than the disclosed embodiments without altering the technical spirit or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be construed as limiting.< / num> < / num> < / num> < / num> < / num> < / num> < / num> < / num> < / num> < / num> < / num>

Claims

1. A system for implementing an AI model that extracts meta information and data information included in a chart. at least one processor; and At least one memory storing instructions or information that cause at least one processor to perform an operation, The action performed by the above command is: A step of inputting the above chart into an image encoder and converting it into a first embedding that can be processed by the AI model; A step of inputting the first embed into the AI model, outputting a second embed including the meta information from the first embed, and outputting a fourth embed including the data information from a third embed including information about an entity included in the second embed; A step of outputting a first data format in which the meta information included in the second embedding is recorded and a second data format in which the data information included in the fourth embedding is recorded, respectively; A system that extracts metadata and data information contained in charts.

2. In claim 1, The data information included in the above fourth embed is distinguished for each group of entities, the system 3. In claim 1, The above meta information includes the title of the chart, the names of the axes, and the names of the entity groups included in the legend.

4. In claim 1, The above data information includes numerical information included in the chart, the system.

5. In claim 1, The above data information is tokenized into a single token for each piece of data information and included in the fourth embedding, A system in which the data information included in the fourth embedding is recorded in a second data format, wherein the tokenized data information is extracted from the single tokens and recorded in the second data format.

6. In claim 4, A system for extracting data information from the above single tokens using a multilayer perceptron (MLP).

7. In claim 5 or 6, A system for extracting data information from the single tokens, wherein the single tokens are input into a predefined repetition template to extract multiple pieces of data information simultaneously.

8. A method for extracting meta information and data information included in a chart, The above chart is input into an image encoder to be converted into a first embedding that can be processed by the AI model; By inputting the first embed into the AI model, a second embed including the meta information is output from the first embed; Outputting a fourth embed containing data information from a third embed containing information about an entity included in the second embed; A method for outputting a first data format in which the meta information included in the second embed is recorded and a second data format in which the data information included in the fourth embed is recorded, respectively.

9. In claim 8, A method in which the data information included in the fourth embed is distinguished for each group of entities.

10. In claim 8, A method wherein the above meta information includes the title of the chart, the names of the axes, and the names of the entity groups included in the legend.

11. In claim 8, A method wherein the above data information includes numerical information included in the above chart.

12. In claim 8, A method in which the data information is tokenized into a single token for each piece of data information and included in the fourth embedding, and in recording the data information included in the fourth embedding in a second data format, the tokenized data information is extracted from the single tokens and recorded in the second data format.

13. In claim 12, A method for extracting data information from the above single tokens using a multilayer perceptron (MLP).

14. In claim 12 or 13, A method for extracting data information from the single tokens, wherein a plurality of pieces of data information are extracted simultaneously by inputting a predefined repetition template into the single tokens.

15. A program stored in a computer-readable recording medium for executing the method of any one of claims 8 to 14, in combination with a computer.

Citation Information

Patent Citations

  • Display verification method for web browser, device, computer equipment and storage medium

    KR1020210108341A

  • Eco-friendly fishing sinker

    KR1020230150498A

  • Method and apparatus of reestablishing PDCP entity for header compression protocol in wireless communication system

    KR1020230157927A

  • Virtual Dialog System Performance Assessment and Enrichment

    US20220245199A1

  • KR20230171842A