Table processing method, apparatus, device, and medium
Patent Information
- Application Number
- CN202311287447.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-07
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-10-07
AI Technical Summary
[0011]根据本公开的一个或多个实施例,通过将表格转化为基于自然语言的文本描述来降低表格理解模型理解表格结构的难度,进而专注于处理表格中的重要内容信息,提高表格理解能力的同时,也增强了模型的泛化能力。
Smart Images

Figure CN117312520B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of natural language processing, deep learning, etc., and particularly to a table processing method, table processing device, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] Table comprehension is an important direction in natural language processing for structured data processing. Its main purpose is to simplify the table processing process by using mechanisms such as deep learning, so that people can obtain key information from complex and discrete table data. It has wide application value in social information retrieval.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a table processing method, a table processing apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of this disclosure, a table processing method is provided, comprising: obtaining a target table and a table understanding query statement for the target table; converting the target table into a table description text based on natural language; and inputting the table understanding query statement and the table description text into a table understanding model to obtain a table understanding answer.
[0007] According to one aspect of this disclosure, a table processing apparatus is provided, comprising: an acquisition unit configured to acquire a target table and a table understanding query statement for the target table; a conversion unit configured to convert the target table into table description text based on natural language; and a generation unit configured to input the table understanding query statement and the table description text into a table understanding model to obtain a table understanding answer.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described above.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described method.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described method when executed by a processor.
[0011] According to one or more embodiments of this disclosure, by converting tables into natural language-based text descriptions, the difficulty for table understanding models to understand table structures is reduced, thereby allowing them to focus on processing important content information in the table, improving table understanding capabilities, and enhancing the model's generalization ability.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;
[0015] Figure 2 A flowchart of a table processing method according to an exemplary embodiment of the present disclosure is shown;
[0016] Figure 3 A flowchart of a table processing method according to an exemplary embodiment of the present disclosure is shown;
[0017] Figure 4 A flowchart is shown for filtering at least one table value and at least one table header in a target table according to an exemplary embodiment of the present disclosure;
[0018] Figure 5A flowchart illustrating the conversion of a target table into natural language-based table description text according to an exemplary embodiment of the present disclosure is shown;
[0019] Figure 6 A schematic diagram illustrating the generation of updated table description text using a text generation model according to an exemplary embodiment of the present disclosure is shown.
[0020] Figure 7 A schematic diagram illustrating the generation of table comprehension answers using a table comprehension model according to an exemplary embodiment of the present disclosure is shown.
[0021] Figure 8 A flowchart illustrating an exemplary embodiment of the present disclosure is provided, showing how table description text and contextual information are input into a text generation model to obtain updated table description text.
[0022] Figure 9 A flowchart illustrating an exemplary embodiment of the present disclosure is provided, showing how to input the text to be recovered and contextual information into a text generation model to obtain connected text;
[0023] Figure 10 A bar chart comparing the results of a table processing method according to an exemplary embodiment of the present disclosure and a table processing method based on row-by-row flattening is shown;
[0024] Figure 11 A structural block diagram of a table processing apparatus according to an exemplary embodiment of the present disclosure is shown; and
[0025] Figure 12 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0027] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0028] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0029] In related technologies, existing table processing methods often use a row-by-row flattening approach to understand tables, which often requires a lot of special methods and data processing to flatten the table structure, which undoubtedly hinders the understanding of the table content.
[0030] To address the aforementioned issues, this disclosure reduces the difficulty for table understanding models to comprehend table structures by converting tables into natural language-based text descriptions. This allows the models to focus on processing important content information within the tables, improving table comprehension capabilities while also enhancing the model's generalization ability.
[0031] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0032] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0033] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the methods of this disclosure.
[0034] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) network.
[0035] exist Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0036] Users can use client devices 101, 102, 103, 104, 105, and / or 106 for human-computer interaction. The client devices provide interfaces that enable users to interact with them. The client devices can also output information to the user through these interfaces. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0037] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0038] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0039] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0040] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0041] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0042] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0043] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a data repository used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be a database, such as a relational database. One or more of these databases may store, update, and retrieve data from and from the database in response to commands.
[0044] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0045] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0046] According to one aspect of this disclosure, a table processing method is provided. For example... Figure 2 As shown, the table processing method includes: step S201, obtaining the target table and the table understanding query statement for the target table; step S202, converting the target table into table description text based on natural language; and step S203, inputting the table understanding query statement and table description text into the table understanding model to obtain the table understanding answer.
[0047] Therefore, by converting tables into natural language-based text descriptions, the difficulty for table understanding models to comprehend table structures is reduced, allowing them to focus on processing important content information within the tables. This not only improves table understanding capabilities but also enhances the model's generalization ability.
[0048] Tabular data differs from ordinary text data, possessing a two-dimensional feature of rows and columns. Different row and column positions impart different attribute information to the text. However, traditional natural language processing systems are often only adept at processing one-way natural language text. Understanding the special information that row and column positions in a table convey to the text has become a common challenge in table understanding systems.
[0049] In one exemplary embodiment, by processing Table 1 using a row-by-row flattening method, the following result can be obtained: "Name|Height / cm|Weight / kg|Date of Birth, Xiaoming|150|50|1998.12.1, Xiaozhang|160|52|1999.1.1". Such a result makes it difficult for a natural language processing system to understand the specific meaning of the table.
[0050] Table 1. Student List of Class 2, Grade 3
[0051] Xiaoming 150 50 1998.12.1 Xiao Zhang 160 52 1999.1.1
[0052] By recombining the discrete text in a table and connecting them by identifying the specific attribute relationships between the table headers and values, the two-dimensional information of the table is transformed into natural language-based text, reducing the difficulty of understanding the table caused by the positional relationships of the cells.
[0053] In an exemplary embodiment, the table values and corresponding column headers in Table 1 can be converted into natural language-based table description text: "For Xiaoming in Class 2, Grade 3, Xiaoming's height is 150cm, weight is 50kg, and date of birth is 1998.12.1; for Xiaozhang in Class 2, Grade 3, Xiaozhang's height is 160cm, weight is 52kg, and date of birth is 1999.1.1." Thus, by combining the table name, the table values in the first row ("Xiaoming", "150", "50", "1998.12.1"), and the corresponding column headers ("Name", "Height / cm", "Weight / kg", "Date of Birth"), we can obtain a text description format that the natural language processing model excels at processing. After processing all cells of the table using the above method, we obtain the overall description of the table. By describing the table values in natural language-based text, the model focuses on processing table information rather than table structure information such as row and column positions. Simultaneously, by combining this with a pre-trained language model, the model's ability to understand tables is improved.
[0054] In some embodiments, in step S201, the target table and the corresponding table understanding query statement can be obtained. The table understanding query statement can be, for example, a query statement entered by the user, or a question obtained through other means, and is not limited here.
[0055] In some embodiments, tables, as data carriers, may contain a great deal of text. Excessively long table sequences not only increase system processing time and affect the accuracy of results, but also, when the length reaches a certain limit (e.g., some models limit input length to 1024 word segments), the portion exceeding the limit is directly discarded, resulting in the loss of some table information during input. To handle excessively long tables, the relevant content for the table comprehension query can be identified within the target table, thereby filtering the necessary rows and columns and removing irrelevant table content.
[0056] According to some embodiments, the target table may include at least one table value and at least one table header corresponding to the column containing the at least one table value. For example... Figure 3 As shown, the table processing method may further include: step S302, before converting the target table into table description text, filtering at least one table value and at least one table header in the target table based on the table understanding query statement. It is understood that... Figure 3 The operations in steps S301, S303-S304 and Figure 2 The operations of steps S201-S203 are similar and will not be described in detail here.
[0057] In some embodiments, such as Figure 4 As shown, step S302 may include: step S401, matching at least one table header with the table understanding query statement for keywords to obtain one or more first table headers among at least one table header; step S402, matching at least one table value with the table understanding query statement for keywords to obtain one or more first table values among at least one table value; and step S403, retaining one or more first table headers in the target table, and retaining the table values in the columns corresponding to one or more first table headers and the table values in the rows where one or more first table values are located.
[0058] Therefore, by using keyword matching, all content that "hardly" matches the table understanding query can be filtered out from the target table for subsequent table understanding.
[0059] In some embodiments, in step S401, after removing some stop words, word-level matching can be performed using table headers and table understanding query statements to obtain one or more first table headers containing the same keywords. In step S402, a similar method can be used to perform word-level matching using table values and table understanding query statements to obtain one or more first table values containing the same keywords. In step S403, one or more first table headers can be retained, along with the table values in the columns corresponding to the one or more first table headers and the table values in the rows (and optionally, the columns containing the one or more first table values). It is understood that all table values in the rows and columns determined in the above manner can be retained, as can the table values at the intersection of rows and columns; this is not limited here.
[0060] According to some embodiments, such as Figure 4 As shown, step S302, before converting the target table into table description text, filtering at least one table value and at least one table header in the target table based on the table understanding query statement may include: step S404, embedding at least one table header into a table header vector and embedding the table understanding query statement into a query vector; step S405, determining one or more second table headers in at least one table header based on the similarity between the table header vector and the query vector; and step S406, retaining one or more second table headers in the target table and retaining the table values in the columns corresponding to one or more second table headers.
[0061] Therefore, through semantic matching, all content that "softly" matches the table understanding query can be found in the target table for subsequent table understanding.
[0062] In an exemplary embodiment, in step S404, all table headers and table-understanding query statements can be mapped to a 384-dimensional dense vector space using a word embedding model. Then, in step S405, the cosine similarity between the table header vectors and the query vectors is calculated, and the K (K>1) table headers with the highest similarity scores are selected as candidate table headers (i.e., one or more second table headers). In step S406, one or more second table headers and the table values in the corresponding columns are retained. It is understood that all table values in the rows and columns determined in the above manner can be retained, as can the table values at the intersection of rows and columns; this is not limited here.
[0063] "Soft" matching effectively solves the omission problem caused by the incomplete coverage of "hard" matching. For example, in response to the question "Who is the tallest student in Class 2, Grade 3?" in the target table of Table 1, "hard" matching would retain the column containing "height" and filter out the main column "name" because there are no keywords in that column that match the query. In contrast, "name" in "soft" matching often has a high degree of matching with "who" in the question, so the system will retain the column containing "name". After two rounds of filtering, the system can retain as many rows and columns as possible that are relevant to the query while removing redundant table information.
[0064] In some embodiments, the target table may include at least one table value and at least one table header corresponding to the column containing the at least one table value. For example... Figure 5 As shown, step S202, converting the target table into table description text based on natural language, may include: step S501, combining at least one table header with at least one table value based on a preset template to obtain table description text.
[0065] Therefore, by combining table headers and values based on preset templates, the table comprehension model can better understand the specific meaning of each table value, thereby improving the accuracy of the table comprehension answers output by the model.
[0066] In some embodiments, a script can be used to combine the table header and table value corresponding to the column containing the table value based on a preset template. It is understood that when implementing the method of this disclosure, the corresponding template can be selected, designed, or generated according to the requirements, and no limitation is made herein.
[0067] In an exemplary embodiment, for Table 2, the following template can be used to combine content: "For <Player> [table value corresponding to the table header <Player>], in <Year> [table value corresponding to the table header <Year>], <Result> is [table value corresponding to the table header <Result>]". It is understood that after filtering the table content, some rows may only include some of the table values corresponding to the table headers. For example... Figure 6 As shown in the upper left and upper middle parts, for example, if the table comprehension query for Table 2 is "Who lost the game?", then according to "hard" matching, the row and column corresponding to the table value "lost" and the table header "result" and its corresponding column can be retained. According to "soft" matching, the table header "player" and its corresponding column can be retained ("player" and "who" in the table comprehension query are semantically matched). For the content of this target table, the above targets can be used to combine the content to obtain the table description text "For player Zhang San, the result is win; for player Li Si, in the year 2011, the result is loss; for player Wang Wu, the result is win".
[0068] Table 2 Annual Tennis Tournament Results
[0069] Zhang San 2011 win Li Si 2011 lose Wang Wu 2013 win
[0070] In some embodiments, in step S203, the obtained table description text and table understanding query statement can be input together into the table understanding model to obtain the table understanding answer for the table understanding query statement output by the model.
[0071] According to some embodiments, the table understanding model can be a large pre-trained language model. Large pre-trained language models (LLMs) are typically pre-trained on corpora ranging from millions to billions of words to acquire as much language knowledge as possible. After converting the table content into natural language-based table description text, the large pre-trained language model can be directly used to process the table understanding query and table description text, leveraging its powerful natural language understanding capabilities and extensive language knowledge to generate accurate table understanding answers.
[0072] In one exemplary embodiment, the table comprehension model can use the T5 model. The T5 model is a neural network model based on the Transformer architecture. Figure 7 As shown, after obtaining the converted table description text, the table comprehension query statement and the table description text can be concatenated into a single text sequence 702 and input into the T5 model. After passing through the model's encoding module 704 and decoding module 706, the question answer 708 is directly output based on the content of the concatenated text sequence. Thus, a text-based approach to solving the table comprehension problem is realized.
[0073] In real social life, when artificially converting tables into natural text, relying solely on the information from a single table may make it difficult to generate fluent text because there is a lack of external information to act as a "bridge" connecting the table content.
[0074] Table 3. Voting Results of the Applying Schools in Each Round
[0075] City A XX middle school 35 63 City B YY Middle School 22 16 City C ZZ middle school 17 8 City D NN middle school 13 -
[0076] The context information corresponding to Table 3 is as follows: After two rounds of voting, XX Middle School in City A was selected to host the middle school science and technology competition.
[0077] As shown in Table 3, the description of "XX Middle School in City A" can be easily obtained from "City" and "City A", and "School" and "XX Middle School". However, the meaning of "35" is not directly derived from the relationship between the table header "First Round" and its first value "35". Therefore, it is necessary to use the contextual information of the table to assist in the conversion from table to text when generating table text. By combining the contextual information in Table 3, "After two rounds of voting, XX Middle School in City A was selected to host the Middle School Science and Technology Competition" and the table title "Voting Results of Each Round for the Applying School", we can obtain the description of the first row of the table as "XX Middle School in City A received 35 votes in the first round and 63 votes in the second round for its application to host the Middle School Science and Technology Competition".
[0078] Based on this idea, this disclosure proposes to use a text generation model to determine the connection relationship between table headers and table values, thereby obtaining a description of the entire table.
[0079] In some embodiments, traditional natural language processing methods typically extract text features using rule-based or feature engineering approaches. However, these methods often neglect contextual information, thus failing to fully understand the meaning and context of the text. Therefore, methods based on deep learning models can be used to perceive the contextual information of the text through attention mechanisms or other means, thereby achieving a better understanding of the text's meaning.
[0080] According to some embodiments, such as Figure 5 As shown, step S202, converting the target table content into a table description text based on natural language, may include: step S502, obtaining the context information of the target table, which includes at least one of the table name and text content related to the target table in the context of the target table; and step S503, inputting the table description text and the context information into a text generation model to obtain an updated table description text. The updated table description text may include connection text generated by the text generation model, which uses natural language to describe the relationship between at least one table value and at least one table header.
[0081] Therefore, by obtaining the contextual information of the target table and using a text generation model to generate connecting text based on the transformed table description text and the contextual information of the target table, the connecting text uses natural language to describe the relationship between the table values and the corresponding table headers in the target table. This allows the updated table description text to incorporate the contextual information of the target table and to describe the relationship between the table values and the corresponding table headers in the target table more accurately, effectively, and fluently. This improves the predictive ability of the updated table description text and enables the downstream table understanding model to generate more accurate table understanding answers based on the updated table description text.
[0082] In some embodiments, in step S503, the table description text can be concatenated with context information and then input into the text generation model. The text generation model can output connecting text that describes the relationship between the table values and corresponding headers in the target table using natural language, or it can directly output an updated table description text that includes the connecting text.
[0083] Understandably, in step S203, the table understanding query statement and the updated table description text can be input into the table understanding model to obtain a table understanding answer.
[0084] In some embodiments, a fill-mask mechanism can be used to fully leverage the context-aware capabilities of the text generation model. Fill-mask can be used to automatically complete masked words or phrases. Specifically, in natural language processing tasks, if a word or phrase in the text is masked (e.g., replaced with a special symbol or mark), the text generation model will attempt to predict the masked word or phrase.
[0085] According to some embodiments, such as Figure 8 As shown, step S503, inputting the table description text and context information into the text generation model to obtain the updated table description text, may include: step S801, inserting a marker to be restored into the table description text to obtain the text to be restored, the marker being used to instruct the text generation model to generate connecting text at the position of the marker; step S802, inputting the text to be restored and context information into the text generation model to obtain the connecting text generated by the text generation model for the marker to be restored; and step S803, inserting the connecting text into the table description text to obtain the updated table description text.
[0086] Therefore, by leveraging the fill-mask mechanism of the text generation model, markers to be restored can be inserted into the table description text to instruct the text generation model to generate connecting text at the locations of these markers. Connecting text generated using the fill-mask mechanism makes the connections between table values and corresponding headers in the target table smoother and more coherent, ultimately resulting in more coherent table description text. This improves the predictive power of the updated table description text, enabling downstream table understanding models to generate more accurate table understanding responses based on the updated table description text.
[0087] In some embodiments, such as Figure 6As shown, the table description text can include multiple text fragments, each fragment consisting of a table value and a corresponding table header. In step S801, a to-be-restored marker can be inserted at the corresponding position in each fragment to instruct the text generation model to generate connection text connecting the table value and the corresponding table header in that fragment at the position of the to-be-restored marker. In step S802, the to-be-restored text with multiple to-be-restored markers inserted and context information can be concatenated and input into the text generation model to obtain the connection text generated by the text generation model for each to-be-restored marker. In step S803, each connection text can be inserted into the table description text at the position of the to-be-restored marker corresponding to that connection text to obtain the updated table description text.
[0088] In an exemplary embodiment, the target table obtained after performing "hard" and "soft" matching on Table 2 can be converted into table description text "For player Zhang San, the result is a win; for player Li Si, in the year 2011, the result is a loss; for player Wang Wu, the result is a win," and a mark to be restored can be inserted into this table description text. <mask>The text to be recovered is "for the player". <mask>Zhang San, <mask>The result was a win; for the players... <mask>Li Si, in <mask>2011 (Year) <mask>The result was a loss; for the player <mask>Wang Wu, <mask>The result is "win". After inputting the text to be recovered into the text generation model, the text generation model can generate the corresponding link text at the position of each mark to be recovered. Inserting the link text into the position of the mark to be recovered in the table description text, the updated table description text can be obtained: "For player [name] Zhang San, the result of [tennis match] is win; for player [name] Li Si, in [year] 2011, the result of [tennis match] is loss; for player [name] Wang Wu, the result of [tennis match] is win".
[0089] As can be seen in the above exemplary embodiments, the connection text is not limited in number of characters and can be blank content "[]", indicating that no corresponding connection text is needed between the table value and the corresponding table header.
[0090] In some embodiments, the insertion position of the mark to be restored may be predetermined in a preset template or determined in the table description text in other ways, which is not limited here.
[0091] In some embodiments, one or more tags to be restored can be inserted into the text fragment obtained by combining each table value and the corresponding table header, thereby allowing the text generation model to freely select appropriate positions from these tags to generate the corresponding connected text, while blank content can be generated in the remaining positions.
[0092] According to some embodiments, such as Figure 9 As shown, step S802, inputting the text to be recovered and context information into the text generation model to obtain the connected text generated by the text generation model for the marker to be recovered, may include: step S901, obtaining at least one candidate connected text generated by the text generation model for the marker to be recovered; step S902, determining a first index for each of the at least one candidate connected text, the first index indicating the fluency of the text obtained after inserting the corresponding candidate connected text into the table description text; and step S903, determining the connected text in the at least one candidate connected text based on the first index of each of the at least one candidate connected text.
[0093] Therefore, by generating at least one candidate linking text and selecting from the candidate linking text based on a first indicator of the fluency of the text, it is possible to obtain linking text of better quality that makes the updated table description text more coherent and smooth, thereby improving the predictive ability of the updated table description text and enabling the downstream table understanding model to generate more accurate table understanding answers based on the updated table description text.
[0094] In some embodiments, the text generation model outputs a probability distribution representing the probability of possible words or phrases (or blanks) given a context, from which we select the top K most likely words or phrases (or blanks) as candidate connected text.
[0095] topK(s) = w i ∈V|s(w i )≥s(w j for all w j ∈V, where i∈[1,K]
[0096] Where s is the prediction score vector of the text generation model for the input text to be recovered and contextual information, and w i V is the i-th possible word or phrase (or blank) in the vocabulary, and V is all possible words or phrases (or blanks) in the model vocabulary.
[0097] According to some embodiments, the first metric can be perplexity to evaluate the overall fluency of the generated table description text after inserting the connecting text. Perplexity represents the uncertainty of the model predicting the next word or character in a given dataset. Specifically, perplexity is the reciprocal of the probability that the model predicts the next word or character in a given text sequence; the closer the value is to 1, the more fluent the generated text. Therefore, for the generated K candidate connecting texts (i.e., K candidate table description texts with inserted connecting texts), the candidate connecting text that results in the lowest perplexity of the corresponding candidate table description text can be selected as the connecting text, and the corresponding updated table description text can be obtained.
[0098] It is understandable that other evaluation metrics can also be used as the primary metric to assess the overall fluency of the table description text after the insertion of linked text; this is not a limitation here.
[0099] In some embodiments, the text generation model can be a large pre-trained language model. Large pre-trained language models typically possess context-aware capabilities, meaning they can make predictions using contextual information rather than just considering the currently input word or phrase. Therefore, by using a large pre-trained language model as the text generation model, the model can fully understand the contextual information of the target table, the table values, and the corresponding headers, and generate accurate, effective, and fluent connecting text describing the relationship between the table values and their corresponding headers. This improves the predictive power of the updated table description text, enabling the downstream table understanding model to generate more accurate table understanding answers based on the updated table description text.
[0100] In some embodiments, the table understanding model and / or text generation model used in the table processing method provided in this disclosure can both be pre-trained language models, and no further training is required for table understanding and / or connected text generation capabilities, thereby reducing the need for additional training resources. In other words, the table processing method transforms the table understanding task into a natural language processing task by converting the content in the target table into a natural language-based table description text, thus fully utilizing the natural language understanding and text generation capabilities of the pre-trained language model.
[0101] In some embodiments, Figure 10 The results of testing on the WTQ, SQA, and Tablefact datasets, comparing the table processing method provided in this disclosure with a table processing method based on row-by-row flattening, with only the table input format changed, are shown. It can be seen that the table processing method provided in this disclosure has a significant improvement in accuracy.
[0102] According to another aspect of this disclosure, a form processing apparatus is provided. For example... Figure 11 The apparatus 1100 includes: an acquisition unit 1110 configured to acquire a target table and a table understanding query statement for the target table; a conversion unit 1120 configured to convert the content of the target table into table description text based on natural language; and a generation unit 1130 configured to input the table understanding query statement and the table description text into a table understanding model to obtain a table understanding answer.
[0103] It is understandable that the operation of units 1110-1130 in device 1100 is related to... Figure 2 The operations of steps S201-S203 are similar and will not be described in detail here.
[0104] According to some embodiments, the target table may include at least one table value and at least one table header corresponding to the column containing the at least one table value. The apparatus 1100 may further include (not shown): a filtering unit configured to filter at least one table value and at least one table header in the target table based on a table understanding query statement before converting the target table into table description text. The filtering unit may include: a first table header determining subunit configured to perform keyword matching between at least one table header and the table understanding query statement to obtain one or more first table headers among the at least one table headers; a first table value determining subunit configured to perform keyword matching between at least one table value and the table understanding query statement to obtain one or more first table values among the at least one table values; and a first retaining subunit configured to retain one or more first table headers in the target table, and retain the table values in the columns corresponding to the one or more first table headers and the table values in the rows containing the one or more first table values.
[0105] According to some embodiments, the filtering unit may include: an embedding subunit configured to embed at least one table header as a table header vector and embed a table comprehension query statement as a query vector; a second table header determination subunit configured to determine one or more second table headers in at least one table header based on the similarity between the table header vector and the query vector; and a second retention subunit configured to retain one or more second table headers in the target table and retain table values in the columns corresponding to the one or more second table headers.
[0106] According to some embodiments, the target table may include at least one table value and at least one table header corresponding to the column containing the at least one table value. The conversion unit 1120 may include a combination subunit configured to combine at least one table header with at least one table value based on a preset template to obtain table description text.
[0107] According to some embodiments, the table comprehension model can be a large pre-trained language model.
[0108] According to some embodiments, the conversion unit 1120 includes: a context information acquisition subunit configured to acquire context information of a target table, the context information including at least one of the table name of the target table and text content related to the target table in the context of the target table; and an update subunit configured to input table description text and context information into a text generation model to obtain updated table description text, wherein the updated table description text includes connection text generated by the text generation model, the connection text using natural language to describe the relationship between at least one table value and at least one table header.
[0109] According to some embodiments, the update subunit may include: an insertion subunit configured to insert a marker to be restored into the table description text to obtain text to be restored, the marker to be restored being used to instruct the text generation model to generate connection text at the position of the marker to be restored; a generation subunit configured to input the text to be restored and context information into the text generation model to obtain connection text generated by the text generation model for the marker to be restored; and an insertion subunit configured to insert the connection text into the table description text to obtain the updated table description text.
[0110] According to some embodiments, the generation subunit may include: a candidate connection text acquisition subunit, configured to acquire at least one candidate connection text generated by a text generation model for a marker to be restored; a first index determination subunit, configured to determine a first index for each of the at least one candidate connection text, the first index indicating the fluency of the text obtained after inserting the corresponding candidate connection text into a table description text; and a connection text determination subunit, configured to determine connection text in the at least one candidate connection text based on the first index for each of the at least one candidate connection text.
[0111] According to some embodiments, the first metric can be perplexity.
[0112] According to some embodiments, the text generation model can be a large pre-trained language model.
[0113] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0114] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0115] refer to Figure 12 The present invention describes a structural block diagram of an electronic device 1200 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0116] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1202 or a computer program loaded from storage unit 1208 into random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.
[0117] Multiple components in device 1200 are connected to I / O interface 1205, including: input unit 1206, output unit 1207, storage unit 1208, and communication unit 1209. Input unit 1206 can be any type of device capable of inputting information to device 1200. Input unit 1206 can receive input numerical or character information and generate key signal input related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 1207 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1208 may include, but is not limited to, a hard disk and an optical disk. Communication unit 1209 allows device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, a modem, network card, infrared communication device, wireless communication transceiver, and / or chipset, such as Bluetooth. TM Devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0118] The computing unit 1201 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning network algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as table processing methods. For example, in some embodiments, the table processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the table processing method described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform table processing methods by any other suitable means (e.g., by means of firmware).
[0119] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0123] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0124] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0125] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0126] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.< / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask>
Claims
1. A table processing method, comprising: Obtain a target table and a table understanding query statement for the target table, wherein the target table includes at least one table value and at least one table header corresponding to the column containing the at least one table value; Converting the target table into a natural language-based table description text includes: Based on a preset template, the at least one table header and the at least one table value are combined to obtain the table description text. The table description text uses natural language to describe the relationship between each of the at least one table value and its corresponding table header. The table description text includes multiple text segments, each segment being obtained by combining a table value and its corresponding table header. Obtain the context information of the target table, the context information including at least one of the table name of the target table and text content related to the target table in the context of the target table; and The table description text and the context information are input into the text generation model to obtain the updated table description text, including: Insert a marker to be restored at the corresponding position in each text segment included in the table description text to obtain the text to be restored. The marker to be restored is used to instruct the text generation model to generate connecting text connecting the table value and the corresponding table header in the text segment at the position of the marker to be restored. The text to be recovered and the context information are input into the text generation model to obtain the connection text generated by the text generation model for the tag to be recovered, including: Obtain at least one candidate link text generated by the text generation model for the tag to be restored; Determine a first index for each of the at least one candidate linking texts, the first index indicating the fluency of the text obtained after inserting the corresponding candidate linking text into the table description text; and Based on a first index of each of the at least one candidate connection texts, determine the connection text among the at least one candidate connection texts; and Insert the connection text into the corresponding position in the table description text to obtain the updated table description text; and The table comprehension query and the updated table description text are concatenated into a text sequence, and the text sequence is input into the table comprehension model to obtain the table comprehension answer.
2. The method according to claim 1, wherein, The first metric is confusion level.
3. The method according to claim 1, further comprising: Before converting the target table into table description text, the query statement is understood based on the table, and at least one table value and at least one table header in the target table are filtered, including: The at least one table header is matched with the table understanding query statement for keywords to obtain one or more first table headers among the at least one table headers; The at least one table value is matched with the table comprehension query statement using keywords to obtain one or more first table values from the at least one table value; and The target table retains the one or more first table headers, and also retains the table values in the columns corresponding to the one or more first table headers and the table values in the rows containing the one or more first table values.
4. The method according to claim 3, wherein, Based on the table, understanding the query statement and filtering at least one table value and at least one table header in the target table includes: The at least one table header is embedded as a header vector, and the table interpretation query statement is embedded as a query vector; Based on the similarity between the header vector and the query vector, one or more second headers are determined from the at least one header; and The target table retains the one or more second table headers, and also retains the table values in the columns corresponding to the one or more second table headers.
5. The method according to any one of claims 1-4, wherein, The table comprehension model is a large pre-trained language model.
6. The method according to any one of claims 1-4, wherein, The text generation model is a large-scale pre-trained language model.
7. A form processing apparatus, comprising: The acquisition unit is configured to acquire a target table and a table understanding query statement for the target table, wherein the target table includes at least one table value and at least one table header corresponding to the column containing the at least one table value; The conversion unit is configured to convert the target table into a natural language-based table description text, including: The combination subunit is configured to combine the at least one table header and the at least one table value based on a preset template to obtain the table description text. The table description text uses natural language to describe the relationship between each of the at least one table value and its corresponding table header. The table description text includes multiple text segments, each segment being obtained by combining a table value and its corresponding table header. A context information acquisition subunit is configured to acquire context information of the target table, the context information including at least one of the table name of the target table and text content related to the target table in the context of the target table; and An update subunit is configured to input the table description text and the context information into a text generation model to obtain an updated table description text. The update subunit includes: The marker insertion subunit is configured to insert a marker to be restored at a corresponding position in each text segment included in the table description text to obtain the text to be restored. The marker to be restored is used to instruct the text generation model to generate connecting text connecting the table values and corresponding table headers in the text segment at the position of the marker to be restored. A generation subunit is configured to input the text to be recovered and the context information into the text generation model to obtain the connection text generated by the text generation model for the marker to be recovered; and A text insertion subunit is configured to insert the connecting text into the corresponding position in the table description text to obtain the updated table description text; and The generation unit is configured to concatenate the table comprehension query statement and the updated table description text into a text sequence, and input the text sequence into the table comprehension model to obtain a table comprehension answer. The generation subunit includes: The candidate connection text acquisition subunit is configured to acquire at least one candidate connection text generated by the text generation model for the marker to be restored. The first indicator determining subunit is configured to determine a first indicator for each of the at least one candidate link text, the first indicator indicating the fluency of the text obtained after inserting the corresponding candidate link text into the table description text; and The connection text determination subunit is configured to determine the connection text among the at least one candidate connection text based on a first index of each of the at least one candidate connection text.
8. The apparatus according to claim 7, wherein, The first metric is confusion level.
9. The apparatus according to claim 7, further comprising: A filtering unit is configured to filter at least one table value and at least one table header in the target table based on the table understanding query statement before converting the target table into table description text. The filtering unit includes: The first header determination sub-unit is configured to perform keyword matching between the at least one header and the table understanding query statement to obtain one or more first headers among the at least one headers; The first table value determination subunit is configured to perform keyword matching between the at least one table value and the table comprehension query statement to obtain one or more first table values from the at least one table value; and The first retention subunit is configured to retain the one or more first table headers in the target table, and to retain the table values in the columns corresponding to the one or more first table headers and the table values in the rows where the one or more first table values are located.
10. The apparatus according to claim 9, wherein, The filtering unit includes: The embedded subunit is configured to embed the at least one table header as a header vector and to embed the table interpretation query statement as a query vector. The second header determination subunit is configured to determine one or more second headers in the at least one header based on the similarity between the header vector and the query vector; and The second retention subunit is configured to retain the one or more second table headers in the target table and retain the table values in the columns corresponding to the one or more second table headers.
11. The apparatus according to any one of claims 7-10, wherein, The table comprehension model is a large pre-trained language model.
12. The apparatus according to any one of claims 7-10, wherein, The text generation model is a large-scale pre-trained language model.
13. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.
Citation Information
Patent Citations
Table description processing method and device, equipment and storage medium
CN110377910A
Multi-table retrieval method and device based on natural language
CN116049354A
Text description generation method and system for table data
CN116484837A