A system and method for inserting data into a database structured based on an image representation of a data table.

A neural network-based system accurately identifies and inserts data from complex tables into structured databases by locating content objects and determining cell positions, addressing the challenges of variable table formats and backgrounds in large-scale data collection.

JP7698626B2Active Publication Date: 2025-06-25NFERENCE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022502444
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-16
Filing Date
2020-07-16
Publication Date
2025-06-25
Estimated Expiration
2040-07-16

AI Technical Summary

Technical Problem

Inserting information into a database accurately and efficiently, particularly in large-scale data collection and storage, is challenging due to the variability in table formats and the absence of clear row or column markers, merged cells, and complex backgrounds.

Method used

A system and method that utilize a neural network model to identify the location of content objects within an image representation of a data table, determine cell locations based on these objects, and insert data into a structured database by associating content objects with classification identifiers, creating rows and columns as needed, and extracting sequence information from graphical sequence objects.

Benefits of technology

Enables accurate and efficient extraction and storage of data from complex tables into structured databases, handling variable formats and backgrounds, and updating data in real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698626000001
    Figure 0007698626000001
  • Figure 0007698626000002
    Figure 0007698626000002
  • Figure 0007698626000003
    Figure 0007698626000003
Patent Text Reader

Abstract

1. A system and method for populating a structured database, the system and method comprising: accessing a graphical representation of a data table including one or more cells arranged in rows and columns; providing the graphical representation as input to a neural network model; running the neural network model to identify a location of a first content object within the graphical representation; identifying a location of a first cell based on the location of the first content object; determining that the first cell belongs to a first row and a first column based on the location of the first cell and the first content object in relation to a plurality of content objects; associating the first content object with one or more classification identifiers; and populating the structured database with the first content object and the one or more classification identifiers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 62 / 874,830, filed on July 16, 2019, entitled “Systems and Methods for Populating a Structured Database Based on an Image Representation of a Data Table”, the entire content of which is incorporated herein by reference.

[0002] This application generally relates to databases, and more specifically, to techniques for inserting data into a structured database based on an image representation of a data table.

Background Art

[0003] Database technology enables vast amounts of data to be digitally stored and accessed in an efficient manner. For example, many emerging “big data” applications are made possible by the development of database technology. Databases can be stored locally within a data center and / or within the cloud. Databases can also be distributed across multiple facilities.

[0004] Databases can be structured in various ways. For example, relational databases model data as a set of tables where each of the data is arranged in rows and columns. Query languages can be used to programmatically access data from a database and manipulate the data stored within the database.

[0005] However, inserting information into a database and keeping that information accurate and up - to - date can be a difficult task. Therefore, it is desirable to develop improved techniques for inserting data into a database, including automated techniques suitable for large - scale collection and storage of information within the database.

SUMMARY OF THE INVENTION

[0006] Systems and methods for inserting data into a structured database based on an image representation of a data table according to embodiments of the present disclosure include accessing, by one or more computer processors, an image representation of a data table, where the data table includes one or more cells arranged in one or more rows and one or more columns, the one or more cells including a first cell belonging to at least one first row and at least one first column, and a first content object is inserted into the first cell; providing, by one or more computer processors, the image representation as an input to a neural network model trained to identify the location of content objects within the image representation; executing, by one or more computer processors, the neural network model to identify the location of the first content object within the image representation; identifying, by one or more computer processors, the location of the first cell based on the location of the first content object; and by one or more computer processors, determining that the first cell belongs to at least one first row and at least one first column based on the location of the first cell and one or more of the first content object in relation to a plurality of content objects associated with the one or more rows and the one or more columns; associating the first content object with one or more classification identifiers; and inserting, by one or more computer processors, information associated with the first content object and the one or more classification identifiers into a structured database based on determining that the first cell belongs to at least one first row and at least one first column, the structured database including at least one data table row associated with at least one first row and at least one data table column associated with at least one first column.

[0007] In some embodiments, the system and method may also include creating one of at least one second column and at least one second row in a structured database based on a determination that a first cell does not belong to at least one first row and at least one first column. In some embodiments, accessing an image representation comprises receiving, by one or more computer processors via a computer network, a digital document, the digital document including a data table; rendering, by one or more computer processors, the digital document as a digital image; and identifying, by one or more computer processors, an image representation of the data table within the rendered digital image. In other embodiments, the location of a first content object includes a first region corresponding to at least a portion of the first content object, and identifying the location of a first cell based on the location of the first content object includes expanding the first region in at least one direction; determining that the expanded first region includes a graphical marker marking one or more of a row boundary and a column boundary; and in response to determining that the expanded first region includes a graphical marker, identifying the expanded first region as corresponding to the location of the first cell.

[0008] In some embodiments, determining that the extended first region includes a graphical marker includes identifying a plurality of pixel positions corresponding to the edge of the extended first region, and for each pixel position within the plurality of pixel positions, determining whether the pixel position is associated with one or more changes in color and intensity along at least one extension direction that exceeds a first predetermined threshold, and determining that the count of the plurality of pixel positions associated with the change in color or intensity exceeds a second predetermined threshold, and in response to determining that the number of the plurality of pixel positions exceeds the second predetermined threshold, determining that the extended first region includes a graphical marker. In other embodiments, the location of the first cell includes a row span along the row axis and a column span along the column axis, and based on the location of the first cell, determining that the first cell belongs to at least one first row and at least one first column includes sorting at least a subset of one or more cells in a data table based on the locations of the plurality of cells, and recursively performing operations for identifying one or more second cells belonging to the first row starting from a selected cell among the subset of the one or more cells, the operations including determining at least one other cell having a row span overlapping the row span of the selected cell, identifying the cell among the at least one other cell that is closest to the selected cell, identifying the closest cell as belonging to at least one first row, selecting the closest cell as the next selected cell, identifying the header row among one or more rows of the data table based on one or more header content objects into which data is inserted into one or more header cells of the header row, determining that the column span of the first cell overlaps the column span of the first header cell among the one or more header cells, and identifying the first cell as belonging to the first column, the first column being associated with the first header cell.

[0009] In some embodiments, identifying a header row from among one or more rows of a data table includes generating one or more text representations corresponding to one or more header content objects and querying a header dictionary with each of the one or more text representations, resulting in a score vector that includes one or more confidence scores corresponding to the one or more text representations, each confidence score being based on the strength of the query, querying, determining a row score based on the score vector, and selecting a header row based on the row score. In other embodiments, determining a row score based on a score vector includes calculating an aggregation metric based on the score vector and one or more of the one or more confidence scores. In yet other embodiments, selecting a header row includes comparing the row score to at least one secondary row score associated with one or more rows of the data table and selecting the header row based on the relative values of the row score and the at least one secondary row score.

[0010] In some embodiments, a system and method include obtaining, by one or more computer processors, a list of excluded header content objects that are not eligible to be part of a header row, determining, by one or more computer processors, whether one or more header content objects inserted with data into one or more header cells of a header row match the list of excluded header content objects, and, if one or more header content objects are on the list of excluded header content objects, identifying, by one or more computer processors, a replacement header row among one or more rows of the data table based on the one or more header content objects inserted with data into one or more header cells of the header row. In other embodiments, a first content object includes a graphical sequence object, and inserting data into a structured database includes extracting sequence information from the graphical sequence object, and information associated with the first content object includes the sequence information.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 9E

Figure 9G

Figure 9G

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 11E

Figure 11F

[0012] The various objectives, features, and advantages of the disclosed subject matter will become more fully apparent from the following detailed description of the disclosed subject matter when considered in conjunction with the following drawings in which like reference numerals identify like elements.

DETAILED DESCRIPTION OF THE INVENTION

[0013] Due to the various formats and structures of tables, extracting information from table data can be a difficult task. The rows and columns of a data table can have variable widths, heights, and spacing. A data table may or may not have row or column markers to identify the boundaries between adjacent rows or columns. Some data tables include merged cells. Additionally, a data table can include a complex background or cell coloring scheme.

[0014] For example, a biopharmaceutical company may provide a web page or downloadable report containing pharmaceutical pipeline information. This information is often presented in tabular form. For example, a pharmaceutical pipeline information table can include various information about products in development, such as drug name, target, mechanism of action, disease, and current phase of development. The phase of development can be presented graphically, for example, using progress bars of different shapes, sizes, and colors. A particular progress bar may span multiple columns, and even if the column markers appear to divide the cell containing the progress bar into multiple columns, the cell should be treated as a merged cell.

[0015] The present disclosure describes systems and methods for extracting information from data tables, such as those described above, and storing them in a structured database for subsequent retrieval and analysis.

[0016] FIG. 1 is a simplified diagram of a system 100 for inserting data into a structured database based on an image representation of a data table according to some embodiments. System 100 includes a plurality of devices 101-109 communicatively coupled via a network 110. Devices 101-109 generally include computer devices or systems such as personal computers, mobile devices, servers, etc. Network 110 may include one or more local area networks (LANs), wide area networks (WANs), wired networks, wireless networks, the Internet, etc. Exemplarily, devices 101-109 may communicate via network 110 using the TCP / IP protocol or other suitable network protocols.

[0017] One or more of devices 101-109 may store and / or access digital documents 121-129 via network 110. For example, as illustrated in FIG. 1, devices 101, 102, and 109 store digital documents 121, 122, and 129, respectively, and device 103 accesses digital documents 121-129 via network 110. Digital documents 121-129 may include web pages, digital files, digital images (including one or more frames of video or animation), etc. Exemplarily, digital documents 121-129 may be formatted as HTML / CSS documents, PDF documents, word processor documents (e.g., Word documents), text documents, slide show presentations (e.g., PowerPoint presentations), image files (e.g., JPEG, PNG, or TIFF images), etc. For efficient storage and / or transmission via network 110, documents 121-129 may be compressed before or during transmission via network 110. Security measures such as encryption, authentication (including multi-factor authentication), SSL, HTTPS, and other security technologies may also be applied.

[0018] According to some embodiments, device 103 can access one or more of digital documents 121-129 by downloading the digital documents 121-129 from devices 101, 102, and 109. Further, one or more of devices 101, 102, or 109 can upload the digital documents 121-129 to device 103. The digital documents 121-129 can be updated at various times. Thus, device 103 can access the digital documents 121-129 multiple times at various intervals (e.g., periodically) to obtain the latest copy.

[0019] At least one of the digital documents 121-129 can include one or more data tables 131-139. For example, the data tables 131-139 may be embedded within the digital documents 121-129, linked from within the digital documents 121-129, etc. The data tables 131-139 can be stored in various formats, such as an image format, a text format (e.g., a CSV or TSV file), a markup language format (e.g., XML or HTML / CSS).

[0020] As shown in FIG. 1, device 103 includes a processor 140 (e.g., one or more hardware processors) coupled to a memory 150 (e.g., one or more non-transitory memories). Memory 150 stores instructions and / or data corresponding to processing pipeline 160 and neural network model 170 (or multiple neural network models). When executed by processor 140, processing pipeline 160 inserts data into database 180 based on the image representations of data tables 131-139. Since digital documents 121-129 are generally stored and accessible in various formats, processing pipeline 160 may convert digital documents 121-129 and / or data tables 131-139 into image representations in preparation for processing. This preliminary conversion step enables processing pipeline 160 to process data tables received in HTML / CSS and PDF formats, for example, using the same technology.

[0021] Database 180 may be configured as a structured database having content organized according to a schema or other logical relationship. For example, database 180 may be a relational database. Although database 180 is shown as being directly coupled to device 103, it should be understood that various other configurations are possible. For example, database 180 may be stored within memory 103 or accessed via network 110, etc.

[0022] During the execution of processing pipeline 160, processor 140 executes neural network model 170. Neural network model 170 is trained to make predictions based on input data. Neural network model 170 includes a plurality of layers of neural network model 170 and a configuration 172 that defines the relationships between the layers. Exemplary examples of layers include an input layer, an output layer, a convolutional layer, a densely connected layer, a merge layer, and the like. In some embodiments, neural network model 170 may be configured as a deep neural network having at least one hidden layer between the input layer and the output layer. The connections between the layers may include feed-forward connections or recurrent connections.

[0023] One or more layers of neural network model 170 are associated with trained model parameters 174. The trained model parameters 174 are a set of parameters (e.g., weights and bias parameters of artificial neurons) that are learned according to a machine learning process. During the machine learning process, labeled training data is provided as input to neural network model 170, and the values of the trained model parameters 174 are repeatedly adjusted until the predictions generated by neural network 170 match the labels corresponding to the desired level of accuracy.

[0024] For improved performance, processor 140 may execute neural network model 170 using a graphical processing unit, a tensor processing unit, an application-specific integrated circuit, or the like.

[0025] FIG. 2 is a simplified diagram of data table 200 according to some embodiments. In some embodiments that are consistent with FIG. 1, data table 200 may generally correspond to at least one of data tables 131-139.

[0026] The data table 220 includes one or more cells 231 - 239 arranged in one or more rows 241 - 249 and one or more columns 251 - 259. Generally, each cell belongs to at least one row and at least one column. Further, one or more of the cells 231 - 239 may correspond to combined cells that occupy multiple rows, multiple columns, or both. For example, as illustrated in FIG. 2, cell 235 corresponds to a combined cell spanning columns 252 - 259.

[0027] Content objects 261 - 269 are inserted into one or more of the cells 231 - 239. The content objects 261 - 269 may include various types of content such as text, graphics, equations, animated content, or combinations thereof.

[0028] According to some embodiments, one or more of the content objects 261 - 269 may include a graphical sequence object. For example, as illustrated in FIG. 2, a content object 269 including a graphical sequence object 270 is inserted into cell 235. The graphical sequence object 270 represents sequence information such as timing or phase information. For example, the graphical sequence object 270 may represent the development stage of a project, the clinical trial phase of a pharmaceutical, etc. In some embodiments, the graphical sequence object 270 may illustrate the sequence information using a progress bar, and the length of the progress bar (e.g., the number of columns the progress bar spans) conveys the sequence information. Generally, graphical sequence objects such as the graphical sequence object 270 may be provided in a wide variety of shapes, sizes, colors, textures, patterns, etc.

[0029] One or more of lines 241 to 249 can be designated as the header row of the data table 200. For example, as shown in FIG. 2, the topmost line 241 is designated as the header row. The content of the header row includes information that explains the content of other rows, such as text labels included in the cells of the columns below the individual cells of the header row. For example, cells 232 and 233 in header row 241 have content objects 261 and 262 inserted with content objects 282 and 284 including header content objects 282 and 284, respectively. Header content object 282 includes information that explains the content of other rows in column 252, and header content object 284 includes information that explains the content of other rows in column 259.

[0030] In some embodiments, adjacent columns or rows of the data table 200 can be delimited using graphical markers such as graphical column marker 292 or graphical row marker 294. Graphical column marker 292 and graphical row marker 294 are illustrated as solid lines in FIG. 2, but many alternatives are possible. For example, the graphical marker can include lines of various styles (e.g., dashed, dotted, double lines, etc.), background colors, or style transitions (e.g., adjacent rows or columns can be delimited by alternating between a light background color and a dark background color, or between different textures). As will be appreciated by those skilled in the art, the graphical marker can be applied in a wide variety of ways depending on the style and content of the data table 200. Some rows and / or columns can include graphical markers, while others can be omitted.

[0031] FIG. 3 is a simplified diagram of a method 300 for inserting data into a structured database based on an image representation of a data table according to some embodiments. According to some embodiments consistent with FIGS. 1 and 2, method 300 can be implemented by a computer processor such as processor 140 based on instructions and / or data stored in a memory such as memory 150.

[0032] In process 301, an image representation of a data table such as data table 200 is accessed. The image representation includes pixel data representing the data table. An exemplary embodiment of accessing the image representation of the data table will be described below with reference to FIG. 4.

[0033] In process 302, a neural network model such as neural network model 170 is accessed. The neural network model is trained to identify the location of content objects in the image representation. Exemplarily, a content object may correspond to a logical group of text, such as a text box. Thus, the neural network model may be trained to identify logical groups of text within the image representation. For example, the neural network model may include a text detector that detects words in the image representation, and a heuristic approach may be used to identify logical groups of the detected words.

[0034] In some embodiments, the neural network model may be trained to directly identify logical groups of text. An example of a neural network model that can identify logical groups of text in this way is the YOLOv3 neural network, which is described in Joseph Redmon and Ali Farhadi, YOLOv3: An Incremental Improvement, Technical report, 2018, which is hereby incorporated by reference in its entirety.

[0035] In some embodiments, the neural network model can be trained using transfer learning to identify one or more types of content objects expected to be found within a data table. For example, the neural network can be trained to identify (1) logical groups of text within a cell (e.g., text boxes), and (2) graphical sequence objects (e.g., progress or phase bars). Subsequent processes of method 300 can be performed for each type of content object identified by the neural network model.

[0036] In process 303, the image representation is provided as an input to the neural network model. Various preprocessing steps can be performed to prepare the image representation for the neural network model. These preprocessing steps can include trimming and / or padding the image representation to fit a predetermined aspect ratio, scaling the dimensions of the image representation to fit a predetermined size, normalizing the color or intensity of the pixels within the image representation, reducing the number of color channels of the image representation (e.g., converting the image representation from color to grayscale), and the like.

[0037] In process 304, the neural network model is executed to identify the location of a first content object within the image representation. The first content object can include a logical group of text, a graphical sequence object, and the like. According to some embodiments, the neural network model can be executed using dedicated computing hardware such as a graphics processing unit (GPU) or an application-specific integrated circuit (ASIC). The location of the first content object can include the coordinates of a point associated with the first content object (e.g., the center position of the first content object), the horizontal and vertical spans of the first content object, a bounding rectangle (or other suitable shape) surrounding the first content object, and the like.

[0038] More generally, executing a neural network model can identify the locations of multiple content objects within an image representation. Process 304 of method 300 and subsequent processes are described with reference to a first content among the multiple content objects, but these processes can be repeated for each of the multiple identified content objects.

[0039] In process 305, the location of the first cell is identified based on the location of the first content object. The first cell corresponds to a cell in a data table into which the first content object has data inserted. Since the first content object is contained within the first cell, the first cell generally corresponds to a region of the image representation that is larger than the first content object. Thus, identifying the location of the first cell can be achieved by expanding the region corresponding to the first content object until the expanded region reaches a boundary associated with the first cell. Exemplary embodiments of methods for identifying the location of the first cell in this way are described below with reference to FIGS. 5 and 6. According to some embodiments, process 305 can be repeated for each of the multiple content objects identified in process 304, resulting in the locations of a corresponding multiple of cells within the data table. In this regard, each of the multiple cells can be associated with a different content object and can have a different location.

[0040] In process 306, the first cell is determined to belong to at least one first row and at least one first column based on the location of the first cell. Generally, a cell within a data table belongs to a single row and a single column. However, the first cell can correspond to a combined cell, in which case the first cell can span multiple rows, multiple columns, or both. Exemplary embodiments of methods for determining that the first cell belongs to at least one first row and at least one first column are described below with reference to FIG. 7.

[0041] In process 307, in a structured database such as database 180, information associated with the first content object is inserted into the data based on determining that the first cell belongs to the first row and the first column. Inserting data into a structured database may include extracting information based on the first content object. For example, when the first content object includes a logical group of text, inserting data into a structured database may include converting the logical group of text from an image representation to a sequence of digital characters.

[0042] When the first content object includes a graphical sequence object, inserting data into a structured database may include extracting sequence information from the graphical sequence object. For example, when the graphical sequence object includes a progress or phase bar, the sequence information may be determined based on the length of the progress or phase bar, or the number of rows or columns that the progress or phase bar spans. In some embodiments, the length of the progress or phase bar is first aligned with the columns or rows of the table before determining the sequence information. In some embodiments, determining the length of the progress or phase bar may include distinguishing between the filled and unfilled portions of the bar and identifying the length of the filled portion. In a scenario where the progress or phase bar spans a portion of a column or row, a percentage of overlap may be determined. For example, if the phase bar spans 60% of the column corresponding to Phase II, it may be determined that Phase II is 60% complete.

[0043] According to some embodiments, one or more processes of method 300 may be repeated until information associated with each content object in the data table is inserted into the structured database. Once the data is inserted, various types of analysis or visualization may then be performed based on the information stored in the structured database. As an illustrative example, in some embodiments, semantic analysis may be performed using the techniques described in U.S. Patent No. 10,360,507, filed on September 22, 2017, entitled "Systems, Methods, and Computer Readable Media for Visualization of Sematic Information and Inference of Temporal Signals Indicating Salient Associations Between Life Science Entities", which is hereby incorporated by reference in its entirety.

[0044] Figure 4 is a simplified diagram of a method 400 for accessing an image representation of a data table according to some embodiments. According to some embodiments that are consistent with FIGS. 1-3, method 400 may be used to implement process 301 of method 300.

[0045] In process 401, digital documents such as digital documents 121-129 are received via a computer network such as network 110. The digital documents may be transmitted and received in various formats. For example, the digital documents may include HTML / CSS documents, image files (e.g., JPEG, PNG, or TIFF images), PDF documents, text or word processing documents, slide show presentations, spreadsheets, and the like.

[0046] In process 402, the digital document is rendered as a digital image. For example, rendering a digital document may include converting the digital document into an array of pixel values that can be used for further processing (and optionally, displayed on a display screen). The rendering engine may be selected to render the digital document into a uniform image format based on the format in which the digital document is received. For example, when the digital document includes an HTML / CSS document, a web browser may be selected to render the document. Similarly, when the digital document includes a PDF document, a PDF viewer may be selected to render the document. In either case, the digital document can be rendered into a uniform digital image format regardless of the format of the received digital document. In this way, flexibility is provided for handling a wide variety of types of received digital documents. In some embodiments, metadata associated with the received digital document (e.g., metadata from a PDF file that describes the contents of a data table included in the PDF file) may be removed from the rendered digital image or otherwise not included.

[0047] In process 403, the image representation of the data table is located within the rendered digital image. One of ordinary skill in the art will understand that a wide variety of object detection techniques can be used to locate the image representation of the data table within the digital image. According to some embodiments, a second neural network model can be trained to detect and localize a data table within a digital image. This second neural network model can then be executed using the rendered digital image as input to predict the location of the image representation of the data table. In an exemplary embodiment, the neural network model can correspond to the SSD512 neural network model that is trained using transfer learning to detect and localize the image representation of the data table. The SSD512 neural network model is described in more detail in Wei Liu et al., SSD: Single Shot MultiBox Detector, European Conference on Computer Vision, 2016, which is hereby incorporated by reference in its entirety.

[0048] In some embodiments, method 400 can be performed multiple times to update the data table over time. For example, when the data table includes phase or progress information that changes or evolves over time, method 400 can be performed periodically to track the indicated phase or progress. Then, a method such as method 300 can be performed each time the data table is updated to insert data into a structured database based on the updated content of the data table.

[0049] FIG. 5 is a simplified diagram of a method 500 for identifying the location of cells based on the location of content objects, according to some embodiments. According to some embodiments that are consistent with FIGS. 1-3, method 500 can be used to implement process 305 of method 300.

[0050] In process 501, a first region corresponding to at least a portion of a first content object is expanded in at least one direction. For example, the first region may correspond to a bounding rectangle surrounding the first content object, such as a box around a logical group of text. In some embodiments, the edges of the bounding rectangle may be aligned to be parallel to the expected directions of the rows and / or columns of a data table. For example, if the rows and columns correspond to the horizontal and vertical axes of an image representation, respectively, the edges of the bounding rectangle may likewise be aligned with the horizontal and vertical axes of the image representation. However, other shapes (e.g., non-rectangular shapes) and / or orientations of the first region are equally applicable to the systems and methods described herein. When the first region corresponds to a bounding rectangle, at least one of the four edges of the bounding rectangle may be shifted outward from the center of the bounding rectangle to expand this first region. The expansion may occur in steps of a predetermined size, for example, in increments of one pixel.

[0051] In process 502, it is determined whether the expanded first region contains a graphical marker that delimits a row boundary or a column boundary. The graphical marker may include a line that delimits a row boundary or a column boundary. The line may generally have any suitable style (e.g., solid, dashed, patterned, colored, etc.). The graphical marker may also include a transition, such as a change in background color or texture, that conveys the row boundary or column boundary. More generally, the graphical marker may include any suitable type of discontinuity that conveys the presence of a boundary between rows or columns at a given location within the image. Various image processing techniques may be used to detect whether such a graphical marker is included within the expanded first region. Exemplary embodiments of a method for determining that the expanded first region contains a graphical marker are described below with reference to FIG. 6.

[0052] In process 503, it is determined whether an extended first region overlaps with a second region corresponding to at least a portion of a second content object. For example, the extended first region can overlap with the second region when the first content object and the second content object are in adjacent rows or columns and there is no graphical marker between the adjacent rows or columns. In these scenarios, the first and second regions can expand during process 501 and continue to increase in size until they overlap with each other. Therefore, comparing the first region with other identified regions within the image containing the second region can be performed to detect cells that do not have graphical markers defining their boundaries.

[0053] In process 504, in response to (a) determining in process 502 that the extended first region includes a graphical marker, or (b) determining in process 503 that the extended first region overlaps with the second region, the extended first region is identified as corresponding to the location of the first cell. In some embodiments, one or more processes of method 500 can be repeated until each boundary of the cell (e.g., the boundary between two rows and the boundary between two columns) is determined in a similar manner.

[0054] FIG. 6 is a simplified diagram of a method 600 for determining that a region includes a graphical marker according to some embodiments. According to some embodiments consistent with FIGS. 1-5, method 600 can be used to implement process 502 of method 500.

[0055] In process 601, a plurality of pixel positions corresponding to the edge of the extended first region are identified. For example, when the extended first region corresponds to an N×M rectangle, the plurality of pixel positions can include N pixels along the right or left edge of the extended first region, or M pixels along the top or bottom edge of the extended first region. According to some embodiments, the extended first region generally corresponds to the extended first region associated with process 501.

[0056] In process 602, for each pixel position, it is determined whether the pixel position is associated with a change in color or intensity along at least one expansion direction exceeding a first predetermined threshold. For example, if a plurality of pixel positions correspond to the left end of an N×M bounding rectangle, each pixel can be compared with its adjacent pixel on the right. During the comparison, the difference between the pixel and the adjacent pixel (e.g., intensity difference, color difference, etc.) can be calculated. The difference may be an absolute difference, a relative difference, or the like. Then, the difference is compared with the first predetermined threshold. The first predetermined threshold is preferably set to a value high enough to avoid false positives (e.g., erroneously detecting the boundaries of a row or column based on a gradual background gradient) and a value low enough to detect a dense type of graphical marker (e.g., a small but sudden transition of the background color between rows).

[0057] In process 603, it is determined whether the count of a plurality of pixel positions associated with a change in color or intensity, as determined in process 602, exceeds a second predetermined threshold. The count can correspond to an absolute count or a relative count of the number of pixels (e.g., a percentage of the total number of pixels). Some types of graphical markers may be continuous (e.g., a solid line), in which case each of the plurality of pixels is likely to be included in the count. However, other types of graphical markers may be discontinuous (e.g., a dashed line), and fewer pixels among the plurality of pixels are likely to be included in the count. Therefore, the second predetermined threshold is preferably set to a numerical value low enough to detect a discontinuous type of graphical marker without introducing false positives.

[0058] In process 604, in response to determining that the count of the plurality of pixel positions exceeds the second predetermined threshold, it is determined that the expanded first region includes a graphical marker. In making this determination, a method such as method 500 can proceed to identify the region as corresponding to the location of the cell, as described in process 504.

[0059] FIG. 7 is a simplified diagram of a method 700 for determining that cells belong to at least one row and at least one column based on the location of the cells, according to some embodiments. According to some embodiments that are consistent with FIGS. 1-3, method 700 may be used to implement process 306 of method 300.

[0060] In process 701, a plurality of cells in a data table (e.g., the plurality of cells identified in process 305 of method 300) are sorted based on their identified locations. For example, the plurality of cells may be sorted in order along a column axis (e.g., from left to right or right to left) and a row axis (e.g., from top to bottom or bottom to top).

[0061] In process 702, one or more of the sorted cells belonging to at least one row are recursively identified. According to some embodiments, recursively identifying one or more cells belonging to at least one row may include recursively performing the following operations starting with a first selected cell: (1) determining a set of cells having a row span that overlaps the row span of the currently selected cell, (2) identifying the closest cell among the set of cells, (3) identifying the closest cell as belonging to at least one row, and (4) selecting the closest cell as the next selected cell. The row span corresponds to the range of positions occupied by cells along a row axis (e.g., the vertical axis of the data table). These operations may be performed from left to right (identifying the cell closest to the right of the selected cell) and from right to left (identifying the cell closest to the left of the selected cell) until each cell of at least one row is identified.

[0062] In process 703, the header row is identified based on one or more header content objects into which data is inserted into one or more header cells of the header row. The header content objects generally describe the content of the corresponding columns, for example, by providing an indicator or a label. Thus, a given column of the data table can be identified based on the corresponding header content object of that column. Subsequently, the corresponding individual cells within that column will have similar content objects that share a common characteristic or data type identified by the header content object of that column. An exemplary embodiment of a method for identifying the header row is described below with reference to FIG. 8.

[0063] In process 704, it is determined that the column span of the first cell overlaps with the column span of at least one first header cell among one or more header cells. The column span corresponds to the range of positions occupied by the cell along the column axis (e.g., the horizontal axis of the data table). When the column spans of different cells overlap, the two cells are likely to belong to the same column. In the case of a merged cell, the column span of the first cell can overlap with multiple header cells.

[0064] In process 705, the first cell is identified as belonging to at least one first column, and at least one first column is associated with at least one first header cell. However, when there is no header row, or when there is no header cell for at least one first column, an alternative approach can be used. For example, at least one first column can be assigned to a default header, such as a placeholder header text without features, when there is no header cell.

[0065] Furthermore, an identifier or label for at least one first column can be predicted and assigned based on semantic analysis of the content of the at least one first column. For example, when there is no header cell for at least one first column, the text contained within the cells of the column can be extracted and analyzed using an entity extraction engine to determine the type of entity contained within the column. In some embodiments, the entity extraction engine can associate the text contained within the cells based on the type of entity, without having a header cell to provide context within the data structure. For example, among other things, the techniques for identifying entity types disclosed in U.S. Patent No. 10,360,507 can be used for this purpose. Then, an identifier or label for the column can be assigned based on the type of entity within the column, and the structured database can have data inserted based on the identifier or label. The identifier or label can be in the form of a qualitative label, such as, among other things, annotating the type of drug, target, disease, mechanism of action, or phase of the trial. In preparation for sending the text to the entity extraction engine, various preprocessing steps can be applied to the text. For example, the text can be sent to a spell correction engine to correct spelling mistakes or irregularly spelled text in preparation for the entity extraction engine. Exemplarily, in the context of pharmaceutical or biomedical applications, the spell correction engine can include a biomedical spell correction engine, and the entity extraction engine can include a biomedical entity extraction engine. Exemplary examples of entity types recognized by biomedical entity extraction can include, but are not limited to, genes, drugs, tissues, diseases, organic chemicals, companies, diagnostic procedures, and physiological functions.

[0066] Figure 8 is a simplified diagram of a method 800 for identifying a header row from among one or more rows of a data table, according to some embodiments. According to some embodiments consistent with FIGS. 1-7, method 800 can be used to implement process 702 of method 700.

[0067] In process 801, one or more text representations corresponding to one or more header content objects are generated. The one or more text representations may include a set of digital characters. According to some embodiments, optical character recognition (OCR) may be used to generate one or more text representations based on an image representation of a data table.

[0068] In process 802, each of the one or more text representations is queried against a header dictionary. This results in a score vector that includes one or more confidence scores corresponding to the one or more text representations. Each confidence score is based on the strength of the query. For example, each confidence score may be determined based on a Levenshtein distance that provides a mechanism for explaining errors and uncertainties (e.g., OCR errors) in previous process steps. Exemplarily, in the context of a pharmaceutical information table, the header dictionary includes entities corresponding to headers that are expected to be included within such a table, such as drug names, diseases / targets, mechanisms of action, phases, etc. In some embodiments, the header dictionary is created by subject matter experts ( "SMEs") to manually identify common entries expected within the data system. The header dictionary may be updated over time to account for new common entity types using either manual or automated text recognition systems. Similar to the above discussion, in some embodiments, semantic analysis may be performed based on information stored within a structured database using, for example, the techniques described in U.S. Patent No. 10,360,507, filed September 22, 2017, entitled "Systems, Methods, and Computer Readable Media for Visualization of Sematic Information and Inference of Temporal Signals Indicating Salient Associations Between Life Science Entities", the entirety of which is incorporated herein by reference.

[0069] In process 803, a row score is determined based on the score vector. The row score is an aggregated metric based on one or more confidence scores that make up the score vector. For example, the row score can be calculated as the sum of the square roots of the score vector.

[0070] In process 804, a header row is selected based on the row score. For example, the row score of the header row can be compared to the row scores of other candidate rows in the data table. Candidate rows can include other rows having a row score greater than a predetermined threshold (e.g., zero). Rows containing a particular type of content object can be excluded from the set of candidate rows. For example, a row containing a graphical sequence object (e.g., a phase bar) may be ineligible to be selected as a header row. Rows can also be excluded as eligible header rows based on whether the row contains a particular content object, which can be defined by an SME using a list of excluded content objects that are ineligible for data insertion into the header row. The header row can then be selected in response to having the highest row score among the candidate rows.

[0071] Figures 9A - 9G are simplified diagrams of pharmaceutical information tables 900a - g according to some embodiments. In some embodiments that are consistent with FIGS. 1 - 8, pharmaceutical information tables 900a - g may correspond to data tables 131 - 139. As shown in FIGS. 9A - 9G, the visual and substantial differences shown between pharmaceutical information tables 900a - g reflect real - world differences in how pharmaceutical information may be distributed. Despite this wide variability in how information is distributed, system 100 and methods 200 - 800 may be configured to automatically parse and interpret pharmaceutical information tables 900a - g and insert data into a structured database such as database 180 based on the information within the tables. In some embodiments, information obtained from pharmaceutical production information tables 900a - g may include columns or rows for a destination website or URL, a drug or development program name, a target population or disease for a trial, a mechanism of action for the test drug, a phase number, a date, or phase information values for a trial such as a sequence, and other information that can be interpreted from pharmaceutical information tables 900a - g.

[0072] Each of the pharmaceutical information tables 900a - g is arranged within rows and columns. Table 900g includes a graphical row marker 912 for separating adjacent rows, table 900c includes a graphical column marker 914 for separating adjacent columns, and tables 900b, 900c, 900e, and 900f include both the graphical row marker 912 and the graphical column marker 914. As illustrated, the graphical markers can include solid lines that are lighter (e.g., table 900c) or darker (e.g., table 900e) than a background color, a sudden change in background color (e.g., table 900d), etc. Table 900a does not include graphical markers, and the other tables use graphical markers inconsistently, separating some rows or columns but not others. For example, table 900c includes a graphical column marker 914 for each row except the top row. In some embodiments, each individual row within a data table can be associated with a single drug or candidate topic that includes the above - described information types. Any additional information associated with the drug or candidate topic can be presented within the table using a separate column in the form of name - value pairs. For example, the name - value pair can derive the name from the header text associated with the identified column, and the value can be the context within an individual cell. In some embodiments, for example, if a column within a data table is identified as non - standard (i.e., not in the heading dictionary), the content of that data table column can be stored within a database table column titled "Other" that includes a list of name - value pairs. The format of the data is [{ 'name': 'column name 1', 'value': 'column value 1'}, { 'name': 'column name 2', 'value': 'column value 2'}]. In some embodiments, the database table can include names that do not exist within the created dictionary of entity types. Additional information can also be stored in the form of metadata associated with the data structuring, including, among other data types, data tables or technologies used to implement a study or a trial.

[0073] In addition, each of the pharmaceutical information tables 900a - g includes a plurality of progress bars 920 indicating the stage of development of a given drug candidate (e.g., discovery, pre - clinical, Phase I, Phase II, Phase III, etc.). The progress bars 920 are illustrated in various styles and include arrows or bars of various shapes and colors. Generally, the progress bars 920 can span multiple cells.

[0074] In some embodiments, a database table can be created by defining a new classification entity, which means defining columns or rows of a data table that have not been previously created as part of a manual heading dictionary. In this way, the system can create a database - structured data entry template based on the recognition of classification entities within a newly identified image representation.

[0075] Figures 10A - 10D are simplified diagrams of pharmaceutical information tables 1000a - d in which a logical group 1010 of text has been automatically identified, according to some embodiments. According to some embodiments consistent with Figures 1 - 8, the location of the logical group 1010 can be identified using a neural network model as described above, for example, with reference to processes 302 - 304. In particular, the annotations in tables 1000a - f correspond to the output generated in process 304 of method 300. The depiction of the pharmaceutical information tables 1000a - d is determined using an experimental system having features consistent with system 100, and the experimental system is configured to implement a method consistent with method 300. The neural network model used to identify the location of the content object corresponds to the YOLOv3 neural network model. Each of the logical groups 1010 is shown with a bounding rectangle (dashed line) around the text.

[0076] Figures 11A - 11F are simplified diagrams of pharmaceutical information tables 1100a - f in which cells are identified as belonging to specific rows and columns. The depiction of pharmaceutical information tables 1100a - f was generated by the experimental system described above with reference to FIG. 10. In particular, the annotations in tables 1100a - f correspond to the output generated in process 305 of method 300. The location of a cell is identified by box 1110 (dashed line), and the location of the progress bar is identified by box 1120 (dashed line). According to some embodiments consistent with FIGS. 1 - 8, boxes 1110 and 1120 can be identified using one or more of process 305, method 500, and / or method 600. Arrow 1130 (dashed line) connects cells identified as belonging to the header row. Arrow 1140 (solid line) connects cells identified as belonging to a given column. Arrow 1150 (dotted line) connects cells identified as belonging to a given row and the phase bar. According to some embodiments consistent with FIGS. 1 - 8, arrows 1140 - 1150 can be identified using process 306, method 700, and / or method 800. In table 1100b, since the first column does not contain a header content object, default header cell 1160 is assigned to the first column.

[0077] The subject matter described in this specification can be implemented in a digital electronic circuit, or in computer software, firmware, hardware, or in combinations thereof, including the structural means disclosed herein and their structural equivalents. The subject matter described in this specification can be implemented as one or more computer program products, such as one or more computer program products tangibly embodied in an information carrier (e.g., in a machine-readable storage device) or embodied in a propagated signal for execution by, or to control the operation of, a data processing apparatus (e.g., a programmable processor, a computer, or multiple computers). A computer program (also known as a program, software, software application, or code) can be written in any form of programming language, including a compiled or interpreted language, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file. The program can be stored in a portion of a file that holds other programs or data, in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program can be executed on one computer or on multiple computers at one site, or it can be distributed across multiple sites and deployed to be interconnected by a communication network.

[0078] The processes and logical flows described herein, including the method steps of the subject matter described herein, can be implemented by one or more programmable processors executing one or more computer programs to perform the functions of the subject matter described herein by operating on input data and generating output. The processes and logical flows can also be implemented by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus of the subject matter described herein can be implemented as special purpose logic circuitry, such as an FPGA or ASIC.

[0079] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, such as magnetic, magneto optical disks, or optical disks. Information carriers suitable for embodying computer program instructions and data include all forms of nonvolatile memory, including by way of example semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto optical disks, and optical disks (e.g., CD and DVD disks). The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0080] To provide for interaction with a user, the subject matter described herein may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user may provide input to the computer. Other types of devices may also be used to provide for interaction with the user. For example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form including acoustic, speech, or tactile input.

[0081] The subject matter described herein may be implemented in a computing system that includes back-end components (e.g., data servers), middleware components (e.g., application servers), or front-end components (e.g., client computers having a graphical user interface or a web browser by which a user may interact with an implementation of the subject matter described herein), or any combination of such back-end, middleware, and front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), such as the Internet.

[0082] It should be understood that the disclosed subject matter is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the drawings. The disclosed subject matter is capable of other embodiments and of being practiced and carried out in various ways. Also, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting.

[0083] Accordingly, those skilled in the art will recognize that the underlying concepts of the present disclosure can be readily utilized as a basis for the design of other structures, methods, and systems for carrying out some of the purposes of the disclosed subject matter. Accordingly, it is important that the claims be regarded as including such equivalent constructions insofar as they do not depart from the spirit and scope of the disclosed subject matter.

[0084] The disclosed subject matter has been described and illustrated in the foregoing exemplary embodiments, but the present disclosure is made only by way of example, and many changes in the details of the implementation of the disclosed subject matter may be made without departing from the spirit and scope of the disclosed subject matter, which is limited only by the following claims.

Claims

Claim 1 A method comprising: accessing, by one or more computer processors, an image representation of a data table, the data table including one or more cells arranged in one or more rows and one or more columns, the one or more cells including a first cell belonging to at least one first row and at least one first column, and a first content object being inserted into the first cell; providing, by the one or more computer processors, the image representation as an input to a neural network model trained to identify locations of content objects within the image representation; executing, by the one or more computer processors, the neural network model to identify a location of the first content object within the image representation; identifying, by the one or more computer processors, a location of the first cell based on the location of the first content object; determining, by the one or more computer processors, based on the one or more locations of the first cell and the first content object with respect to a plurality of content objects associated with the one or more rows and the one or more columns, that the first cell belongs to the at least one first row and the first column; associating, by the one or more computer processors, the first content object with one or more classification identifiers based on semantic analysis of the first content object; inserting, by the one or more computer processors, information associated with the first content object and information associated with the one or more classification identifiers into a structured database based on determining that the first cell belongs to the at least one first row and the at least one first column, the structured database including at least one data table row associated with the at least one first row and at least one data table column associated with the at least one first column; The location of the first content object includes a first region corresponding to at least a portion of the first content object, identifying the location of the first cell based on the location of the first content object, extending the first region in at least one direction, determining that the extended first region includes a graphical marker that marks one or more of the row boundaries and column boundaries, identifying the extended first region as corresponding to the location of the first cell in response to determining that the extended first region includes the graphical marker, the method comprising: **Claim 2** creating one of at least one second column and at least one second row in the structured database based on determining that the first cell does not belong to the at least one first row and the at least one first column, further comprising the method according to claim 1. **Claim 3** accessing the image representation, receiving, by the one or more computer processors, a digital document via a computer network, the digital document including the data table, rendering, by the one or more computer processors, the digital document as a digital image, locating, by the one or more computer processors, the image representation of the data table in the rendered digital image, the method according to claim 1. **Claim 4** determining that the extended first region includes the graphical marker, identifying a plurality of pixel positions corresponding to the edge of the extended first region, for each pixel position within the plurality of pixel positions, determining whether the pixel position is associated with one or more changes in color and intensity along the at least one direction exceeding a first predetermined threshold, determining that the count of the plurality of pixel positions associated with the change in color or intensity exceeds a second predetermined threshold. Determining that the extended first region includes the graphical marker in response to determining that the count of the plurality of pixel positions exceeds the second predetermined threshold, the method according to claim 1, comprising:

5. The location of the first cell includes a row span corresponding to a range of positions occupied by cells along the row axis and a column span corresponding to a range of positions occupied by cells along the column axis, Based on the location of the first cell, it is determined that the first cell belongs to the at least one first row and the at least one first column, Sorting the one or more cells in the data table based on the plurality of locations of the one or more cells, Starting from a selected cell among the one or more cells, recursively performing an operation for identifying one or more second cells belonging to the first row, the operation comprising: Determining at least one other cell having a row span overlapping the row span of the selected cell, Identifying the cell among the at least one other cell that is closest to the selected cell, Identifying the closest cell as belonging to the at least one first row, Selecting the closest cell as the next selected cell, Identifying the header row among the one or more rows of the data table based on one or more header content objects in which data is inserted into one or more header cells of the header row, Determining that the column span of the first cell overlaps the column span of a first header cell among the one or more header cells, Identifying the first cell as belonging to the first column, the first column being associated with the first header cell, Identifying the header row from among the one or more rows of the data table, Generating one or more text representations corresponding to the one or more header content objects by optical character recognition, Collating each of the one or more text representations with a header dictionary, resulting in a score vector including one or more confidence scores corresponding to the one or more text representations, each confidence score being based on the degree of the collation, Determining a row score based on the score vector, selecting the header row based on the row score, The method according to claim 1, wherein determining a row score based on the score vector includes calculating a sum of one or more square roots of the score vector and one or more of the one or more confidence scores. **Claim 6** selecting the header row includes comparing a plurality of row scores associated with the one or more rows of the data table with each other, selecting, as the header row, the row having the largest row score, the method according to claim 5. **Claim 7** obtaining, by the one or more computer processors, a list of excluded content objects that are not eligible to be part of the header row, determining, by the one or more computer processors, whether the one or more rows of the data table include the excluded content object, further comprising identifying, by the one or more computer processors, the header row from the one or more rows of the data table with the excluded rows that include the excluded content object, the method according to claim 5. **Claim 8** the first content object includes a graphical sequence object in which sequence information representing progress or a phase is illustrated using a progress bar, inserting data into the structured database includes extracting the sequence information from the graphical sequence object, the method according to claim 1, wherein the information associated with the first content object includes the sequence information. **Claim 9** A computing system for inserting data into a structured data set, accessing an image representation of a data table, the data table including one or more cells arranged in one or more rows and one or more columns, the one or more cells including a first cell belonging to at least one first row and at least one first column, and a first content object being inserted into the first cell, providing the image representation as an input to a neural network model trained to identify the location of content objects within the image representation Execute the neural network model to identify the location of the first content object within the image representation; Identify the location of the first cell based on the location of the first content object; Based on the locations of the first cell and the first content object with respect to the one or more rows and the one or more columns, determine that the first cell belongs to the at least one first row and the first column; Associate the first content object with one or more classification identifiers based on semantic analysis of the first content object; Based on determining that the first cell belongs to the at least one first row and the at least one first column, insert data into a structured database for the information associated with the first content object and the information associated with the one or more classification identifiers, wherein the structured database includes at least one data table row associated with the at least one first row and at least one data table column associated with the at least one first column; The location of the first content object includes a first region corresponding to at least a portion of the first content object; Identifying the location of the first cell based on the location of the first content object includes: Expanding the first region in at least one direction; Determining that the expanded first region includes a graphical marker marking one of a row boundary and a column boundary; In response to determining that the expanded first region includes the graphical marker, identifying the expanded first region as corresponding to the location of the first cell. Claim 10 The processor is The computing system according to claim 9, further configured to create one of at least one second column and at least one second row in the structured database based on determining that the first cell does not belong to the at least one first row and the at least one first column. **Claim 11** Accessing the image representation is Receiving, by the one or more computer processors, a digital document via a computer network, the digital document including the data table; Rendering, by the one or more computer processors, the digital document as a digital image; Locating, by the one or more computer processors, the image representation of the data table in the rendered digital image, the computing system according to claim 9. **Claim 12** Determining that the extended first region includes the graphical marker is Identifying a plurality of pixel positions corresponding to an edge of the extended first region; For each pixel position within the plurality of pixel positions, determining whether the pixel position is associated with one or more changes in color and intensity along at least one direction exceeding a first predetermined threshold; Determining that a count of the plurality of pixel positions associated with the change in color or intensity exceeds a second predetermined threshold; Determining that the extended first region includes the graphical marker in response to determining that the count of the plurality of pixel positions exceeds the second predetermined threshold, the computing system according to claim 9. **Claim 13** The location of the first cell includes a row span corresponding to a range of positions occupied by cells along the row axis and a column span corresponding to a range of positions occupied by cells along the column axis, Based on the location of the first cell, determining that the first cell belongs to the at least one first row and the at least one first column; Sorting the one or more cells in the data table based on the plurality of locations of the one or more cells. Starting from a selected cell among the one or more cells, recursively performing an operation for identifying one or more second cells belonging to the first row, wherein the one or more second cells include the first cell, and recursively performing the operation, and the operation includes determining at least one other cell having a row span overlapping with the row span of the selected cell; identifying the cell closest to the first cell among the one or more cells; identifying the closest cell as belonging to the at least one first row; selecting the closest cell as the next selected cell; identifying the header row among the one or more rows of the data table based on one or more header content objects in which data is inserted into one or more header cells of the header row; determining that the column span of the first cell overlaps with the column span of a first header cell among the one or more header cells; identifying the first cell as belonging to the first column, wherein the first column is associated with the first header cell; identifying the header row from among the one or more rows of the data table; generating, by optical character recognition, one or more text representations corresponding to the one or more header content objects; matching each of the one or more text representations against a header dictionary, resulting in a score vector including one or more confidence scores corresponding to the one or more text representations, each confidence score being based on the degree of the matching; determining a row score based on the score vector; selecting the header row based on the row score, including The computing system according to claim 9, wherein determining a row score based on the score vector includes calculating a sum of one or more square roots of the score vector and one or more of the confidence scores.

14. Selecting the header row includes comparing a plurality of row scores associated with the one or more rows of the data table with each other; selecting, as the header row, the row having the largest row score, the computing system according to claim 13.

15. obtaining, by the one or more computer processors, a list of excluded content objects that are not eligible to be part of the header row; determining, by the one or more computer processors, whether the one or more rows of the data table include the excluded content objects; identifying, by the one or more computer processors, the header row from the one or more rows of the data table with the rows including the excluded content objects excluded, the computing system of claim 13 further comprising.

16. the first content object includes a graphical sequence object in which sequence information representing progress or a phase is illustrated using a progress bar; inserting data into the structured database includes extracting sequence information from the graphical sequence object; the information associated with the first content object includes the sequence information, the computing system of claim 9.

Citation Information

Patent Citations

  • Method for recognizing table

    JP2001331763A

  • Low-resolution OCR for document acquired by camera

    JP2005346707A