A framework for extracting information from text layout.
An image-based method using OCR and neural networks addresses the inefficiencies of conventional data extraction by accurately organizing information into searchable and logically segmented output files, enhancing database and knowledge base integration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BRISTOL MYERS SQUIBB CO
- Filing Date
- 2022-04-20
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional systems struggle to accurately and efficiently extract data from diverse and varied file formats, such as texts, images, and charts, for seamless integration into databases and knowledge bases, making the process error-prone and time-consuming.
An image-based process utilizing multiple models, including optical character recognition (OCR) and neural networks, to extract and organize information with precise coordinates, generating an output file that is searchable and logically segmented, excluding irrelevant content like headers and footers.
The solution ensures accurate extraction and organization of information into logical sections, enabling efficient searching and knowledge extraction, improving the integration of data into databases and knowledge bases.
Smart Images

Figure 0007855086000001 
Figure 0007855086000002 
Figure 0007855086000003
Abstract
Description
Background Art
[0001] The extraction of information is an important aspect of creating a searchable knowledge base or database. Furthermore, information extraction and knowledge base creation are the ability to understand the data in a file and the ability to extract information therefrom. Information can be extracted from files such as texts, images, charts, graphs, etc. Files can also be in various formats and have various arrangements. As a result, it may be difficult to accurately extract data files. Furthermore, it may be difficult to mine files on a large scale in search of automatically processable information. Furthermore, conventional systems cannot extract data from files in the same way that a human reads a file.
Summary of the Invention
[0002] Provided herein are embodiments of a system, apparatus, device, method, and / or computer program product for extracting information from a file, and / or combinations thereof, alternative combinations thereof.
[0003] A particular embodiment includes a method for extracting information. The method includes receiving a first-form file having information and a plurality of regions of interest (ROIs), and converting the file into an image. The method further includes generating a first output using a first model, having a first set of information extracted from the image and a first set of coordinates of the first set of information in the image. The method further includes generating a second output using a second model, having a second set of coordinates of each of the plurality of ROIs in the image. The method further includes generating a third output using a third model, having a second set of information extracted from the image and a third set of coordinates of the second set of information in the image. The method further includes combining the first output and the third output to generate the information and a plurality of coordinates contained in the file, the plurality of coordinates having the coordinates of the information in the image. The method also includes generating a second-form output file having a plurality of sections, each section of which corresponds to an ROI of the plurality of ROIs, and each section of which is included in the output file based on the coordinates of the ROIs in the second set of coordinates corresponding to each section. The method includes inputting the portion of the information, determined to correspond to each section, into each of the multiple sections of the output file, based on the coordinates corresponding to the portion of the information and the coordinates of each section. The second format makes the information in the output file searchable while it is displayed on a graphical user interface (GUI) or stored in a data storage device.
[0004] In some embodiments, the information has one or more words, and generating the first output and the third output involves generating bounding boxes that surround the one or more words in the image.
[0005] In some embodiments, the third coordinate set forms a bounding box surrounding each of the plurality of ROIs.
[0006] In some embodiments, the first output is generated using optical character recognition (OCR).
[0007] In some embodiments, the third output is generated using a neural network.
[0008] In some embodiments, the output file is used by one or more machine learning models to generate a knowledge base.
[0009] In some embodiments, the information is selectable within the output file.
[0010] In some embodiments, the method further includes displaying the output file and the image.
[0011] In some embodiments, the information comprises words and / or images. Generating the information by combining the first output with the third output may include retaining the images contained in the first output. The method may include identifying one or more words in the first output that share the same coordinates as one or more words in the third output, and determining the similarity level between the one or more words in the first output and the one or more words in the third output.
[0012] The method may further include assigning a first priority value to the first output and a second priority value to the third output. The method may further include including one or more words from the third output of the plurality of words based on the second priority value of the third output and on the similarity level being greater than a predetermined threshold. The method may further include excluding one or more words from the first output of the plurality of words based on the first priority value of the first output and on the similarity level being greater than a predetermined threshold. The method may further include identifying a first data format for one or more words in the first output and a second data format for one or more words in the second output. The method may further include including one or more words from the first output of the plurality of words based on the first data format for one or more words in the first output. The method may further include excluding one or more words from the third output of the plurality of words based on the second data format for one or more words in the third output.
[0013] Other specific embodiments are systems for extracting information. The system includes memory, the memory having instructions and a processor coupled to the memory. The processor is configured to execute the instructions, which, when executed, cause the processor to receive a file of a first form having information and a plurality of regions of interest (ROIs), and to convert the file into an image. When executed, the instructions cause the processor to generate a first output using a first model, having a first set of information extracted from the image and a first set of coordinates of the first set of information in the image. When executed, the instructions further cause the processor to generate a second output using a second model, having a second set of coordinates of each of the plurality of ROIs in the image. When executed, the instructions further cause the processor to generate a third output using a third model, having a second set of information extracted from the image and a third set of coordinates of the second set of information in the image. When the instruction is executed, it further causes the processor to combine the first output and the third output to generate the information contained in the file and a plurality of coordinates. The plurality of coordinates have coordinates of the information in the image. When the instruction is executed, it further causes the processor to use the second output to generate a second-format output file having a plurality of sections. Each of the plurality of sections corresponds to an ROI of a plurality of ROIs, and each of the plurality of sections is included in the output file based on the coordinates of the ROI in the second coordinate set corresponding to each section. When the instruction is executed, it further causes the processor to input the portion of the information determined to correspond to each section into each of the plurality of sections in the output file, based on the coordinates corresponding to the portion of the information and the coordinates of each section. The second format makes the information in the output file searchable while it is displayed on a graphical user interface (GUI) or stored in a data storage device.
[0014] In some embodiments, the information has one or more words, and generating the first output or the third output comprises generating bounding boxes surrounding each of the one or more words in the image.
[0015] In some embodiments, the third coordinate set forms a bounding box surrounding each of the plurality of ROIs.
[0016] In some embodiments, the first output is generated using optical character recognition (OCR).
[0017] In some embodiments, the third output is generated using a neural network.
[0018] In some embodiments, the output file is used by one or more machine learning models to generate a knowledge base.
[0019] In some embodiments, the information is selectable within the output file.
[0020] In some embodiments, when the instruction is executed, it further causes the processor to display the output file and the image.
[0021] In some embodiments, the information comprises words and / or images. Generating the information by combining the first output with the third output includes holding the images contained in the first output. When the instruction is executed, the processor may further cause the processor to identify one or more words that share the same coordinates as one or more words in the third output and to determine the similarity level between the one or more words in the first output and the one or more words in the third output.
[0022] In some embodiments, when the instruction is executed, the processor may further cause the processor to assign a first priority value to the first output and a second priority value to the third output. When the instruction is executed, the processor may further cause the processor to include one or more words from the third output of the plurality of words based on the second priority value of the third output and on the similarity level being greater than a predetermined threshold. When the instruction is executed, the processor may further cause the processor to exclude one or more words from the first output of the plurality of words based on the first priority value of the first output and on the similarity level being greater than a predetermined threshold. When the instruction is executed, the processor may further cause the processor to identify a first data format for one or more words in the first output and a second data format for one or more words in the second output. When the instruction is executed, the processor may further cause the processor to include the one or more words from the first output of the multiple words based on the first data format of the one or more words in the first output. When the instruction is executed, the processor may further cause the processor to exclude the one or more words from the third output of the multiple words based on the second data format of the one or more words in the third output.
[0023] Other specific embodiments include a non-temporary machine-readable medium that stores instructions and, when executed by at least one computing device, causes at least one computing device to perform an operation: the operation comprises receiving a file of a first form having information and regions of interest (ROIs), and converting the file into an image. The operation further comprises, using a first model, generating a first output having a first set of information extracted from the image and a first set of coordinates of the first set of information in the image. The operation further comprises, using a second model, generating a second output having a second set of coordinates of each of a plurality of ROIs in the image. The operation further comprises, using a third model, generating a third output having a second set of information extracted from the image and a third set of coordinates of the second set of information in the image. The operation further comprises, combining the first output and the third output to generate the information contained in the file and a plurality of coordinates, the plurality of coordinates having the coordinates of the information in the image. The operation further comprises generating a second-format output file having multiple sections using the second output. Each of the multiple sections corresponds to an ROI of the multiple ROIs, and each section of the multiple sections is included in the output file based on the coordinates of the ROIs in the second coordinate set corresponding to each section. The operation further comprises inputting the portion of the information determined to correspond to each section into each of the multiple sections in the output file, based on the coordinates corresponding to the portion of the information and the coordinates of each section. The second format makes the information in the output file searchable while it is displayed in a graphical user interface (GUI) or stored in a data storage device.
[0024] In some embodiments, the information comprises one or more words, and generating the first output or the third output comprises generating bounding boxes surrounding each of the one or more words in the image.
[0025] In some embodiments, the third coordinate set forms a bounding box surrounding each of the plurality of ROIs.
[0026] In some embodiments, the first output is generated using optical character recognition (OCR).
[0027] In some embodiments, the third output is generated using a neural network.
[0028] In some embodiments, the output file is used by one or more machine learning models to generate a knowledge base.
[0029] In some embodiments, the information is selectable within the output file.
[0030] In some embodiments, the operation further comprises displaying the output file and the image.
[0031] In some embodiments, the information comprises words and / or images. Generating the information by synthesizing the first output and the third output includes retaining the images included in the first output. The operation may further comprise identifying one or more words in the first output that share the same coordinates as one or more words in the third output, and determining a similarity level between the one or more words in the first output and the one or more words in the third output.
[0032] In some embodiments, the operation may further comprise assigning a first priority value to the first output and a second priority value to the third output. The operation may further comprise including one or more words from the third output of the plurality of words based on the second priority value of the third output and on the similarity level being greater than a predetermined threshold. The operation may further comprise excluding one or more words from the first output of the plurality of words based on the first priority value of the first output and on the similarity level being greater than a predetermined threshold. The operation may further comprise identifying a first data format for one or more words in the first output and a second data format for one or more words in the second output. The operation may further comprise including one or more words from the first output of the plurality of words based on the first data format for one or more words in the first output. The operation may further include removing one or more words from the third output among the multiple words based on the second data format of the one or more words in the third output.
[0033] The accompanying drawings are incorporated herein by reference and constitute part of this specification, illustrating this disclosure, further illustrating the intent of this disclosure together with this specification, and enabling those skilled in the relevant art to manufacture and utilize this disclosure. [Brief explanation of the drawing]
[0034] [Figure 1] Figure 1 is a block diagram of a system for extracting information from a file, according to several embodiments. [Figure 2] Figure 2 is a block diagram of the data flow in a system for extracting information from a file, according to several embodiments. [Figure 3] Figure 3 shows examples of images with multiple sections according to several embodiments. [Figure 4]Figure 4 shows bounding boxes surrounding the first section of an image in several embodiments. [Figure 5] Figure 5 shows output files from several embodiments. [Figure 6] Figure 6 shows output files displayed in a graphical user interface (GUI) according to several embodiments. [Figure 7] Figure 7 is a flowchart showing the process of extracting information from a file according to several embodiments. [Figure 8] Figure 8 is a block diagram of an example of the device components according to several embodiments. [Modes for carrying out the invention]
[0035] The first appearance of an element in a drawing is generally indicated by the leftmost digit or the number corresponding to its reference number. In a drawing, similar reference numbers may also indicate identical or functionally similar elements.
[0036] Provided herein are embodiments of systems, apparatuses, devices, methods, and / or computer program products for extracting information from files, and / or combinations thereof, or alternative combinations thereof.
[0037] As mentioned above, accurately extracting information from files can be an error-prone and time-consuming process. For example, creating a database or knowledge base may require extracting and understanding information from files that can be hundreds or even thousands of pages long. Files may contain text, images, charts, graphs, etc. Furthermore, files can be in different formats and in diverse and varied arrangements. In this respect, traditional systems have not been able to accurately and efficiently extract data in a way that allows for seamless integration into databases and knowledge bases.
[0038] The embodiments described herein address these challenges using an image-based process, synthesizing the outputs of various models, which have information extracted from files along with location data and annotated ROIs, into an output file that can be imported into a database or knowledge base. This ensures that the output file contains information that is accurately mapped to the original file. Furthermore, the information in the output file is searchable while it is displayed in a GUI or stored in a data storage device. The information in the output file may also be used to extract knowledge from the information using machine learning models that can also be used to input into a database or ontology. Thus, the embodiments described herein provide an automated method for accurately extracting information from files for efficient access.
[0039] In a particular embodiment, the server receives a file of a first format containing information and multiple regions of interest (ROIs). The server converts the file into an image. Furthermore, the server generates a first output using a first model, which contains a first set of information extracted from the image and a first set of coordinates for the first set of information within the image. The server generates a second output using a second model, which contains a second set of coordinates for each of the multiple ROIs within the image. Furthermore, the server generates a third output using a third model, which contains a second set of information extracted from the image and a third set of coordinates for the second set of information within the image. The server further combines the first and third outputs to generate information and multiple coordinates contained in the file. The multiple coordinates contain the coordinates of the information within the image. The server uses the second output to generate an output file of a second format containing multiple sections. Each section of the multiple sections corresponds to an ROI of the multiple ROIs, and each section of the multiple sections is contained in the output file based on the coordinates of the ROIs in the second set of coordinates corresponding to each section. Furthermore, the server inputs the portion of information corresponding to each section into each section of the output file, based on the coordinates corresponding to the portion of information and the coordinates of each section. The second format makes the information in the output file searchable while it is displayed in the graphical user interface (GUI) or stored in the data storage device.
[0040] The embodiments described herein generate an output file containing information (e.g., text, images, charts, graphs) that is accurately extracted from a file and segmented based on the ROI from which the information was extracted. The embodiments described herein improve the accuracy of the information extracted from the file by organizing the information within the logical sections extracted from the file without the need for existing organized metadata. Furthermore, the information is extracted in the intended order. Additionally, irrelevant information such as headers, footers, page numbers, copyrights, and table of contents may be excluded from the output file.
[0041] The output file may be displayed to the user. In addition, the output file may be imported into a database so that the information extracted from the file and its location within the file can be efficiently searched. Furthermore, other machine learning models may use the information in the output file to identify knowledge when generating a knowledge base.
[0042] Figure 1 is a block diagram of a system for extracting data from a file, according to several embodiments. The system comprises a server 100, a client device 110, and a data storage device 120. The devices of the system may be connected to a network. For example, the devices of the system may be connected via wired, wireless, or a combination of wired and wireless. In one example embodiment, one or more parts of the network may be an ad-hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), part of the internet, a public switched telephone network (PSTN), a mobile phone network, a wireless network, a WiFi network, a WiMAX network, or any other network format, or a combination of two or more such networks. Alternatively, the server 100, client device 110, and data storage device 120 may be a single physical or virtual machine.
[0043] In some embodiments, the server 100 and the data storage device 120 may belong to a cloud computing environment. In other embodiments, the server 100 may belong to a cloud computing environment while the data storage device 120 belongs outside the cloud computing environment. Furthermore, in other embodiments, the server 100 may belong outside the cloud computing environment while the data storage device 120 belongs to a cloud computing environment.
[0044] The data storage device 120 may be one or more databases configured to store structured data and / or unstructured data. The data storage device 120 may also store files from which information has been extracted, output files, training files for training a model that extracts page layouts, and the like.
[0045] In some embodiments, the user may use the client device 110 to transmit or upload files for information extraction to the server 100. Alternatively, the client device 110 may transmit a request to the server to extract information from one or more files stored in the database 120.
[0046] Server 100 may receive files from client device 110. Alternatively, server 100 may retrieve one or more files from database 120 as indicated in the request transmitted by client device 110.
[0047] Server 100 may include an extraction engine 102. The extraction engine 102 may convert the file into an image. The extraction engine 102 may implement one or more models to extract information from a specific file. Furthermore, the extraction engine 102 may synthesize the outputs from one or more models to generate an output file having the extracted information organized into logical sections extracted from the file. The output file may be stored in a database 120 so that the information within the output file can be searched. The extraction and generation of the output file will be described in more detail with reference to Figure 2.
[0048] Figure 2 is a block diagram of the data flow in a system for extracting information from a file, according to several embodiments. Figure 2 will be explained in relation to Figure 1.
[0049] In some embodiments, the server 100 may receive a file 200 (e.g., a document) and a request to extract information from the file 200. Alternatively, the server 100 may retrieve the file 200 from the database 120 in response to receiving a request from the client device 110 to extract information from the file 200. The file 200 may contain information such as text (e.g., words), images, charts, graphs, etc.
[0050] The extraction engine 102 may determine whether the file is in a common format. The common format may be, for example, a Portable Document Format (PDF), but is not limited to this. If the file is not in a common format, the extraction engine 102 may convert the file to a common format (e.g., PDF) and make it file 202. For example, the extraction engine 102 may use C# to convert a non-PDF file to PDF format. Hereafter in this specification, PDF will be referred to as the common format, but those skilled in the art will understand that other common formats may be used instead.
[0051] In addition, the extraction engine 102 may display / convert each page of the PDF file 202 into an image and use it as image 204. For example, the extraction engine 102 may implement PyMuPDF to display / convert each page of the PDF file 202 into an image. The extraction engine 102 also maintains a list of each page within the PDF file 202. This allows for the analysis of countless pages within the PDF file 202.
[0052] The extraction engine 202 may implement various models to extract information from image 204. For example, the extraction engine 202 may implement a text extraction optical character recognition (OCR) model 206, a page layout analysis model 208, and a PDF image reading model 210. The extraction engine 202 may additionally or in combination with several models to extract information from image 204. The extraction engine 202 may simultaneously or in parallel transmit image 204 to the text extraction OCR model 206, the page layout analysis model 208, and the PDF image reading model 210 to extract information from image 204.
[0053] The text extraction OCR model 206 may perform an OCR process on image 204 to extract text from image 204 and each word containing each position of text within image 204. The text extraction OCR model 206 may extract text and each position of text within image 204. For example, the text extraction OCR model 206 may generate bounding boxes surrounding each word within image 204. The text extraction OCR model 206 may identify the coordinates of the bounding boxes. For example, the text extraction OCR model 206 may implement one or more of the following, but is not limited to: AMAZON TEXTRACT, TESERACT, GOOGLE VISION, MICROSOFT AZURE READ API, etc. The text extraction OCR model 206 may generate a first output. The first output comprises information detected in image 204 (e.g., text) and a first set of coordinates identifying the positions of the information detected in image 204.
[0054] The page layout analysis model 208 may identify ROIs in image 204. ROIs may be sections in image 204, such as titles, summaries, chapters, paragraphs, tables, images, etc. The page layout analysis model 208 may implement a neural network to identify ROIs within image 204. The neural network may be a convolutional neural network (CNN), such as fast R-CNN, faster R-CNN, or Inception, but is not limited to these examples.
[0055] A neural network may be trained to identify ROIs in images with various page layouts. For example, a CNN may be trained in two phases: a forward phase and an inverse phase. The forward phase has (multiple) convolutional layers, (multiple) pooling layers, and a fully connected layer. To train the CNN, the extraction engine 102 may instruct the CNN to identify ROIs in input images of training files retrieved from the database 120.
[0056] A convolutional layer may apply filters to the input image to generate a feature map. In particular, the convolutional layer of the CNN may perform feature extraction from the input image. Features have parts of the input image. For example, features may be different edges or shapes of the input image. The CNN may extract different forms of features to generate different forms of feature maps. For example, the CNN may apply an array of numbers (e.g., a kernel) across different parts of the input image. A kernel is also called a filter. As mentioned above, different forms of filters may be applied to the input image to generate different feature maps. For example, a filter that identifies shapes in the input image may be different from a filter that detects edges. Therefore, a kernel different from the one that detects edges may be applied to identify shapes in the input image. Each kernel may have a different array of numbers. The values of the filters or kernels may be assigned randomly and optimized over time (e.g., using gradient descent). A kernel may be applied as a moving window across different parts of the input image. The kernel can be summed with a specific part of the input image to produce an output value. The output value may be included in the feature map. The feature map may contain output values from different kernels applied to different parts of the input image. The generated feature map may be a two-dimensional array.
[0057] A pooling layer may generate a reduced feature map. In particular, in the pooling layer, the CNN may reduce the dimension of each feature map generated in the convolutional layer. The CNN may extract a specific portion of the feature map and discard the rest. Image pooling retains important features and discards unimportant features. For example, a feature map may have activated and unactivated locations. Activated locations may contain detected features, while unactivated locations may indicate that a portion of the image does not contain features. Pooling may remove the unactivated locations. In this way, the size of the image is reduced. The CNN may perform these operations using maximum pooling or mean pooling in the pooling layer. Maximum pooling retains the larger values of a portion of the feature map and discards the remaining values. Mean pooling retains the average of different portions of the feature map. Thus, the CNN may generate a reduced feature map from each of the feature maps generated in the convolutional layer.
[0058] The CNN may have additional convolutional layers. In the additional convolutional layers, the CNN may generate additional feature maps based on the reduced feature maps generated in the pooling layers. Furthermore, the CNN may have additional pooling layers. In the additional pooling layers, the CNN may generate further reduced feature maps based on the feature maps generated in the additional convolutional layers.
[0059] The convolutional layer may apply a rectified linear unit (ReLU) function to the input image. The ReLU function is applied to the image to remove linearity. For example, the ReLU function may remove all black elements from the image, leaving only gray and white. This makes the colors in the input image change more abruptly and removes linearity from the input image.
[0060] Convolutional and pooling layers may be used for feature learning. Feature learning enables the CNN to identify necessary features in the input image and, as a result, accurately classify the input image. Consequently, by optimizing the convolutional and pooling layers, the CNN may apply appropriate filters to the input image and extract the features necessary to classify the input image.
[0061] The fully connected layer may then classify the image features using weights and biases to generate an output. The CNN may identify ROIs in the input image. The fully connected layer may generate bounding boxes surrounding the identified ROIs. Furthermore, the fully connected layer may identify the coordinates of the identified regions of interest in the input image.
[0062] In particular, in the fully connected layer, the CNN may flatten the reduced feature maps generated in the pooling layer into a one-dimensional array (or vector). The fully connected layer is a neural network. The CNN may perform a linear transformation on the one-dimensional array. The CNN may perform the linear transformation by applying weights and biases to the one-dimensional array. Initially, the weights and biases may be initialized randomly and optimized over time. The CNN may perform the linear transformation, such as a function of the activation layer (e.g., softmax or sigmoid), to identify the ROI of the input image.
[0063] In the reverse phase, the CNN may use backpropagation to determine whether it was able to accurately identify the ROI. Backpropagation may optimize the input parameters so that the CNN can classify the text more accurately. The input parameters may include values such as kernel, weights, and biases. Gradient descent may be used to optimize the parameters. In particular, gradient descent may be used to optimize the CNN's identification of ROIs in the input image.
[0064] Gradient descent is an iterative process for optimizing a CNN. Gradient descent may involve updating the CNN's parameters and having the CNN identify each ROI in the input image based on the updated parameters, thereby validating the identification of ROIs in the input image.
[0065] As a result, in some respects, the CNN may use backpropagation to validate the classification of the input images. In particular, in some respects, a subject expert (e.g., a pathologist) may determine whether model 104 correctly identified the ROIs in the input images. The subject expert may provide feedback on the accuracy of the identified ROIs. Instead of, or in addition to, feedback, the CNN may validate the identified ROIs using metadata associated with each image. For example, the metadata may be ROIs that were accurately identified for the input images.
[0066] The CNN may compare the ROIs identified by the CNN in the input image with information contained in metadata or feedback from subject experts. Based on the comparison, the CNN may update the filter values, weights, and biases using gradient descent and run the forward phase again on the input image.
[0067] The extraction engine 102 may instruct the CNN to identify the ROI of each image in a training file that has different versions of each image. The CNN iteratively optimizes its parameters, identifying the ROI of each image in the training file throughout the forward phase and verifying the ROI through the reverse phase, until it reaches a desired accuracy threshold.
[0068] Once the CNN reaches the required accuracy threshold, it may be considered sufficiently trained. Once the CNN is sufficiently trained, the page layout analysis model 208 may use the CNN to identify ROIs in image 204. The CNN may also identify and extract ROI identifiers in image 204. The page layout analysis model 208 may use the CNN to generate a second output having a second set of coordinates and ROI identifiers for each ROI in image 204. The CNN may be continuously improved after identifying ROIs in a particular image based on feedback received from the user.
[0069] In some embodiments, the page layout analysis model 208 may use a faster-R-CNN with image segmentation to detect and locate ROIs within image 204. ROIs may be tables, images, paragraphs, chapter titles, etc.
[0070] As an example, though not limited to specific cases, the PDF image reading model 210 extracts text from image 204 using PyMuPDF. The PDF image reading model 210 outputs bounding boxes surrounding each detected word in image 204. If the PDF image reading model 210 encounters non-word information in image 204, it extracts pixel information. The reading PDF image model 210 generates a third output having a second word set of words detected in image 204 and a third set of coordinates for each word in the third word set within image 204.
[0071] Furthermore, the extraction engine 102 may implement a composite analysis model 212. The composite analysis model 212 may be configured to generate an output file by combining the first, second, and third outputs. In particular, the first output from the text extraction OCR model 206 and the third output from the PDF image reading model 210 have text from image 204 and the coordinates of each part of the text within image 204. There may be some or considerable overlap between the first and third outputs. That is, the text extraction OCR model 206 and the PDF image reading model 210 extracted information from the same location within image 204. The composite analysis model 212 may need to determine which information should be retained and which information should be deleted.
[0072] For this purpose, the synthetic analysis model 212 may assign priorities to the information in the first and third outputs. For example, if the synthetic analysis model 212 determines that the text extraction OCR model 206 and the PDF image reading model 210 extracted text from the same location in image 204, the synthetic analysis model 212 may determine the similarity level between the text in the first output at a specific location in image 204 and the text in the third output at a specific location in image 204. For example, the synthetic analysis model 212 may determine how similar the words in the text are between the first output at a specific location in image 204 and the third output at a specific location in image 204.
[0073] Furthermore, the synthetic analysis model 212 may determine the data format of the extracted information in the first and second outputs. For example, if the similarity level is greater than a predetermined threshold and the text does not contain embedded fonts or images (e.g., the text is raw text), the synthetic analysis model 212 may retain the text from the third output (e.g., from the PDF image reading model 210) and delete the text from the first output (e.g., from the text extraction OCR model 206). If the text contains embedded fonts or images, the synthetic analysis model 212 may retain the text from the first output (e.g., from the text extraction OCR model 206) and delete the text from the third output (e.g., from the PDF image reading model 210). For extracted information such as images, charts, graphs, etc., the synthetic analysis model 212 may retain the information from the first output (e.g., from the text extraction OCR model 206) and delete the information from the third output (e.g., from the PDF image reading model 210). By combining the first and second outputs, the synthetic analysis model 212 may generate the complete contents of image 204 in the correct positions or in the desired order.
[0074] The synthetic analysis model 212 may generate an output file having sections. Each section may correspond to an identified ROI from the second output. The sections may be organized in the output file based on the coordinates and identifiers of the ROIs, as shown in the second output. For example, if image 204 has chapter 1 and chapter 2, the second output may have the coordinates of chapter 1 and chapter 2 in image 204. The synthetic analysis model 212 may use the coordinates of chapter 1 and chapter 2 in the second output to generate the first and second sections in the output file.
[0075] The synthetic analysis model 212 may input information extracted from image 204 into sections within the output file. That is, the synthetic analysis model 212 may input information determined to correspond to each section into each section of the output file, using the information from the first and second outputs synthesized by the synthetic analysis model 212, the coordinates of the information in image 204 as shown in the first and second outputs, and the coordinates of the ROI as shown in the third output.
[0076] In some embodiments, the first or third output may contain text from images, charts, and tables in image 204. The second output may contain coordinates of images, charts, and tables. The synthetic analysis model 212 may input text extracted from images, charts, and tables in the first or third output to sections in the output file corresponding to images, charts, and tables in the second output.
[0077] The output file may contain information extracted from image 204, organized into logical sections extracted from image 204. The output file may also contain coordinates of the information and logical sections mapped to image 204. Furthermore, the output file may contain symbols indicating spaces, paragraphs, tabs, indents, etc., within the information. In some embodiments, the text extraction OCR model 206, the page layout analysis model 208, and the PDF image reading model 210 may be configured to exclude headers, footers, table of contents, and other information that is not necessary for the user.
[0078] The output file may be a JavaScript Object Notation (JSON) file. Therefore, the output file may be human-readable. Furthermore, other machine learning models may use the output file for further processing. In particular, machine learning models may identify / extract knowledge from the output file and use the output file to build a knowledge base. The machine learning model may be one or more of the following: named entity recognition (NER) model 214, embedding model 216, summary generation model 218, categorization model 220, etc.
[0079] In some embodiments, the output file may be stored in a data storage device 120. The information in the output file may be queried. Also, since the output file contains the coordinates of information in the image 204, the user can query the output file to identify the information in the image 204.
[0080] Figure 3 shows an example of an image having multiple sections according to several embodiments. Figure 3 will be described with reference to Figures 1 and 2.
[0081] Image 300 may be a file converted into image 300 by the extraction engine 102. Image 300 may have text divided into two sections: a first section 304 and a second section 306. The first section 304 may be a section containing text about dogs. The second section 306 may be a section containing text about cats.
[0082] The extraction engine 102 may instruct the text extraction OCR model 206, the page layout model 208, and the PDF image reading model 210 to extract information from the image 300. In this scenario, the information may be text about dogs and cats.
[0083] The text extraction OCR model 206 and the PDF image reading model 210 may each detect and extract text from image 300. Furthermore, the text extraction OCR model 206 and the PDF image reading model 210 may each generate bounding boxes around the detected words. For example, the text extraction OCR model 206 and the PDF image reading model 210 may each generate a bounding box 302 around the word "jumped". Furthermore, the text extraction OCR model 206 and the PDF image reading model 210 may each determine the coordinates of the bounding box 302 in image 300. The coordinates are in the format (x,y), where x and y are pixels in image 300. The text extraction OCR model 206 and the PDF image reading model 210 may each extract the word "jumped" and the coordinates of the bounding box 302. The text extraction OCR model 206 and the PDF image reading model 210 may each include the extracted words and coordinates in their first and third outputs, respectively.
[0084] Figure 4 illustrates the bounding box surrounding the first section 304 of image 300 in several embodiments. Figure 4 will be described with reference to Figures 1 to 3.
[0085] As described above, the extraction engine 102 may instruct the page layout analysis model 208 to identify the ROI of image 300. The ROI may be either the first section 304 or the second section 306. The page layout analysis model 208 may use a neural network such as a CNN to identify the first section 304 and the second section 306 in image 300.
[0086] For example, the page layout analysis model 208 may, in response to identifying the first section 204, generate a bounding box 400 around the first section 304. The page layout analysis model 208 may determine the coordinates of the bounding box 400 within the image 300. Furthermore, the page layout analysis model 208 may determine an identifier for the first section 304 (e.g., "I. Dogs"). The page layout analysis model 208 may include the coordinates of the bounding box 400 and the identifier for the first section 304 in a second output.
[0087] Figure 5 illustrates output files according to several embodiments. Figure 5 will be explained with reference to Figures 1 to 4.
[0088] The synthetic analysis model 212 may generate an output file 500 using the first output of the text extraction OCR model 206, the second output generated by the page layout analysis model 208, and the third output generated by the PDF image reading model 210. In particular, the synthetic analysis model 212 may generate information contained in the image 300 by combining the extracted text with the coordinates of the bounding boxes of the text extracted by the first and second outputs.
[0089] Furthermore, the synthetic analysis model 212 may use a third output to generate sections 502 and 504 in the output file 500. Section 502 may correspond to the first section 304 in image 300. Section 504 may correspond to the second section 306 in image 300. The synthetic analysis model 212 may generate section 502 based on the bounding box of the first section 304 (e.g., bounding box 400) and the identifier of the first section 304, and section 504 based on the bounding box of the second section 306 and the identifier of the second section 306.
[0090] The synthetic analysis model 212 may identify a sentence from a synthesized sentence corresponding to section 502 based on the coordinates of the sentence and the coordinates of the first section 304. The synthetic analysis model 212 may be input with the sentence corresponding to section 502. Similarly, the synthetic analysis model 212 may identify a sentence from a synthesized sentence corresponding to section 504 based on the coordinates of the sentence and the coordinates of the second section 306. The synthetic analysis model 212 may be input with the sentence corresponding to section 504.
[0091] Figure 6 illustrates output files displayed in a graphical user interface (GUI) according to several embodiments. Figure 6 will be explained with reference to Figures 1 to 5.
[0092] The user may use the client device 110 to display the output file 500 on the GUI 600. The GUI 600 has parts 602 and 604. Part 602 may have an image 300, and part 604 may have the output file 500.
[0093] The text in output file 500 may be selectable and searchable. Furthermore, output file 500 may contain the coordinates of the text in image 300. Therefore, the text in output file 500 may be mapped to the relevant parts of image 300. If the user selects the word "jumped" in output file 500, the corresponding word 608 can be highlighted in image 300 based on the coordinates of word 608 "jumped" in image 300. In this way, image 300 may be effectively searched using output file 500.
[0094] In another example, image 300 may be displayed in GUI 600, while output file 500 may not be displayed in GUI 600. However, when a user attempts to search for image 300, the search may be performed on output file 500, and relevant elements of image 300 may be highlighted based on the search results.
[0095] Figure 7 is a flowchart illustrating the process of extracting information from a file according to several embodiments. Method 700 may be executed by processing logic and may comprise hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.) and software (e.g., instructions executed by a processing device), or a combination thereof.
[0096] Method 700 will be described with reference to Figures 1 and 2. However, Method 700 is not limited to the embodiments described herein.
[0097] In operation 702, the server 100 receives a file 200 containing information and multiple regions of interest (ROIs). The information may be text, images, graphs, charts, etc. The ROIs may be sections such as titles, summaries, chapters, sections, etc. The file may also be received from a client device 110 in response to a request to extract information from the file.
[0098] In operation 704, the extraction engine 102 converts the file to image 204. The extraction engine 102 may also determine whether file 200 is a PDF. If not, the extraction engine 102 may convert file 200 to a PDF file 202. Furthermore, the extraction engine 102 may convert PDF file 202 to image 204.
[0099] In operation 706, the extraction engine 102 generates a first output having a first information set extracted from image 204 and a first coordinate set of the first information set within image 204. The extraction engine 102 may also instruct the text extraction OCR model 206 to generate the first output. The text extraction OCR model 206 may use OCR to extract the first information set and the first coordinate set of the first information set. For example, the first coordinate set may be bounding boxes surrounding each word, each image, each chart, each graph, etc. The first coordinate set indicates the location of the information within image 204.
[0100] In operation 708, the extraction engine 102 generates a second output having a second set of coordinates for each ROI in image 204. The extraction engine 102 may instruct the page layout analysis model 208 to detect and identify the set of coordinates for each ROI. The page layout analysis model 208 may implement a neural network to detect and identify each ROI in image 204. The page layout analysis model 208 may create bounding boxes surrounding each ROI in image 204. The second set of coordinates may correspond to the bounding boxes in image 204. The second set of coordinates indicates the location of the ROI in image 204.
[0101] In operation 710, the extraction engine 102 generates a third set having a second information set extracted from the image and a third coordinate set of the second information set within the image. The extraction engine 102 may also instruct the PDF image reading model 210 to generate a third output. For example, the third coordinate set may be a bounding box surrounding each word, each image, each chart, each graph, etc.
[0102] In operation 712, the extraction engine 102 combines the first output and the third output to generate information contained in the file and multiple coordinates. The multiple coordinates represent the coordinates of the information in the image. The extraction engine 102 may also instruct the synthesis analysis model 212 to combine the first output and the third output to generate information.
[0103] In operation 714, the extraction engine 102 generates an output file containing sections. The synthetic analysis model 212 may also generate an output file. Each section corresponds to an ROI in image 204. Each section is included in the output file based on the coordinates of the ROI corresponding to each section.
[0104] In operation 716, the extraction engine 102 inputs the portion of information determined to correspond to each section into each section of the output file, based on the portion of information and the coordinates of each section. The synthetic analysis model 212 inputs the sections into the output file. The output file may contain information extracted from files and logically organized like a file. Furthermore, the output file may be searchable.
[0105] Various embodiments may be implemented using one or more computer systems, such as the computer system 800 shown in Figure 8. The computer system 800 can be used, for example, to implement the method 700 shown in Figure 7. Furthermore, the computer system 800 may comprise at least some of the server 100, client devices 110, and data storage devices 120, as shown in Figure 1. For example, the computer system 800 routes communications to various applications. The computer system 800 may be any computer capable of performing the functions disclosed herein.
[0106] The computer system 800 may be any well-known computer capable of performing the functions disclosed herein.
[0107] The computer system 800 has one or more processors (also called a central processing unit or CPU), such as processor 804. Processor 804 is connected to a communication interface or bus 806.
[0108] One or more processors 804 may each be a graphics processing unit (GPU). In embodiments, a GPU is a processor of special electronic circuits designed to process mathematically intensive applications. A GPU may have a parallel architecture that efficiently processes large data blocks, such as mathematically intensive data commonly found in computer graphics applications, images, videos, etc.
[0109] The computer system 800 also has user input / output devices 803 such as a monitor, keyboard, and pointing device, and communicates with the communication infrastructure 806 through the user input / output devices 802.
[0110] The computer system 800 also has main memory or first memory 808, such as random access memory (RAM). The main memory 808 may have one or more cache levels. The main memory 808 stores control logic (i.e., computer software) and data.
[0111] The computer system 800 may also have one or more secondary storage devices or secondary memory 810. The secondary memory 810 may have, for example, a hard disk drive 812 and / or a removable storage device 814. The removable storage device 814 is a floppy disk, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0112] The removable storage device 814 may interact with the removable storage unit 818. The removable storage unit 818 has a usable or readable storage device that stores computer software (control logic) and / or data. The removable storage unit 818 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, or any other computer data storage device. The removable storage device 814 reads from and / or writes to the removable storage unit 818 in a well-known manner.
[0113] According to the exemplary embodiment, the secondary memory 810 may be made accessible by the computer system 800 to computer programs and / or other instructions and / or data, including other means, devices, or other approaches. Such means, devices, or other approaches may include, for example, a removable storage unit 822 and interface 820. Examples of a removable storage unit 822 and interface 820 may include a program cartridge and cartridge interface (such as those found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other interface associated with a removable storage unit.
[0114] The computer system 800 may further have a communication or network interface 824. The communication interface 824 enables the computer system 800 to communicate with and interact with any combination of remote devices, remote networks, remote entities, etc. (referred to individually and collectively by reference numeral 828). For example, the communication interface 824 enables the computer system 800 to communicate with a remote device 828 via a communication circuit 826, which may be wired and / or wireless and may include any combination of LAN, WAN, Internet, etc. Control logic and / or data may be transmitted to and from the computer system 800 via the communication circuit 826.
[0115] In embodiments, a tangible non-temporary device or product having a usable or readable medium for a tangible non-temporary computer in which control logic (software) is stored is also referred to herein as a computer program product or program storage device. This includes, but is not limited to, a tangible product embodying a computer system 800, a main memory 808, a secondary memory 810, removable storage units 818 and 822, or any combination thereof. When such control logic is executed by one or more data processing devices (such as computer system 800), it causes such data processing devices to operate as disclosed herein.
[0116] Methods for creating and using embodiments of the disclosure using data processing devices, computer systems, and / or computer structures other than those shown in Figure 8, based on the teachings contained herein, will be obvious to those skilled in the art. In particular, embodiments may operate with software, hardware, and / or operating system implementations other than those disclosed herein.
[0117] It should be understood that the section of detailed description, rather than any other section, is intended for use in interpreting the claims. Other sections may describe, but not all, one or more exemplary embodiments intended by the inventors, and are therefore not intended to limit in any way the claims of this disclosure or the appended claims.
[0118] This disclosure describes exemplary embodiments for exemplary fields and uses, but it should be understood that this disclosure is not limited thereto. Other embodiments and modifications thereof are possible and are included within the scope and spirit of this disclosure. For example, without limiting the generality of this paragraph, embodiments are not limited to software, hardware, firmware, and / or entities shown in the figures and / or described herein. Furthermore, embodiments (whether expressly described herein or not) may have significant utility for fields and uses beyond the examples described herein.
[0119] This specification describes embodiments with the help of functional blocks that demonstrate the implementation of specific functions and their relationships. The boundaries of these functional blocks are arbitrarily defined in this specification for the sake of clarity. Alternative boundaries may be defined as long as the specific functions and relationships (or equivalents) are properly performed. Furthermore, alternative embodiments may execute functional blocks, steps, operations, methods, etc., in an order different from that described herein.
[0120] In this specification, references to “a certain embodiment,” “one embodiment,” “an exemplary embodiment,” or similar expressions indicate that the embodiments described may include certain features, structures, or characteristics, but not all embodiments may necessarily include certain features, structures, or characteristics. Furthermore, such expressions do not necessarily refer to the same embodiment. Moreover, where certain features, structures, or characteristics are described in relation to an embodiment, it would be within the knowledge of those skilled in the art to incorporate such features, structures, or characteristics into other embodiments, whether or not they are explicitly mentioned or described herein. Furthermore, some embodiments may be described using the expressions “connected” and “linked” and their derivatives. These terms are not necessarily intended to be synonymous. For example, some embodiments may be described using the terms “connected” and / or “linked” to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “linked” may also mean that two or more elements are not in direct contact with each other but are still cooperating or acting on each other.
[0121] The breadth and scope of this disclosure should not be limited by any of the exemplary embodiments described above, but should be defined solely in accordance with the following claims and their equivalents.
Claims
1. A method for extracting information, The processor receives a file of a first form, which contains information and multiple regions of interest (ROIs). The aforementioned processor converts the aforementioned file into an image. The processor generates a first output using the first model, which includes a first information set extracted from the image and a first coordinate set of the first information set within the image. The processor generates a second output using a second model, which has a second set of coordinates for each of the multiple ROIs in the image. The processor generates a third output using a third model, which includes a second information set extracted from the image and a third coordinate set of the second information set within the image. The processor combines the first output and the third output to generate a plurality of coordinates having the information contained in the file and the coordinates of the information in the image. The processor generates a second type of output file having multiple sections using the second output, where each section of the multiple sections corresponds to an ROI of the multiple ROIs, and the output file contains the coordinates of the ROIs in the second coordinate set corresponding to each section. The processor inputs the portion of the information that has been determined to correspond to each section into each of the multiple sections in the output file, based on the coordinates corresponding to the portion of the information and the coordinates of each section. Includes, The second format makes the information in the output file searchable while it is displayed on the graphical user interface (GUI) or stored in the data storage device. method.
2. The aforementioned information comprises one or more words, Generating the first output or the third output includes generating bounding boxes surrounding each of the one or more words in the image. The method according to claim 1.
3. The third coordinate set forms a bounding box surrounding each of the plurality of ROIs. The method according to claim 1.
4. The first output is generated using optical character recognition (OCR). The method according to claim 1.
5. The third output is generated using a neural network. The method according to claim 1.
6. One or more machine learning models generate a knowledge base using the output file. The method according to claim 1.
7. The aforementioned information is searchable within the output file. The method according to claim 1.
8. The processor further includes displaying the output file and the image. The method according to claim 1.
9. The aforementioned information includes words and / or images, The method according to claim 1.
10. Combining the first output with the third output to generate the information includes holding the image included in the first output. The method according to claim 9.
11. The processor identifies one or more words in the first output that share the same coordinates as one or more words in the third output. The processor determines the similarity level between one or more words in the first output and one or more words in the third output. Further including, The method according to claim 9.
12. The processor assigns a first priority value to the first output and a second priority value to the third output. The processor includes one or more words from the third output among the plurality of words based on the second priority value of the third output and on the similarity level being greater than a predetermined threshold, The processor removes one or more words from the first output among the multiple words based on the first priority value of the first output and on the similarity level being greater than a predetermined threshold. Further including, The method according to claim 11.
13. The processor identifies the first data format of one or more words in the first output and the second data format of one or more words in the third output. The processor includes, based on the first data format of the one or more words in the first output, the one or more words in the multiple words from the first output, The processor removes one or more words from the third output based on the second data format of one or more words in the third output, Further including, The method according to claim 11.
14. It is a system for extracting information, Memory containing instructions, A processor connected to the memory and configured to execute the instructions, Equipped with, When the instruction is executed, the processor will: To receive a first type of file containing information and multiple regions of interest (ROIs), Convert the aforementioned file into an image, Using the first model, generate a first output having a first information set extracted from the image and a first coordinate set of the first information set in the image. Using the second model, a second output is generated having a second set of coordinates for each of the multiple ROIs in the image. Using the third model, a third output is generated having a second information set extracted from the image and a third coordinate set of the second information set within the image. The first output and the third output are combined to generate the information contained in the file and the multiple coordinates having the coordinates of the information in the image. Using the second output, generate a second format output file having multiple sections, where each section of the multiple sections corresponds to an ROI of the multiple ROIs, and each section of the multiple sections is included in the output file based on the coordinates of the ROI in the second coordinate set corresponding to each section. Based on the coordinates corresponding to the portion of the information and the coordinates of each section, the portion of the information determined to correspond to each section is input to each of the multiple sections in the output file. Includes, The second format makes the information in the output file searchable while it is displayed on a graphical user interface (GUI) or stored in a data storage device. system.
15. The information comprises one or more words, and generating the first output or the third output comprises generating bounding boxes surrounding each of the one or more words in the image. The system according to claim 14.
16. The third coordinate set forms a bounding box surrounding each of the plurality of ROIs. The system according to claim 14.
17. The first output is generated using optical character recognition (OCR). The system according to claim 14.
18. The third output is generated using a neural network. The system according to claim 14.
19. The aforementioned output file is used by one or more machine learning models to generate a knowledge base. The system according to claim 14.
20. The aforementioned information is selectable within the output file. The system according to claim 14.
21. When the instruction is executed, it further causes the processor to display the output file and the image. The system according to claim 14.
22. The aforementioned information includes words and / or images. The system according to claim 14.
23. Combining the first output with the third output to generate the information includes holding the image included in the first output. The system according to claim 22.
24. When the aforementioned instruction is executed, the processor further: Identify one or more words in the first output that share the same coordinates as one or more words in the third output, To determine the similarity level between one or more words in the first output and one or more words in the third output, The system according to claim 22.
25. When the aforementioned instruction is executed, the processor further: Assign the first priority value to the first output and the second priority value to the third output. Based on the second priority value of the third output and on the similarity level being greater than a predetermined threshold, one or more words from the third output among the plurality of words are included. Based on the first priority value of the first output and on the similarity level being greater than a predetermined threshold, one or more words are excluded from the first output among the multiple words. The system according to claim 24.
26. When the aforementioned instruction is executed, the processor further: The first data format of one or more words in the first output and the second data format of one or more words in the third output are identified. Based on the first data format of the one or more words in the first output, the one or more words from the first output are included in the multiple words. Based on the second data format of one or more words in the third output, one or more words are excluded from the third output among the multiple words. The system according to claim 24.
27. A non-temporary machine-readable medium that stores instructions and, when executed by at least one computing device, causes at least one computing device to perform an operation. The aforementioned operation is, Receiving a file of a first format containing information and a region of interest (ROI), Converting the aforementioned file into an image, Using the first model, a first output is generated having a first information set extracted from the image and a first coordinate set of the first information set within the image. Using the second model, a second output is generated having a second set of coordinates for each of the multiple ROIs in the image. Using a third model, a third output is generated having a second information set extracted from the image and a third coordinate set of the second information set within the image. The first output and the third output are combined to generate a plurality of coordinates having the information contained in the file and the coordinates of the information in the image. Using the second output, generate a second format output file having multiple sections, where each section of the multiple sections corresponds to an ROI of the multiple ROIs, and each section of the multiple sections is included in the output file based on the coordinates of the ROI in the second coordinate set corresponding to each section. Based on the coordinates corresponding to the portion of the information and the coordinates of each section, the portion of the information determined to correspond to each section is input to each section of the multiple sections in the output file. Includes, The second format makes the information in the output file searchable while it is displayed on the graphical user interface (GUI) or stored in the data storage device. Non-temporary machine-readable media.
28. The information comprises one or more words, and generating the first output or the third output comprises generating bounding boxes surrounding each of the one or more words in the image. The non-temporary machine-readable medium according to claim 27.
29. The third coordinate set forms a bounding box surrounding each of the plurality of ROIs. The non-temporary machine-readable medium according to claim 27.
30. The first output is generated using optical character recognition (OCR). The non-temporary machine-readable medium according to claim 27.
31. The third output is generated using a neural network. The non-temporary machine-readable medium according to claim 27.
32. The output file is used by one or more machine learning models to generate a knowledge base. The non-temporary machine-readable medium according to claim 27.
33. The aforementioned information is selectable within the output file. The non-temporary machine-readable medium according to claim 27.
34. The operation further comprises displaying the output file and the image. The non-temporary machine-readable medium according to claim 27.
35. The aforementioned information includes words and / or images, The non-temporary machine-readable medium according to claim 27.
36. Combining the first output with the third output to generate the information includes holding the image included in the first output. The non-temporary machine-readable medium according to claim 35.
37. The aforementioned operation is, Identifying one or more words in the first output that share the same mark as one or more words in the third output, Determining the similarity level between one or more words in the first output and one or more words in the third output, Furthermore, The non-temporary machine-readable medium according to claim 35.
38. The aforementioned operation is, Assigning a first priority value to the first output and a second priority value to the third output, Based on the second priority value of the third output and on the similarity level being greater than a predetermined threshold, to include one or more words from the third output among the plurality of words, Based on the first priority value of the first output and on the similarity level being greater than a predetermined threshold, one or more words from the plurality of words are excluded from the first output. Furthermore, The non-temporary machine-readable medium according to claim 37.
39. The aforementioned operation is, Identifying the first data format of one or more words in the first output and the second data format of one or more words in the third output, Based on the first data format of the one or more words in the first output, include the one or more words from the first output among the multiple words, Based on the second data format of the one or more words in the third output, remove the one or more words from the third output among the multiple words. Furthermore, The non-temporary machine-readable medium according to claim 37.
Citation Information
Patent Citations
System and method for providing access to multimodal content in a technical document
EP3961425A1
Form type learning system and image processing apparatus
JP2019109562A
Image processing system, image processing method, and image processing apparatus
JP2020101843A
Image processing system, image processing method and program
JP2021039424A
Image processing method and image processing system
JP2021502628A