System and method for evaluating the quality of visual representations

The system uses AI and machine learning to analyze and improve the accessibility, readability, and explainability of scientific figures by detecting and correcting issues like color blindness and missing legends, enhancing the clarity and inclusivity of scientific publications.

WO2025184481A1PCT designated stage Publication Date: 2025-09-04THE REGENTS OF THE UNIVERSITY OF COLORADO
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017808
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Scientific figures in publications often suffer from accessibility, readability, and explainability issues, such as color-blind safety, low contrast, and lack of captions, which hinder understanding and interpretation by readers.

Method used

A system and method using artificial intelligence and machine learning algorithms to analyze digital documents, classify images into groups based on accessibility, readability, and explainability, and automatically detect and address issues like color blindness, low contrast, and missing legends or captions.

Benefits of technology

The system achieves high accuracy in detecting and addressing accessibility, readability, and explainability issues in a large sample of publications, improving the clarity and inclusivity of visual representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000028_0000
    Figure 00000028_0000
  • Figure 00000029_0000
    Figure 00000029_0000
  • Figure 00000029_0001
    Figure 00000029_0001
Patent Text Reader

Abstract

A method for analyzing digital documents including one or more images includes reading through a digital document to identify one or more images, classifying a portion of the images into a first group and a second group, for a first quality of visual representation, analyzing the images to determine a level of the first quality of visual representation, for the second quality of visual representation, denoising the first group and resizing the denoised first group to a standard size and analyzing the resized first group to determine a level of the second quality, for the third quality of visual representation, analyzing the second group to determine a level of the third quality of visual representation, and in a case where the level of the first, second, or third quality is lower than a respective threshold, notifying the level of the first, second, or third quality.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR EVALUATING THE QUALITY OF VISUAL REPRESENTATIONSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to United States Provisional Patent Application Serial No. 63 / 559,393 filed on February 29, 2024, and entitled "System and Method for Evaluating the Quality of Visual Representations,” which is expressly incorporated herein by reference in its entirety.GOVERNMENT RIGHTS

[0002] This invention was made with government support under grant number ORI1R 180041 , awarded by the U.S. Office of Research Integrity-. The government has certain rights in the invention.FIELD

[0003] This disclosure relates to systems and methods for evaluating the quality- of visual representations, and more particularly for evaluating the quality of visual representations based on how accessible, readable, and explainable documents are.BACKGROUND

[0004] Figures are an essential part of scientific publications because they can present complex data relationships to readers in an efficient manner. However, figures can have several issues that reduce their communication quality. For example, when they contain color-blind issues, they preclude readers from understanding the underlying trends or even make them misinterpret results. In some disciplines, editors and readers might not pay enough attention to figures, partially because it is time-consuming. Using computational methods to help flag common patterns in figures could thus be important.

[0005] The subject matter claimed herein is not limited to embodiments that solve any disadvantages or that operate only in environments such as those described above. Rather, this background is only provided to illustrate one exemplary technology area where some aspects described herein may be practiced.BRIEF SUMMARY

[0006] Disclosed embodiments include systems and methods for evaluating the quality of visual representations. The quality of visual representations are evaluated based on how accessible(e.g., color-blind safe), readable (e.g., good contrast), and explainable (e.g., presence of captions and legends) documents are.

[0007] In accordance with various aspects of the present disclosure, the techniques described herein relate to a method for analyzing digital documents including one or more images, the method including: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second quality of visual representation and a second group is classified for a third quality of visual representation; for a first qualify of visual representation: analyzing the one or more images to determine a level of the first qualify of visual representation; for the second qualify of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second qualify of visual representation; for the third qualify of visual representation: analyzing the second group to determine a level of the third qualify of visual representation; and in a case where the level of the first, second, or third qualify of visual representation is lower than a respective threshold, notifying the level of the first, second, or third qualify of visual representation.

[0008] In accordance with various aspects of the present disclosure, the techniques described herein relate to an apparatus analyzing a digital document including one or more images, the apparatus including: one or more processors; and a memory' including instructions that, when executed by the one or more processors, cause the apparatus to perform operations for analyzing digital documents including one or more images, the operations including: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second quality of visual representation and a second group is classified for a third qualify of visual representation; for a first qualify of visual representation: analyzing the one or more images to determine a level of the first qualify of visual representation; for the second qualify’ of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second qualify of visual representation; for the third qualify of visual representation: analyzing the second group to determine a level of the third quality of visual representation; and in a case where the level of the first, second, or third qualify of visual representation is lower than a respective threshold, notifying the level of the first, second, or third qualify of visual representation.

[0009] In accordance with various aspects of the present disclosure, the techniques described herein relate to a nontransitory computer-readable medium including instructions that, when executed by a computer, cause the computer to perform operations for analyzing digital documentsincluding one or more images, the operations including: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second qualify of visual representation and a second group is classified for a third qualify of visual representation; for a first qualify of visual representation: analyzing the one or more images to determine a level of the first qualify of visual representation; for the second qualify of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second qualify of visual representation; for the third qualify of visual representation: analyzing the second group to determine a level of the third qualify' of visual representation; and in a case where the level of the first, second, or third qualify of visual representation is lower than a respective threshold, notifying the level of the first, second, or third qualify of visual representation.

[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify’ key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0011] Additional features and advantages will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of the teachings herein. Features and advantages of the present disclosure may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. Features of the present disclosure will become more fully apparent from the following description and appended claims, or may be learned by the practice of the present disclosure as set forth hereinafter.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to describe the manner in which at least some of the advantages and features of the present disclosure may be obtained, a more particular description of aspects of the present disclosure will be rendered by reference to specific aspects thereof which are illustrated in the appended drawings. Understanding that these drawings depict only typical aspects of the present disclosure and are not therefore to be considered to be limiting of its scope, aspects of the present disclosure will be described and explained with additional specificity and detail through the use of the accompanying drawings.

[0013] FIG. 1 illustrates a block diagram of a method for analyzing images in documents according to various aspects of the present disclosure.

[0014] FIG. 2A illustrates a graphical representation having no accessibility issues according to various aspects of the present disclosure.

[0015] FIG. 2B illustrates a graphical representation having no accessibility issues according to various aspects of the present disclosure.

[0016] FIG. 2C illustrates two graphical representations of chemical structures having no accessibility issues according to various aspects of the present disclosure.

[0017] FIG. 2D illustrates a graphical representation having accessibility issues according to various aspects of the present disclosure.

[0018] FIG. 2E illustrates a graphical representation having accessibility issues according to various aspects of the present disclosure.

[0019] FIG. 2F illustrates two graphical representations having accessibility issues according to various aspects of the present disclosure.

[0020] FIG. 3A illustrates a graphical representation having no readability issues according to various aspects of the present disclosure.

[0021] FIG. 3B illustrates two graphical representations having no readability issues according to various aspects of the present disclosure.

[0022] FIG. 3C illustrates a graphical representation of a chemical structure having no readability issues according to various aspects of the present disclosure.

[0023] FIG. 3D illustrates three scanned images having readability issues according to various aspects of the present disclosure.

[0024] FIG. 3E illustrates an angiographic image of blood vessels having readability issues according to various aspects of the present disclosure.

[0025] FIG. 3F illustrates two X-ray images having readability issues according to various aspects of the present disclosure.

[0026] FIG. 3G illustrates a scanned image having readability issues according to various aspects of the present disclosure.

[0027] FIG. 4A illustrates three data plots having no explainability issues according to various aspects of the present disclosure.

[0028] FIG. 4B illustrates a data plot having no explainability issues according to various aspects of the present disclosure.

[0029] FIG. 4C illustrates two data plots having no explainability issues according to various aspects of the present disclosure.

[0030] FIG. 4D illustrates six data plots having explainability issues according to various aspects of the present disclosure.

[0031] FIG. 4E illustrates four bar graphs and two data plots having explainability issues according to various aspects of the present disclosure.

[0032] FIG. 5A illustrates seven bar graphs of documents based on bibliometric factors according to various aspects of the present disclosure.

[0033] FIG. 5B illustrates a table of example bibliometric factors of documents according to various aspects of the present disclosure.

[0034] FIG. 6 illustrates correlated features of documents according to various aspects of the present disclosure.

[0035] FIG. 7 illustrates a table of standardized coefficients on documents with respect to bibliometric factors according to various aspects of the present disclosure.

[0036] FIG. 8 illustrates statistical marginal effects on documents having accessibility issues with respect to bibliometric factors according to various aspects of the present disclosure.

[0037] FIG. 9 illustrates a table of standardized coefficients on documents having readability issues with respect to bibliometric factors according to various aspects of the present disclosure.

[0038] FIG. 10 illustrates statistical marginal effects on documents having readability issues with respect to bibliometric factors according to various aspects of the present disclosure.

[0039] FIG. 11 illustrates statistical marginal effects on documents having explainability issues with respect to a bibliometric factor according to various aspects of the present disclosure.

[0040] FIG. 12 illustrates a flow chart for detecting accessibility issues according to various aspects of the present disclosure.DETAILED DESCRIPTION

[0041] Embodiments of the present disclosure generally relate to evaluating the quality of visual representations. More particularly, at least some embodiments of the present disclosure relate to systems, hardware, software, computer-readable media, and methods for detecting the quality of visual representations issues in publications. The quality of visual representations are evaluated based on how accessible (e.g., color-blind safe), readable (e g., good contrast), and explainable (e.g.. presence of captions and legends) publications are.

[0042] Publications including research papers in various areas of technologies inherently include data plots, images, bar graphs, tables, and any other graphical or visual representations because they are critical part in visually presenting data relationships or increasing reader’s understanding of the context in an efficient way. These visual representations, however, in publications might have accessibility, readability, and explainability (ARE) issues that prevent readers from clearly reading or fully understanding the context thereof. Thus, the presentdisclosure provides ways to automatically detect the ARE issues in the publications so as to inform authors / publishers of visual representations having the ARE issues in the publications.

[0043] Computational techniques to measure these features and analyze a large sample of publications may be developed from open access publications. Disclosed systems and methods combine computer and human vision research principles, achieving high accuracy in detecting the ARE problems. In the sample, around 20.6% of publications are estimated to contain either one of accessibility, readability, or explainability issues.

[0044] Three classifiers may be disclosed herein to detect whether a panel within a figure has color-blind unsafe problems, low light / contrast problems, and explainability7problems that the figure has insufficient legend or captions. Disclosed systems and methods may be validated on a hand-annotated training dataset and simulated image datasets, achieving high accuracy on these three tasks. Artificial intelligence or machine learning algorithms or models may be used to reinforce the accuracy thereof. For example, sample documents may include a sample of 70,000+ publications and about 300,000 figures from the PubMed Open Access Subset. The results show that around 2% of all figures contain accessibility issues, 3% of diagnostic figures contain readability issues, and 23% of line charts contain explainability issues. Following may be analyzed: whether these issues are associated with bibliometric factors such as ranking of the journal, seniority of the researcher, country, and field. This present disclosure may be applied as good publication practices in various areas of technologies.

[0045] Various aspects of the present disclosure, such as the examples disclosed herein, may be beneficial in a variety of respects. For example, and as will be apparent from the present disclosure, one or more embodiments of the present disclosure may provide one or more advantageous and unexpected effects, in any combination, some examples of which are set forth below. It should be noted that such effects are neither intended, nor should be construed, to limit the scope of the present disclosure in any way. It should further be noted that nothing herein should be construed as constituting an essential or indispensable element of any embodiment. Rather, various aspects of the present disclosure may be combined in a variety7of ways so as to define yet further embodiments. Such further embodiments are considered as being within the scope of this disclosure. As well, none of the embodiments embraced within the scope of this disclosure should be construed as resolving, or being limited to the resolution of, any particular problems. Nor should any such embodiments be construed to implement, or be limited to implementation of, any particular technical effects or solutions. Finally, it is not required that any embodiment implement any of the advantageous and unexpected effects disclosed herein.

[0046] Nothing herein should be construed as teaching or suggesting that any aspect of any embodiment of the present disclosure could or would be performed, practically or otherwise, in the mind of a human. Further, and unless explicitly indicated otherwise, the disclosed methods, processes, and operations, are contemplated as being implemented by computing systems that may comprise hardware and / or software. That is, such methods, processes, and operations, are defined as being computer-implemented.

[0047] Now turning to FIG. 1, illustrated is a block diagram for a method 100 for detecting ARE issues in publications or digital documents according to various aspects of the present disclosure. The method 100 may be performed by artificial intelligence or machine learning models, which may have been trained with testing datasets. In aspects, the testing datasets may be acquired from the PubMed Open Access Subset or obtained from public / private publishers or libraries. For example, the testing datasets may include 70,000+ publications and 300,000 figures from the PubMed Open Access Subset. The testing datasets are, however, not limited to the PubMed Open Access Subset and can include any other sources of publications or digital / paper documents.

[0048] The method 100 may start by accessing documents at step 105. The documents may be research papers, dissertations, medical publications, patent publications, scholarly publications, legal documents, digital documents, or any other types of documents in various formats. In aspects, the documents may be in word, pdf. html. image (e.g., jpeg, tiff, png, gif, bmp, psd or any other image formats), email, or any other electronic format. In a case where the documents are in paper, the documents may be scanned to images or converted to pdf or any other suitable electronic format so that computers may be able to read and analyze them.

[0049] The method 100 may include step 110, at which pages of the documents are read through to access or identify images (e.g., bar graphs, data plots, diagnostic images, tables, or any other graphical representations of relationships .) The images may be color images, black or white images, medical diagnostic images, representations of chemical compounds, three-dimensional images, correlational images, or etc.

[0050] The method 100 may optionally include step 115, at which the images are randomly sampled in a case where there are too many images to be processed. In a case where one image includes more than one figure or is a compound image, which are common in publications to show relevant information and results together, the compound image may be separated into subplots at step 115. The separation may be performed by a feature extractor, which may be trained by a convolutional neural network-based model (e.g.. Resnet-152 v2, pre-trained on ImageNet). Its top layer may be trained with a compound figure classification dataset from ImageCLEF 2016.

[0051] In aspects, compound figures may be separated into each subplot by a fine-tuned convolutional neural network (YOLO v4, pre-trained on MS COCO dataset) with a subfigure separation dataset from ImageCLEF 2016.

[0052] The method 100 may further include step 120, at which the subplots and other images are saved or collected at a storage medium or buffer.

[0053] After step 120. the method 100 may parallelly process each of the ARE issues, as illustrated in FIG. 1. In an aspect, processes for the ARE issues may be serially performed or in any combination or order. Steps 125 and 130 are for the accessibility issues, step 135-150 are for the readability7issues, and steps 155-170 are for the explainability issues.

[0054] For the accessibility7issues, all subplots and all images are used in step 125. Accessibility in scientific articles is closely related to being “accessible information technology” artifacts, which is defined by the Americans with Disability Act (ADA) as “technology that can be used by people with a wide range of abilities and disabilities. It incorporates the principles of universal design, whereby each user can interact with the technology in ways that work best for him or her.” Accessibility issues are rooted in color combinations or color maps, such as the rainbow color map and can affect how various color-blind readers (such as red-green colorblindness and blue-yellow colorblindness) understand a figure and research findings. This issue needs attention from authors and publishers because, in some regions, 14% of middle-aged populations are color-blind, and from 1.69% to 8.73% of the population is colorblind.

[0055] Specifically, when readers have color blindness, they are unable to distinguish one color from another. For example, deuteranomaly is the most common type of red-green color vision deficiency. Thus, readers having deuteranomaly are unable to distinguish red from green so that they tend to perceive green regions as red regions or vice versa. For another example, tritanomaly makes it hard to tell the difference between blue and green and between yellow and red. Thus, a combination of blue and green colors or yellow and red colors makes it harder for readers with tritanomaly to distinguish one color from another color. At step 130, the images and subplots are analyzed to identify whether they have colors related to color blindness issues.

[0056] In a case where colors related to color blindness issues exist in the images or the subplots, the method 100 may further include step 175, at which the accessibility issues at an image or a subplot may be notified so that a proper measure may be applied to such images and subplots to address the accessibility7issues. The proper measure may include replacing the pair of colors related to the color blindness issues with colors not related to the color blindness issues. For example, red and green color combination may be replaced with orange and blue color combination.

[0057] Now turning back to the readability issues, the method 100 may include step 135, at which a portion of the images and subplots are classified for the readability issues into a first group. The readability issues prevent readers from correctly parsing information presented in an image or subplot. Among the factors affecting this ability may be low light and low contrast images and high complexity images. Low light and low contrast issues are especially worrisome among elderly readers because they have difficulty reading low-contrast images, and are likely to make readers lose details within the context. High complexity images typically need more attention and effort for readers to understand, and this complexity may be quantified by computational methods.

[0058] The method 100 may further include step 140, at which the first group of images and subplots is identified, collected, and saved. The first group may be then analyzed to find out which images and / or subplots have low light at step 145 and have low spatial frequency at step 150.

[0059] In a case where an image or a subplot is identified as having a low-light or low-contrast, the image or the subplot is notified at step 175 with the readability' issues.

[0060] Now turning back to the explainability issues, The method 100 may include step 155, at which a portion of the images and subplots are classified for the explainability issues into a second group. Explainability' may be broadly understood as being able to produce an explanation of the trends and factors observed in a data plot. For example, some line charts or bar charts need legend or caption to assist readers to understand them because of their usage of colors and graphic design. Also, biological images can be hard to interpret if they do not have a scale bar. Explainability issues affect all the population because they7relate to whether a figure contains a good caption or legend. A recent study found that between 5 to 17 percent of scientific figures do not provide enough information to explain the colors inside them. Yet, this previously mentioned study has a relatively small sample of figures, and a large-scale analysis on this issue might estimate the prevalence of explainability issues more accurately.

[0061] The method 100 may further include step 160, at which the second group of images and subplots is identified, collected, and saved. The second group may be then analyzed to find out which images and / or do not have legends at step 165 and captions at step 170.

[0062] In a case where an image or a subplot is identified as missing legends or captions, the image or the subplot is notified at step 175 with the explainability' issues.

[0063] In aspects, the notification step 175 may be performed by outputting a list of pairs, each pair including an image or subplot and one or more related issues from among the ARE issues.

[0064] Specifically, in this simulation, the followings are randomly selected: 300,000 figures from 71.508 publications in PudMed Open Access, a subset of PubMed Central, containing millions of publications. Compound figure classification and separation have been applied to get 788,028 subplots to analyze each image or subplot. The accessibility of all subplots has been assessed with computer vision techniques. The readability of diagnostic figures has been estimated by using a fine-tuned deep learning network (e.g., ResNet50 v2). Finally, the explainability of line charts have been estimated by fine-tuning a ResNetl52 v2 classifier using the annotations.

[0065] To train the classification and detection models for explainability analysis, 1,407 line charts for legend detection and 1,454 line charts have been manually annotated for legend necessary’ classification. In legend detection, where the legend was in the chart have been annotated. For legend neediness, line charts as needed have been annotated for a legend if there is more than one line or symbol.

[0066] A high-quality feature extractor has been applied based on a convolutional neural network (e g., Resnet-152 v2, pre-trained on ImageNet) to classify figures into bar charts, line charts, scatter charts, heatmap charts, box charts, area charts, radar plots, maps, pie charts, tables, pareto charts, Venn diagrams, violin charts, diagnostic figures, etc. These categorized and separated subplots are analyzed at steps 130, 145, 150, 165, and 170 to find the ARE issues and to train and fine-tune artificial intelligence or machine learning models.

[0067] In various aspects, the method 100 may further include a fixing step to automatically address the identified ARE issues based on the results. Specifically, when the ARE issues are identified, the image may be touched up, modified, and corrected, and to generate another image, which does not include any ARE issues.

[0068] In various aspects, a fixing tool may be incorporated into viewers of the images so that the viewers may automatically generate and display images, which do not have any ARE issues to screen readers. For example, when there are green and red color combination, the viewers may replace the green and red color combination with the blue and yellow color combination. Or the viewers may be part of statistical packages that could warn the readers if an image contains one or more ARE issues.

[0069] FIGS. 2A-2F illustrate example data plots, tables, chemical structures, maps, three- dimensional graphs related to accessibility issues according to various aspects of the present disclosure. Specifically, FIGS. 2A-2C show a forest plot, a design diagram, and chemical structure, respectively. 95% confidence intervals are displayed with a mean in the center of each range of FIG. 2A in black, and FIG. 2B also shows the diagram in black with gray degradation. In other words, FIGS. 2A and 2B have a monotonic color, and thus have no accessibility issues.

[0070] Even though FIG. 2C shows black gradation only, portions 240 and 245 are meant to be in blue color, and the other portions of protected and free oligoyne amphiphiles are in black. Since the black color and the bule color are not related to any color blindness issues, the protected and free oligoyne amphiphiles do not have any accessibility' issues.

[0071] The lower portion of FIG. 2C, oligoyne rotaxanes, have two color portions: blue portions 250 and 256 and orange portions 253. Since the blue and orange color combination is not related to any color blindness issues, the chemical structure of oligoyne rotaxanes does not have any accessibility issues.

[0072] Now referring to FIG. 2D, illustrated are a graph 260 including three scatter graphs 263, 266. and 269. The top scatter graph 263 is in blue, the middle scatter graph 266 is in red, and the bottom scatter graph 269 is in green. In this case, the red and green color combination 266 and 269 may appear to be the same color to readers having deuteranomaly. Thus, the graph 260 may be identified as having accessibility issues at step 130 of FIG. 1.

[0073] Likewise, as illustrated in FIG. 2E, a map 270 includes Australia 273 in green and Queensland QLD 276 and New South Wales NSW in red. Also, the map 270 includes New Zealand in green and Auckland (AK), Bay of Plenty (BP) and Mid-Canterbury (MC) in red. Due to the combination of green and red, the map 270 may be identified as having accessibility issues at step 130 of FIG. 1.

[0074] FIG. 2F illustrates a heat map 280 and a three-dimensional surface plot 290. The heat map 280 is a two-dimensional representation of the three-dimensional surface plot 290 in different colors, which include green 286 and red 283. Likewise, the three-dimensional surface plot 290 also includes portions in green 296 and red 293. As such, both the heat map 280 and the three- dimensional surface plot 290 may be identified as having accessibility issues at step 130 of FIG. 1.

[0075] Now turning to the readability issues, a convolutional neural network may be finetuned with a low light image dataset, which the classifier has classified from the 300,000 figures from 71,508 publications in PudMed Open Access, as having the readability' issues. Because some scientific figures may be different from natural scenes in the training dataset, diagnostic figures in detecting readability issues may be used to train the artificial intelligence or machine learning models. In aspects, since figures of table or texts could be misclassified as diagnostic figures, the tables and texts may be removed from the training dataset for the readability issues.

[0076] Specifically, when the spatial frequency of an object in images is too high (e.g., greater than 30 of spatial frequency or 60 of spatial frequency), details of some objects can be hard for readers to view even with high contrast. Thus, in the image analysis (e.g.. step 150 of FIG. 1), the spatial frequency of one image may be estimated by transforming the image with Fast FourierTransform and by estimating the spatial frequency of pixels. Then, it is measured if one image has a large area (greater than half of the image) with a high spatial frequency. Taking the low-light image classifier and spatial frequency analysis together, if an image is classified as a low-light image by the classifier and also contains a large area of high spatial frequency, such images with readability issues can be considered.

[0077] The low light analysis may be done by removing images of table and text by computing the size of white background in images and removing images with large w hile background from analysis. Low light image classifier trained on a set of low light images in nature scenes is applied. High spatial frequency analysis may be done by transforming images into magnitude spectrum with fast Fourier transform and checking the size of area in images with high spatial frequency (the default may be 60). If the low light image classifier classify an image as low light and it has a large size of area with high spatial frequency, it is considered such images as low" light and contrast image.

[0078] FIG. 3A illustrates a linkage disequilibrium (LD) plot 300, also known as an LD heatmap or Haploview plot, commonly used in genetics and genomics to visualize correlations between genetic variants (single nucleotide polymorphisms (SNPs)). The LD plot 300 includes three color regions: an orange color region 303, a yellow" color region 306, and a red color region 309. Even though the red color is included in the LD plot 300, there is no green corresponding to red in deuteranomaly. Thus, the LD plot 300 may be identified as having no accessibility issues at step 130 of FIG. 1. Further, since there is no low"-light region, the LD plot 300 may be identified as having no readability issues at step 145 of FIG. 1 . Furthermore, since the boundaries of the LD plot 300 are simple and straight, the LD plot 300 may be identified as having no spatial frequency issues at step 150 of FIG. 1.

[0079] Now turning to FIGS. 3B and 3C. illustrated are an image 310 for analyzing protein fragment and a chemical structure diagram 320. Both 310 and 320 are in black and have high light and suitable spatial frequency. Thus, FIGS. 3B and 3C may be identified as having no spatial frequency issues at step 150 of FIG. 1.

[0080] On the other hands, FIGS. 3D-3F illustrate images having readability issues. Specifically, FIG. 3D shows medical imaging scans 300 including one ultrasound image A and two X-ray images B and C. The white arrowhead in the ultrasound image A shows low" light, while black arrowheads in the X-ray images B and C show" high contrast. Thus, the ultrasound image A in the scans 330 may be identified as having low light at step 145 of FIG. 1. and the X-ray images B and C may be identified as having high spatial frequency issues at step 150 of FIG. 1.

[0081] Now turning to FIG. 3E, illustrated is a coronary angiography scan 340 showing blood vessels. Due to low light and seemly high spatial frequencies, the scan 340 may be identified as having low light at step 145 of FIG. 1 , and the X-ray images B and C may be identified as having high spatial frequency issues at step 150 of FIG. 1.

[0082] FIG. 3E illustrates an image 350 showing dental and craniofacial X-rays, which may be identified as having low light at step 140 of FIG. 1 or high spatial frequency issues at step 150 of FIG. 1. Likewise, a scan image 360 as illustrated in FIG. 3G may be also identified as having low light at step 140 of FIG. 1 or high spatial frequency issues at step 150 of FIG. 1.

[0001] Now turning to the explainability issues, two parts of analysis may be needed to estimate the explainability issues: legend detection and caption analysis. Legend detection may be performed by a deep neural network model (e.g.. YOLO-v4. pre-trained on MS COCO dataset) on human-annotated figures. The deep neural network model may be fine-tuned to identify legends on scientific figures. Some compound figures may only contain one legend, which applies to all subplots therein. Thus, when it is determined that a legend exists in an original compound figure, the legend is considered as existing for all subplots in the original compound figure.

[0083] The legend may not be necessary when there is only one line in the graph. To filter this situation out, a convolution neural network (e.g., ResNetl52v2, pre-trained on ImageNet) on human-annotated charts is fine-tuned to classify if figures need legend or not (0.73 precision, 0.81 recall on the testing dataset).

[0084] Caption may exist in images and textually explain the image. Thus, a parser may be used to extract textual captions from images and check whether color or symbol explanations exist (e.g., blue, red, green, dashed line, solid line, triangle, square, etc.) in the images. Further, a line chart may be classified to have explainability issues if a legend and explanation in the caption for legend-needed charts is classified.

[0085] FIG. 4 A illustrates three data plots 400, 405, and 410. The data plots 405 and 410 show trajectory of wide-type (Wt) particles based on the captions 402 and 406, and the data plot 410 shows mean squared displacement (MSD) vs. time plot based on the caption 414 and the legend 412. Thus, three data plots 400, 405. and 410 may be identified as having no explainability issues at steps 165 and 170 of FIG. 1.

[0086] Likewise, FIGS. 4B and 4C illustrate captions near horizontal and vertical axes, thereby explaining the relationship of data plots 420, 430, and 435. Further, since each of the data plots 420, 430, and 435 has one data curve, no legend is required.

[0087] On the other hand, data plots or bar graphs illustrated in FIGS. 4D and 4E may be considered as having the explainability issues. Specifically, an image 440 in FIG. 4D shows sixdata plots a-f. Captions of each data plot have "G / Go" and “Ig (CFU / g),” which may be acronyms but are not spelled out. Thus, six data plots a-f are not self-explanatory without the acronyms being spelled out and may be identified as having the explainability issue at step 165 of FIG. 1 .

[0088] FIG. 4E illustrates two sets 460 and 470 of sublots. Set 460 includes two histograms and one scatter data plot 465. Histograms have sufficient to understand the horizontal and vertical axes. However, “p=2.9e-11” requires explanation, which does not exist in the set 460. Further, even though the titles of the histograms, “stress response probesets” and “all other probesets,” and the horizontal axis caption “expression ratio (log scale) are provided,” it is not self-explanatory to identify the relationship between the title and the horizontal axis caption.

[0089] The scatter data plot 465 is not self-explanatory either because it is not clear what each data point represents.

[0090] The set 470 and the scatter data plot 475 have issues similar to those of the set 460 and the scatter data plot 465. Thus, each subplot in FIG. 4E may be identified as having explainability issues at steps 165 and 170 of FIG. 1.

[0091] Now referring back to the documents used in training the artificial intelligence or machine learning models for analyzing the ARE issues, the documents have several bibliometric factors in Table 1 as illustrated in FIG. 5B. The bibliometric factors may include field of study, the number of publications, average h-index of authors, journal rank, author academic age, and journal age. The list of these bibliometric factors is not limited thereto but may include other factors (e.g., countries of publications). These factors are shown in the horizontal axes of histograms 500 as illustrated in FIG. 5A.

[0092] After matching publications in the sample to Microsoft Academic Graph (MAG) and removing outliers, 57.837 publications from 1,818 journals from 1966 until 2018 have been analyzed. The three most popular fields are Biology, Medicine, and Chemistry.

[0093] According to FIG. 5A, documents produced in the United States comprise close to 40% among all the documents, and China produces about 15%, and so on. The journal rank is based on the PageRank of the citations to the journal, as calculated by the MAG. The average rank of journals is 10,353. with a standard deviation (SD) = 1,638. The average number of publications by a journal is 10,193.22, with a minimum of 248 and a maximum of 276,186. The average age of a journal is 37.41, with the newest being seven years old and the oldest being 222 years old. The average h-index of authors in journals is 12.47, with an SD of 4.19. Finally, the average academic age is 20.48 years, with an SD of 5.79 years.

[0094] FIG. 6 illustrates a correlation matrix 600 of predictors. The most correlated features are the journal’s rank, number of publications, and journal age. Journal rank is numerically highif the journal is not cited as often as a journal with a rank numerically smaller. These correlations are negative, meaning that top-cited journal cites produce more papers and are older.

[0095] Accessibility issues may be analyzed in scientific publications and 2% of scientific figures (788,028 figures in the sample) are determined to contain color-blind unsafe figures. Multiple linear regression analysis may be used to examine whether or not journals’ bibliometric factors associated with the percentage of their publications with accessibility issues. It is found that the journals in business and in engineering fields published the least and most color-blind unsafe figures, respectively, in Table 2 as illustrated in FIG. 7.

[0096] FIG. 8 illustrates statistical marginal effects on documents having accessibility issues with respect to bibliometric factors. The vertical axes represent proportion of publications with accessibility issues and the horizontal axes represent author’s academic age, average h-index of authors, age of journal, and journal rank for regression trend lines 810-840, respectively. For example, journals’ average author h-index has the biggest coefficient (standardized coef = 0.20, t(3895) = 11.43, p < 0.001) based on the regression trend line 820. suggesting that journals with highly cited authors have a higher proportion of articles with accessibility issues. Authors’ academic age shows the negative effect (standardized coef = -0.057, t(3895) = -3.41, p < 0.001) based on the regression trend line 830. In other words, older journals have a lower proportion of articles with accessibility issues (standardized coef = -0.0493. t(3895) = -2.6251, p = 0.0087). Likewise, based on the regression trend lines 830 and 840, the author’s academic age and journal rank have the negative effect similar to the author’s academic age.

[0097] Now referring to FIG. 9, illustrated is Table 3 of standardized coefficients on documents having readability issues with respect to bibliometric factors. Readability issues (e.g., low light and contrast images) may alfect older readers. The elderly population is likely to have difficulty reading images with low contrast or low contrast. Artificial intelligence and / or machine learning models may be developed to automatically assess low contrast and low light in images. 3% of medical diagnostic figures (259,351 diagnostic figures) may have low light and contrast. Multiple linear regression analysis may be used to examine whether or not journals’ bibliometric factors associated with the percentage of their publications with readability issues.

[0098] Table 3, as illustrated in FIG. 9, shows statistical marginal effects on documents having readability issues with respect to bibliometric factors. Based on the statistical measures in Table 3, journals in physics and business fields published the most and least figures with readability issues, respectively. Journals’ average author h-index has the biggest coefficient (standardized coef = 0.20, t(3228) = 10.20, p < 0.001), suggesting that journals with highly cited authors have a higherproportion of articles with readability issues. That is also evidenced in a regression trend line 1010 of FIG. 10. Also, a regression trend line 1020 of FIG. 10 shows that older journals have a higher proportion of articles with readability issues (standardized coef = 0.066, t(3228) = 3. 16, p = 0.002).

[0099] Another issue in scientific figures is explainability: some figures have no legend even if they have multiple colors, symbols, or lines or have no caption. In this simulation, line charts may be focused on only because they typically need an explanation. A method may be developed to split he classification into detecting legend and whether the caption exists and contains legend- related information. The explainability issues with line charts may be very high: around 23% of them (22,065 line charts) lacking legends and color explanations in their captions. While the method may not be perfect, in the worst-case scenario, the prevention of explainability issues may be predicted to be surprisingly high still. Multiple linear regression analysis may be used to examine if journals’ bibliometric factors associate with the percentage of their publications with explainability issues. Based on FIG. 11, journals’ average author h-index may have the biggest coefficient (standardized coef = -0.20, t(2191) = -8.77, p < 0.001), suggesting that journals with highly-cited scientists have a lower proportion of articles with explainability issues.

[0100] Now referring to FIG. 12, illustrated is a flow chart of a method 1200 for analyzing documents to detect accessibility' issues according to various aspects of the present disclosure. The method 1200 may include step 1210, at which images or plots are denoised and resized to a standard size. By resizing to the standard size, every image, plot, and subplot can be compared with the same restriction.

[0101] The method 1200 may further include step 1220, at which a color at each pixel is identified, and step 1230, at which distribution of pixel colors is computed. Based on the pixel color distribution, it is determined at step 1240 whether or not red and green colors are contained in the color distribution. The green and red colors may be based on deuteranomaly.

[0102] In a case where it is determined that the red and green colors are not contained in the color distribution at step 1240, the image is identified as having no deuteranomaly issues or being deuteranomaly safe at step 1250.

[0103] In another case where it is determined that the red and green colors are contained in the color distribution at step 1240, the method 1200 may further include step 1260, at which the image is processed by simulating deuteranomaly vision on the image. This simulation may shift red and green hues to match based on the perception of deuteranomalous vision. Thereby, the contrast between red and green can be reduced or look more similar to each other.

[0104] The method 1200 may further include step 1270, at which it is determined whether or not the read area in the image disappears. If the read area disappears, the method 1200 may furtherinclude step 1280 to indicate or output that the image is deuteranomaly unsafe. Otherwise, the image is identified as deuteranomaly safe.

[0105] In aspects, since tritanomaly makes it hard to tell the difference between blue and green and between yellow and red, step 1240 may also determine whether or not a combination of blue and green colors or yellow and red colors are contained in the color distribution instead of red and green colors. In this case, the image may be identified as tritanomaly safe at steps 1250 and 1290 or unsafe at step 1280.

[0106] In this present disclosure, artificial intelligence or machine learning models may be used to analyze accessibility, readability, and explainability issues. The models may be based on a combination of classifiers for accessibility and readability issues and parsers for explainabili ty issues. According to the positive or negative effects illustrated in FIGS. 7-11, the ARE issues may be predicted based on several bibliometric factors at the journal level.

[0107] Systems and methods disclosed herein may be implemented by one or more computers. Computing system functionality can be enhanced by a computing systems’ ability to be interconnected to other computing devices via network connections. Network connections may include, but are not limited to, connections via wireless connections including satellite, Ethernet, cellular connections, or wired connections including even computer to computer connections through serial, parallel, USB, or other connections. The connections allow a computing system to access services at other computing systems and to quickly and efficiently receive application data from other computing systems.

[0108] Interconnection of computing systems has facilitated distributed computing systems, such as so-called “cloud” computing systems. In this description, “cloud computing” may be systems or resources for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, services, etc.) that can be provisioned and released with reduced management effort or service provider interaction. A cloud model can be composed of various characteristics (e.g., on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc.), service models (e.g.. Software as a Service (“SaaS”), Platform as a Service (“PaaS”), Infrastructure as a Service (“laaS”), and deployment models (e.g., private cloud, community cloud, public cloud, hybrid cloud, etc.).

[0109] Cloud and remote based service applications are prevalent. Such applications are hosted on public and private remote systems such as clouds and usually offer a plurality of web based services for communicating back and forth with clients.

[0110] Many computers are intended to be used by direct user interaction with the computer. As such, computers have input hardware and software user interfaces to facilitate user interaction. For example, a modem general-purpose computer may include a keyboard, mouse, touchpad, camera, etc. for allowing a user to input data into the computer. In addition, various software user interfaces may be available. Examples of software user interfaces include graphical user interfaces, text command line based user interface, function key or hot key user interfaces, and the like.[OHl] Disclosed aspects may comprise or utilize a special purpose or general -purpose computer including computer hardware, as discussed in greater detail below. Disclosed aspects also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computer-executable instructions are physical storage media. Computer-readable media that carry' computer-executable instructions are transmission media. Thus, by way of example, and not limitation, aspects of the present disclosure can comprise at least two distinctly different kinds of computer-readable media: physical computer-readable storage media and transmission computer-readable media. The transmission computer-readable media that carry computer-executable instructions may include signals, carrier waves, and propagating signals.

[0112] On the other hand, the physical computer-readable storage media may be volatile memory, which requires power to maintain stored information. The physical computer-readable storage media may be non-volatile memory', which retains stored information when the non-volatile memory is not powered. In some aspects, the non-volatile memory may include flash memory, dynamic random-access memory (DRAM), ferroelectric random-access memory (FRAM), or phase-change random access memory (PRAM). In some aspects, the computer- readable media may include, by way of non-limiting examples, CD-ROMs, DVDs, flash memory' devices, magnetic disk drives, magnetic tapes drives, optical disk drives, and cloud computingbased storage. In some aspects, the computer-readable media may be a combination of devices such as those disclosed herein.

[0113] The physical computer-readable media may' include executable instructions (e.g., codes, programs, algorithms, etc.). The executable instructions represent instructions that are executable by the processor. Further, the computer-readable media may exclude signals, carrier waves, and propagating signals.

[0114] Generally, a processor executes executable instructions stored in the computer- readable media. The processor may include, without limitation, Field-Programmable Gate Arrays (“FPGAs”), Program-Specific or Application-Specific Integrated Circuits (‘'ASICs”), Program- Specific Standard Products (“ASSPs”), System-On-A-Chip Systems (“SOCs”), Complex Programmable Logic Devices (“CPLDs”), Central Processing Units (“CPU”), Graphical Processing Units (“GPU”), or any other type of programmable hardware by performing the basic arithmetic, logical, control and input / output (I / O) operations specified by the instructions. As used herein, terms such as “executable module,” “executable component,” “component,” “module,” or “engine” may refer to the processor or to software objects, routines, or methods that may be executed by the processor. The different components, modules, engines, and services described herein may be implemented as objects, codes, programs, or libraries that the processor executes.

[0115] A general-purpose computer, special purpose computer, or special purpose processing device also includes a display, which may be a cathode ray tube (CRT), a liquid crystal display (LCD), light emitting diode (LED), or an organic light emitting diode (OLED) display. In some aspects, the OLED display is a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display. In other aspects, the display may be a touch screen, through which alphanumerals may be input or entered. In still other aspects, the display may be a hologram, through which users may enter data by touching or swiping space.

[0116] Data or commands may be entered via an input device in the special purpose or general-purpose computer. The input device may be a keyboard, a mouse, a touch screen, or a hologram keyboard.

[0117] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry' program code in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above are also included within the scope of computer-readable media.

[0118] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission computer-readable mediate physical computer-readable storage media (or vice versa). For example, computer-executable instructions or data structures received over a networkor data link can be buffered in RAM within a network interface module (e.g., a ‘‘NIC’'), and then eventually transferred to computer system RAM and / or to less volatile computer-readable physical storage media at a computer system. Thus, computer-readable physical storage media can be included in computer system components that also (or even primarily) utilize transmission media.

[0119] Computer-executable instructions comprise, for example, instructions and data which cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. The computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0120] Those skilled in the art will appreciate that the present disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones tablets, mobile devices, smartphones. PDAs, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory' storage devices.

[0121] Alternatively, or in addition, the functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include FPGAs, ASICs, ASSPs, SOCs, CPLDs, etc.Example Implementations

[0122] In view of the foregoing, the present disclosure relates, for example and without being limited thereto, to the following aspects:

[0123] Clause 1. A method for analyzing digital documents including one or more images, the method comprising: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second quality of visual representation and a second group isclassified for a third quality of visual representation; for a first quality of visual representation: analyzing the one or more images to determine a level of the first quality of visual representation; for the second quality of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second quality of visual representation; for the third quality' of visual representation: analyzing the second group to determine a level of the third quality of visual representation; and in a case where the level of the first, second, or third quality of visual representation is lower than a respective threshold, notifying the level of the first, second, or third quality of visual representation.

[0124] Clause 2. The method according to clause 1, wherein the first quality of visual representation is related to accessibility.

[0125] Clause 3. The method according to clause 2, wherein the accessibility indicates that colors related to a color-blindness exist in the one or more images.

[0126] Clause 4. The method according to clauses 2 and 3, wherein analyzing the one or more images comprises determining whether the colors related to the color-blindness disappears to an eye of a person with the color-blindness.

[0127] Clause 5. The method according to clauses 1-4, wherein the level of the second quality of visual representation is related to readability.

[0128] Clause 6. The method according to clause 5, wherein the readability indicates that a reader is prevented from correctly parsing information presented in an image in the first group.

[0129] Clause 7. The method according to clauses 5 and 6, in a case where the level of the second quality of visual representation is lower than a threshold, further comprising: adjusting a level of light or contrast to increase the level of the second quality of visual representation.

[0130] Clause 8. The method according to clauses 1-7, wherein the third quality of visual representation is related to explainability.

[0131] Clause 9. The method according to clause 8, wherein the explainability indicates that an image in the second group does not contain a caption or legend.

[0132] Clause 10. The method according to clauses 1-9, wherein notifying the level of the first, second, or third quality of visual representation comprises outputting a list of pairs of an image, of which the level of the first, second, or third visual representation is lower than the respective threshold, and a corresponding issue of the level of the first, second, or third quality of visual representation.

[0133] Clause 11. An apparatus analyzing a digital document including one or more images, the apparatus comprising: one or more processors; and a memory including instructions that, when executed by the one or more processors, cause the apparatus to perform operations for analyzingdigital documents including one or more images, the operations comprising: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second quality of visual representation and a second group is classified for a third qualify of visual representation; for a first qualify of visual representation: analyzing the one or more images to determine a level of the first quality of visual representation; for the second quality of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second qualify of visual representation; for the third qualify of visual representation: analyzing the second group to determine a level of the third quality of visual representation; and in a case where the level of the first, second, or third qualify of visual representation is lower than a respective threshold, notifying the level of the first, second, or third qualify of visual representation.

[0134] Clause 12. The apparatus according to clause 11, wherein the first qualify of visual representation is related to accessibility.

[0135] Clause 13. The apparatus according to clause 12. wherein the accessibility indicates that colors related to a color-blindness exist in the one or more images.

[0136] Clause 14. The apparatus according to clauses 12 and 13, wherein analyzing the one or more images comprises determining whether the colors related to the color-blindness disappears to an eye of a person with the color-blindness.

[0137] Clause 15. The apparatus according to clauses 11-14, wherein the level of the second qualify of visual representation is related to readability.

[0138] Clause 16. The apparatus according to clause 15, wherein the readability indicates that a reader is prevented from correctly parsing information presented in an image in the first group.

[0139] Clause 17. The apparatus according to clauses 15 and 16, wherein, in a case where the level of the second qualify of visual representation is lower than a threshold, the operations further comprise: adjusting a level of light or contrast to increase the level of the second qualify of visual representation.

[0140] Clause 18. The apparatus according to clause 11-17, wherein the third qualify of visual representation is related to explainabilify.

[0141] Clause 19. The apparatus according to clause 18, wherein the explainabilify indicates that an image in the second group does not contain a caption or legend.

[0142] Clause 20. A nontransitory computer-readable medium including instructions that, when executed by a computer, cause the computer to perform operations for analyzing digital documents including one or more images, the operations comprising: reading through a digitaldocument to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second quality of visual representation and a second group is classified for a third quality of visual representation; for a first qualify of visual representation: analyzing the one or more images to determine a level of the first qualify of visual representation; for the second quality of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second quality of visual representation; for the third quality of visual representation: analyzing the second group to determine a level of the third quality of visual representation; and in a case where the level of the first, second, or third qualify of visual representation is lower than a respective threshold, notifying the level of the first, second, or third quality of visual representation.

[0143] The present disclosure may be embodied in other specific forms without departing from its spirit or characteristics. The described aspects are to be considered in all respects only as illustrative and not restrictive. The scope of the present disclosure is, therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

CLAIMSWhat is claimed is:

1. A method for analyzing digital documents including one or more images, the method comprising: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second quality of visual representation and a second group is classified for a third quality of visual representation; for a first quality of visual representation: analyzing the one or more images to determine a level of the first quality of visual representation; for the second quality of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second quality of visual representation; for the third quality of visual representation: analyzing the second group to determine a level of the third quality of visual representation; and in a case where the level of the first, second, or third quality of visual representation is lower than a respective threshold, notifying the level of the first, second, or third quality of visual representation.

2. The method according to claim 1, wherein the first quality of visual representation is related to accessibility.

3. The method according to claim 2, wherein the accessibility indicates that colors related to a color-blindness exist in the one or more images.

4. The method according to claim 3, wherein analyzing the one or more images comprises determining whether the colors related to the color-blindness disappears to an eye of a person with the color-blindness.

5. The method according to claim 1, wherein the level of the second quality of visual representation is related to readability.

6. The method according to claim 5, wherein the readability indicates that a reader is prevented from correctly parsing information presented in an image in the first group.

7. The method according to claim 6, in a case where the level of the second quality of visual representation is lower than a threshold, further comprising: adjusting a level of light or contrast to increase the level of the second quality' of visual representation.

8. The method according to claim 1, wherein the third quality of visual representation is related to explainability.

9. The method according to claim 8. wherein the explainability indicates that an image in the second group does not contain a caption or legend.

10. The method according to claim 1, wherein notifying the level of the first, second, or third quality of visual representation comprises outputting a list of pairs of an image, of which the level of the first, second, or third visual representation is lower than the respective threshold, and a corresponding issue of the level of the first, second, or third quality of visual representation.

11. An apparatus analyzing a digital document including one or more images, the apparatus comprising: one or more processors; and a memory' including instructions that, when executed by the one or more processors, cause the apparatus to perform operations for analyzing digital documents including one or more images, the operations comprising: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second quality of visual representation and a second group is classified for a third quality of visual representation; for a first quality of visual representation:analyzing the one or more images to determine a level of the first uality of visual representation; for the second quality of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second quality of visual representation; for the third quality of visual representation: analyzing the second group to determine a level of the third quality of visual representation; and in a case where the level of the first, second, or third quality of visual representation is lower than a respective threshold, notifying the level of the first, second, or third quality of visual representation.

12. The apparatus according to claim 11, wherein the first quality of visual representation is related to accessibility.

13. The apparatus according to claim 12, wherein the accessibility indicates that colors related to a color-blindness exist in the one or more images.

14. The apparatus according to claim 13, wherein analyzing the one or more images comprises determining whether the colors related to the color-blindness disappears to an eye of a person with the color-blindness.

15. The apparatus according to claim 11, wherein the level of the second quality of visual representation is related to readability.

16. The apparatus according to claim 15, wherein the readability indicates that a reader is prevented from correctly parsing information presented in an image in the first group.

17. The apparatus according to claim 16, wherein, in a case where the level of the second quality of visual representation is lower than a threshold, the operations further comprise: adjusting a level of light or contrast to increase the level of the second quality of visual representation.

18. The apparatus according to claim 11. wherein the third quality of visual representation is related to explainability.

19. The apparatus according to claim 18, wherein the explainability indicates that an image in the second group does not contain a caption or legend.

20. A nontransitory computer-readable medium including instructions that, when executed by a computer, cause the computer to perform operations for analyzing digital documents including one or more images, the operations comprising: reading through a digital document to identify one or more images therein; classifying a portion of the one or more images into a first group and a second group, wherein the first group is classified for a second qualify of visual representation and a second group is classified for a third qualify of visual representation; for a first qualify of visual representation: analyzing the one or more images to determine a level of the first qualify of visual representation; for the second qualify' of visual representation: denoising the first group and resizing the denoised first group to a standard size; and analyzing the resized first group to determine a level of the second qualify of visual representation; for the third qualify of visual representation: analyzing the second group to determine a level of the third qualify of visual representation; and in a case where the level of the first, second, or third qualify of visual representation is lower than a respective threshold, notifying the level of the first, second, or third qualify' of visual representation.

Citation Information

Patent Citations

  • Applying a segmentation engine to different mappings of a digital image

    US20080310715A1

  • Systems and methods for classifying objects in digital images captured using mobile devices

    US20150339526A1

  • Method and device for classifying scanned documents

    US20170351914A1

  • Utilizing machine learning and image filtering techniques to detect and analyze handwritten text

    US20210374455A1