Determining attributes of elements of displayable content and adding them to accessibility tree

By adding the role attributes of image elements in the accessibility tree, the problem of the lack of semantic information in the accessibility tree in the prior art is solved, and the complete accessibility of these contents is achieved.

CN119948479APending Publication Date: 2025-05-06GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071822.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-14
Filing Date
2023-10-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, many documents and displayable content lack the necessary semantic information in traditional accessibility trees, resulting in the incomplete accessibility of these contents in accessibility software.

Method used

By receiving images, use the layout extraction model to generate a list of elements, including bounding boxes and role properties, and add these role properties to nodes in the accessibility tree to supplement the missing semantic information.

Benefits of technology

Complete accessibility to images and other displayable content is achieved, ensuring that accessibility software can effectively navigate and access these content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948479A_ABST
    Figure CN119948479A_ABST
Patent Text Reader

Abstract

A method may receive an image representing displayable content for display by an application. A method may execute a layout extraction model that uses the image as an input and generates a list of elements of the image as an output, the list of elements including at least a bounding box defining portions of the image, and role attributes. A method may use the list of elements to add the role attributes to nodes in an accessibility tree.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is a continuation of and claims priority to U.S. application No. 18 / 046,898, filed on October 14, 2022, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] The present description relates to methods of adding properties or nodes to an accessibility tree to make additional displayable content accessible. Background Art

[0003] The present description generally relates to methods of generating roles and / or nodes for elements that can display content and adding those roles and / or nodes to an accessibility tree. The accessibility tree is used to support accessibility features in software such as screen readers that provide access to content for people with different sets of abilities. Traditionally, many documents, such as web documents, are heavily oriented toward providing content to users visually. Accessibility software such as screen readers can be used, for example, by people with no vision or low vision to provide content in an alternative format. The accessibility software relies on the accessibility tree being intact to facilitate providing content in an accessible format. Summary of the invention

[0004] The present disclosure describes a way to add information representing elements from displayable content to an accessibility tree that did not previously exist in the accessibility tree. The present disclosure describes an accessibility infrastructure that receives an image that represents displayable content for display by an application. A layout extraction model generates a list of elements for the image. Each element includes a bounding box that defines the portion of the image where the element is located, and a role attribute associated with the element. The accessibility infrastructure then adds the role attribute to the node in the accessibility tree. In some examples, the accessibility infrastructure may also add nodes or other attributes to the accessibility tree.

[0005] In some aspects, the technology described herein relates to a method comprising: receiving an image that represents displayable content for display by an application; executing a layout extraction model that uses the image as input and generates an element list of the image as output, the element list including at least a bounding box that defines a portion of the image, and a role attribute; and adding the role attribute to a node in an accessibility tree using the element list.

[0006] In some aspects, the technology described herein relates to a system comprising: an image receiving module configured to receive an image representing displayable content for display by an application; a layout extraction model configured to use the image as input and generate an element list of the image as output, the element list comprising at least a bounding box defining a portion of the image, and a role attribute; and an accessibility tree enhancement module configured to add the role attribute to a node in an accessibility tree using the element list.

[0007] In some aspects, the technology described herein relates to a computing device comprising: a processor; and a memory configured with instructions for: receiving an image representing displayable content for display by an application; executing a layout extraction model using the image as input and generating as output a list of elements of the image, the list of elements including at least a bounding box defining a portion of the image, and a role attribute; and adding the role attribute to a node in an accessibility tree using the list of elements. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1A A document accessibility scenario 100A is depicted.

[0009] Figure 1B A document accessibility scenario 100B is depicted.

[0010] Figure 1C An image-to-document accessibility scenario 100C is depicted according to examples described throughout this disclosure.

[0011] Figure 1D Depicted are displayable content including text according to examples described in this disclosure.

[0012] Figure 1E Depicted is an accessibility tree according to examples described in this disclosure.

[0013] Figure 2 A user device 202 is depicted according to examples described throughout this disclosure.

[0014] Figure 3 A flowchart 300 is depicted according to examples described throughout this disclosure.

[0015] Figure 4A A method 400A is depicted according to examples described throughout this disclosure.

[0016] Figure 4B A method 400B is depicted according to examples described throughout this disclosure.

[0017] Figure 4C A method 400C is depicted according to examples described throughout this disclosure.

[0018] Figure 4D A method 400D is depicted according to examples described throughout this disclosure. DETAILED DESCRIPTION

[0019] The present disclosure describes a manner in which information representing elements from displayable content is added to an accessibility tree that did not previously exist in the accessibility tree. The accessibility tree is used by accessibility software to provide access to displayable content. When the accessibility tree associated with displayable content is incomplete, the accessibility software cannot provide full access to the displayable content. For example, an image element in web displayable content may not include information about its semantic function or content that can be ported into the accessibility tree. In another example, a PDF file may not include text that can be ported into the accessibility tree. In both cases, the accessibility software will not be able to provide access to these image elements / documents.

[0020] The present disclosure describes an accessibility infrastructure that receives an image that represents displayable content for display by an application. A layout extraction model generates a list of elements for the image. Each element includes a bounding box that defines the portion of the image where the element is located, and a role attribute associated with the element. The accessibility infrastructure then adds the role attribute to a node in an accessibility tree. In some examples, the accessibility infrastructure may also add nodes or other attributes to the accessibility tree.

[0021] Many documents and many types of displayable content are primarily designed to be presented visually to users via software. Visual cues in many documents help users navigate the content. For example, subheadings, changes in the formatting of text, tables, or buttons all have semantic meanings that help users navigate the information. When this information is missing from the accessibility tree, the displayable content becomes difficult to navigate.

[0022] Content creators may sometimes take extra steps to create displayable content or documents that include the necessary information needed to generate a complete accessibility tree. For example, in HTML, creators can include HTML attributes that are well understood by the accessibility tree generation function in the browser. Alternatively, creators can include Accessible Rich Internet Applications (ARIA) tags in HTML to specify accessibility tree nodes or roles. However, many creators do not realize that they sometimes need to take extra steps to ensure that their content includes the necessary information to generate a complete accessibility tree. Unfortunately, this means that some displayable content and documents cannot be accessed by users via accessibility software.

[0023] Figure 1A A document accessibility scenario 100A is depicted. The example document accessibility scenario 100A depicts two branches: a first branch provides displayable content 110 for visual display, and a second branch provides an accessible version of the displayable content 110 through accessibility software 112 .

[0024] Document accessibility scenario 100A starts with HTML code 102 that includes displayable content. In the example, HTML code 102 can be executed on any web browser. Example HTML code 102 generates an image of a puppy with a name, a button with the word "Donate!", and an image of a letter icon with a link to an email address. However, HTML code 102 is not intended to be limiting. In the example, HTML code 102 can display any possible content. Instead of HTML code 102. Document accessibility scenario 100A can receive displayable content and any other format designed for display in an application executed on a user device.

[0025] The browser may use the HTML code 102 to generate a document object model DOM 104. The DOM 104 is a data structure representing the content specified by the HTML code 102. In the document accessibility scenario 100A, the DOM 104 includes nodes for images, buttons, and links.

[0026] When presenting information visually, the browser may generate a Cascading Style Sheet CSS 103 from the HTML code 102. The CSS 103 may be used to generate a CCS object model that provides information about the layout of displayed content (not shown).

[0027] The visual content processing module 106 can receive the DOM 104 and CSS 103 as input and generate a render tree. Next, the layout software can calculate the exact position and size of each object on the render tree. Finally, the drawing software can take the final render tree and generate pixels to visually display the content. Figure 1A In the example of , the visual version of displayable content 110 is displayed to include an image of a dog, a button with text saying "Donate!", and an icon representing a letter linked to an email address.

[0028] Accessibility software may alternatively provide an accessible version of the content via a second branch of the document accessibility scenario 100A. An application providing access to the content (in the example of the document accessibility scenario 100A, a browser) may make an API call to the accessibility infrastructure to provide an accessible version of the displayable content.

[0029] An application that provides access to content can use DOM 104 to generate an accessibility tree 108. Accessibility tree 108 is a data structure that maps out content that can be navigated by accessibility software. Accessibility tree 108 includes node 114. In an example, 114 may include more than one node. In the example of document accessibility scenario 100A, accessibility tree 108 includes three nodes, all of which share a single parent node, which is the root node. However, this is not intended to be limiting. In an example, the accessibility tree may include any number of nodes 114 connected via any possible configuration of parent nodes and child nodes. In an example, the root node may include a document node.

[0030] Each node 114 of the accessibility tree 108 includes attributes 116. In an example, the attributes 116 may include any number of attributes 116. For example, Figure 1A Each node 114 of the example accessibility tree 108 depicted in includes two attributes: role and name. However, this is not intended to be limiting. In the example, each node 114 may include any number of attributes 116. In the example, each node 114 may include a role attribute, a name attribute, a state attribute, and a description attribute. In the example, the nodes 114 of the accessibility tree 108 may include further attributes, such as a parent node id attribute, a child node id attribute, a position attribute. The nodes 114 may also include other metadata including details about text, location, how the individual nodes are represented, or any other features.

[0031] The role attribute may describe a semantic role, such as, for example, a title, button, table, paragraph, text box, combo box, list box, image, etc. In an example, the role attribute may include roles defined by any accessibility tree standard or definition, such as those defined by the ARIA standard. In a further example, the role attribute may include any classification of a set of elements in a display that have any semantic meaning.

[0032] In an example, the name attribute can be related to what the node represents. For example, the name can represent the text on a button, or a clue as to what is depicted in an image. In an example where the role attribute of the node is related to text, the name field can include that text.

[0033] Accessibility software 112 uses accessibility tree 108 to generate an accessible version of the displayable content. In the example of document accessibility scenario 100A, accessibility software 112 generates an audible version of the content that includes reading out the name attribute of each node followed by the role attribute: "Puppy image", "Donate button", and "Email link". However, Figure 1A The examples of are not intended to be limiting. In the examples, other types of accessibility software may be used. In the case of screen reader type accessibility software 112, the screen reader may use additional attributes connected to each node 114 to provide additional audible clues.

[0034] Although the document accessibility scenario 100A provides an example of converting displayable content in the form of HTML code into visual content and accessible content, this is not intended to be limiting. In the example, the document accessibility scenario 100A can convert displayable content in any other format into visual content and / or accessible content.

[0035] The challenge with accessibility in generating displayable content for display on a web browser is that it requires developers to use HTML tags and ARIA attributes in a way that matches the intent of their code. Lack of awareness and priority among developers, and failure to prioritize accessibility, is a major barrier to achieving an accessible web and accessible content. Prior assistive technologies often lack layout elements, such as custom controls (e.g., a button with a custom class). Tags) and common icon controls (e.g., browser back buttons). Prior assistive technologies also did not generate accessibility trees for image-based documents such as PDFs.

[0036] The problem is Figure 1B 100B is illustrated in FIG. 100B using document accessibility scenario 100B. Document accessibility scenario 100B uses different HTML code 120 to generate substantially the same displayable content 110. HTML code 120 is an example of HTML code in which the content creator did not include HTML tags or ARIA attributes to facilitate creation of a complete accessibility tree. For example, HTML code 120 provides an image "abc.jpg" without a name that provides a clue as to what the content of the image file is. HTML code 120 further provides a button link that is an image of a button "def.gif" instead of an HTML button tag. Finally, HTML code 120 provides an image "ghi.jpg" as a link to an email address without a link HTML tag.

[0037] Therefore, the DOM 122 generated from the HTML code 120 looks very different from the DOM 104. The DOM 104 includes an image, a button, and a link node, while the DOM 122 includes three image nodes. When the DOM 122 is transformed into the accessibility tree 124, each of the three example elements of the HTML code 120 has a role attribute of image and the name of the corresponding image file, which name is arbitrarily assigned in this case. Therefore, when the accessibility software 112 generates an audible version of the HTML code 120, the accessibility software 112 will read "image", "image", and "image". Therefore, the accessible version of the HTML code 120 does not include all the information available in the displayable content 110. Therefore, the document accessibility scenario 100A does not provide sufficient accessibility to the displayable content 110.

[0038] Figure 1C The document accessibility scenario 100C depicted in FIG. 1 also provides access to the HTML code 120 from the document accessibility scenario 100B. However, in the document accessibility scenario 100C, access to the HTML code 120 is significantly improved by including an accessibility infrastructure 126. The accessibility infrastructure 126 receives an image, generates a bounding box around the portion of the image that includes individual elements, and identifies the role attributes of the individual components. This is shown in FIG. Figure 1C Visually depicted in. Figure 1C In the example of FIG. 1 , displayable content 110 including a web page display is output from visual content processing module 106 and received at accessibility infrastructure 126. Displayable content 110 includes Figure 1A and Figure 1B : a dog image, a donate button, and a letter icon representing an email link. The accessibility infrastructure 126 generates at least one bounding box 130 within the image that includes the displayable content 110. In the example of the document accessibility scenario 100C, the accessibility infrastructure 126 identifies three instances of the bounding box 130 represented by the dashed lines, one bounding box 130 for each of the three respective elements. The accessibility infrastructure 126 then determines a role attribute for each respective element defined by the bounding box. Finally, the accessibility infrastructure 126 adds the role attribute to the node 114 in the accessibility tree 128. However, this is not intended to be limiting. In the example, adding the role attribute to the node 114 may include revising the role attribute. For example, the role attribute "image" may be revised to "button."

[0039] The resulting accessibility tree 128 can be compared to Figure 1B The accessibility tree 124 depicted in includes more detail and / or accuracy, making the displayable content accessible to the accessibility software 112. The accessibility software 112, which may be a screen reader, receives the accessibility tree 128 and may say that the image includes "Puppy image", "Donate! button", and "Email link". The accessibility infrastructure 126 is described in further detail below.

[0040] Figure 2 A user device 202 operable to perform the methods described herein is depicted. In an example, the user device 202 may include a laptop computer, a desktop computer, a handheld device such as a tablet computer, a mobile phone, a wrist-worn device such as a smart watch, or any other user device that provides access to content including images.

[0041] The user device 202 includes a processor 204, a communication interface 206, and a memory 208. In an example, the processor 204 may include multiple processors, and the memory 208 may include multiple memories. The processor 204 may be configured by instructions to perform the accessibility infrastructure described in the present disclosure. These instructions may include non-transitory computer-readable instructions stored in the memory 208 and called from the memory.

[0042] The communication interface 206 of the user device 202 may be operable to facilitate communication between the user device 202 and a server or another computing device. In an example, the communication interface 206 may utilize a short-range wireless communication protocol such as Bluetooth, Wi-Fi, Zigbee™, or any other wireless or wired communication method.

[0043] The memory 208 includes an application 210. In an example, the application may include a web browser, such as Google Chrome, Firefox, Microsoft Edge, or any other web browser. In an example, the application may include any application operable to display a document to a user, such as a PDF viewer, a word processor, etc. In an example, the application may include any application operable to provide a user with access to a document or displayable content.

[0044] Application 210 includes accessibility infrastructure 212. Accessibility infrastructure 212 is operable to receive an image, identify elements in the image via bounding boxes, generate role attributes for the elements, and update the role attributes for nodes in an accessibility tree.

[0045] In an example, the accessibility infrastructure 212 can execute one or more software modules. In an example, the accessibility infrastructure 212 can execute any combination of the image receiving module 222, the layout extraction model 224, the OCR model 226, the icon model 228, the other models 230, and / or the accessibility tree enhancement module 232. Figure 2 In the example of , the image receiving module 222, the layout extraction model 224, the OCR model 226, the icon model 228, the other models 230, and the accessibility tree enhancement module 232 are included in a library 220 separate from the application 210. In the example, the library 220 may include separate software executable by the accessibility infrastructure 212 via, for example, one or more API calls. However, Figure 2 The examples are not intended to be limiting. In the examples, any one of the image receiving module 222, the layout extraction model 224, the OCR model 226, the icon model 228, the other models 230, and the accessibility tree enhancement module 232 can be incorporated into the object code or executable code of the accessibility infrastructure 212 or the library 220.

[0046] In the example, any of the image receiving module 222, the layout extraction model 224, the OCR model 226, the icon model 228, the other models 230, and the accessibility tree enhancement module 232 can be executed on a server (not depicted) that is available to the accessibility infrastructure 212 via the network and communication interface 206. In the case where any of the image receiving module 222, the layout extraction model 224, the OCR model 226, the icon model 228, the other models 230, and the accessibility tree enhancement module 232 can be executed on the server, a control is provided to the user that allows the user to select what image data or other displayable content data can be sent to the server. In addition, before storing or using certain data, the data can be processed in one or more ways so that user information is removed. For example, the user identity can be processed so that the user information of the user cannot be determined, or the user's geographic location in which the location information is obtained can be generalized (such as generalized to the city, zip code, or state level) so that the user's specific location cannot be determined based on the image data or displayable content data submitted to the server. Thus, users can control what information is collected about them, how that information is used, and what information is provided to them.

[0047] Figure 3 A flowchart 300 is depicted. The flowchart 300 depicts execution of the accessibility infrastructure 126. The flowchart 300 begins when the application 210 sends an image 302 to the accessibility infrastructure 126. In an example, the image 302 may be received in response to determining that the accessibility tree of the displayable content 110 has a node 114 with an image or document role. In an example, the image 302 may be received in response to determining that the node 114 is missing an attribute 116 or has a null value for the attribute 116. For example, as described above with respect to the example accessibility tree 124, each of the three nodes includes an image without a name attribute. In an example of receiving a single document for display, such as a PDF document, the accessibility tree may include only a root node representing the document without an attribute. In an example, the image 302 may be received in response to a user command. For example, a user using a screen reader to navigate the displayable content 110 within the application 210 may suspect that their screen reader is missing information and initiate the steps of the flowchart 300 via a user command.

[0048] Flowchart 300 depicts layout extraction model 224 receiving image 302. For example, layout extraction model 224 may receive image 302 via image receiving module 222. Image receiving module 222 is operable to receive an image from application 210.

[0049] The image 302 may represent displayable content 110 for display by the application 210. In an example, the image 302 may include any type of image file, including but not limited to PDF, JPEG, JPG, GIF, TIF, BMP, or any other image format that undergoes at least one of rendering, layout, or drawing in order to be displayed on a computer display via the visual content processing module 106. In an example, the displayable content 110 may include any combination of HTML elements designed for display in an application. For example, the displayable content may include a custom control (e.g., a custom class used as a button) or a custom HTML element. Labels) and common icon controls (e.g., browser back buttons). Custom controls can include images, buttons, text boxes, tables, radio buttons, titles, paragraphs, links, etc. The examples provided are not intended to be limiting. In the examples, images can include any type of displayable content.

[0050] The layout extraction model 224 is operable to use the image 302 as input and generate as output an element list 306 of the image 302. The element list 306 includes at least a bounding box 130 defining a portion of the image 302 and a role attribute associated with the bounding box 130.

[0051] The layout extraction model 224 may include a machine learning model that is trained, for example using supervised or semi-supervised training, to take an image 302 and identify a list of elements from the image 302 , each element including a bounding box 130 and an associated role attribute.

[0052] In an example, the layout extraction model 224 may further receive the DOM 304, and executing the layout extraction model may further include using the document object model as input. In an example, the DOM 304 may be generated by the application 210. In an example, the DOM 304 may be used by the layout extraction model 224 to generate the element list 306.

[0053] In an example, it can be determined that the role attribute is text-related. The text-related attributes can include a title, a paragraph, a word, a combo box, a list box, a text box, a static text, or any other text-related displayable content that mainly includes text with formatting.

[0054] In response to determining that the portion of image 302 includes character attributes associated with text, layout extraction model 224 can execute OCR model 226 on the portion of image 302 and analyze the output of OCR model 226 to generate text spans. A text span can include a text string of characters.

[0055] In an example, the OCR model 226 may include a machine learning model that is trained, for example using supervised or semi-supervised training, to receive the portion of the image 302 and identify text within the portion of the image 302 .

[0056] In an example, the OCR model 226 can output an OCR text tree. The OCR text tree can include a root node representing an image document. The child nodes of the root node can include paragraph nodes. The child nodes of the paragraph node can include sentence nodes. The child nodes of the sentence node can include word nodes. The child nodes of the word node can include character nodes. Metadata can be further included with the nodes of the OCR text tree, and the metadata includes information about text location and text formatting.

[0057] The OCR text tree is not in a format that can be stitched into the accessibility tree 312. Therefore, in an example, the layout extraction model 224 can perform further processing on the OCR text tree to generate text spans. In an example, analyzing the OCR text tree output by the OCR model 226 to generate text spans can further include determining the formatting applied to consecutive character groups in the output of the OCR model. The text spans can then be set to consecutive character groups with the same formatting.

[0058] For example, Figure 1D An example displayable content 140 is depicted. The displayable content 140 may include a PDF file displayed on a web browser. The displayable content 140 includes a title, three headings, and corresponding paragraphs under each of the three headings. The layout extraction model 224 has drawn a bounding box 130 around the entire area of ​​the displayable content 140 including the text, thereby capturing substantially all of the displayable content 140. However, in the example, the layout extraction model 224 may instead draw a bounding box around each or any adjacent combination of the title, three headings, and three paragraphs.

[0059] In the example, the OCR model 226 can seek to identify the text within the bounding box 130 by starting from the upper left corner and sweeping to the right along the surface area of ​​similar lines, repeating this action for the sequential lower lines until the OCR model 226 reaches the bottom of the bounding box 130 to find the text. In this way, the OCR model 226 can mimic reading English. However, this is not intended to be limiting. For other languages, the OCR model 226 can evaluate the areas of the bounding box 130 in a different order, such as from right to left for Hebrew. In the example, the OCR model 226 can use any other method to evaluate the area represented by the bounding box 130 for text.

[0060] The OCR text tree output by the OCR model 226 for the displayable content 140 may include the text represented in the displayable content 140, as well as metadata identifying the formatting of the text. The formatting may include, for example, font size, font type, font color, font styles such as bold, italic, underline, strikethrough, subscript, and superscript. By traversing the OCR text tree, the layout extraction model 224 may thus evaluate all text found in the bounding box 130 starting from the upper left corner and sweeping to the right in subsequent lines until the bottom of the bounding box 130.

[0061] In an example, layout extraction model 224 can determine formatting to apply to consecutive groups of characters while traversing the OCR text tree. For example, the title of displayable content 140, "A Guide to British Cheese Varieties," includes consecutive groups of characters that all have the same formatting. Subsequent lines of text in displayable content 140 have different formatting. Therefore, the characters of the title can be grouped together into a single span of text by layout extraction model 224 to be included in accessibility tree 312 separately from other text found in displayable content 140.

[0062] In an example, the layout extraction model 224 may set a name attribute to the text span. In an example, the accessibility infrastructure 212 may access the name attribute to share the content of the text span with the user. However, in other examples, the layout extraction model 224 may generate a child node from the node 310. The node 310 may have a role attribute of a title, a paragraph, or any other attribute related to text. The child node may have a role attribute of static text and a name attribute set to the text span.

[0063] Figure 1E An example accessibility tree 132 generated by the accessibility infrastructure 212 based on the displayable content 140 is depicted. The accessibility tree 132 includes a root node 152 having a role attribute of a document. In an example, the root node 152 may instead include a role attribute of an image.

[0064] Root node 152 includes child node 154. Child node 154 may include text-related role attributes. In an example, child node 154 may have a role attribute of title or static text. In an example, child node 154 may have a name attribute of "A Guide to British Cheese Varieties".

[0065] The child node 154 includes three further child nodes 156. The child nodes 156 may each include a corresponding text-related role attribute, such as a title or static text. In the example, the child nodes 156 may include name attributes of "Stilton", "CornishYarg" and "Red Leicester".

[0066] The child nodes 156 may each include a corresponding child node 158. In an example, the child nodes 158 may each include a corresponding text-related role attribute, such as paragraph or static text. In an example, the name attribute of the child nodes 158 may each include Figure 1D The contents of the paragraphs found under their corresponding parent node headings.

[0067] In an example, the child nodes 158 may each include a corresponding child node 160. In an example, the child nodes 160 may each include a corresponding role attribute of the static text and include Figure 1D The corresponding name attribute of the content of the paragraph under their corresponding parent node title. The child node 160 includes Figure 1D In the case of a name attribute in the content of a paragraph, the child node 158 may not include Figure 1D The name attribute of the content of the paragraph.

[0068] In an example, the OCR model 226 may determine text-related role attributes of the displayable content 140, such as, for example, a title, a heading, or a paragraph. In an example, the layout extraction model 224 may determine the text-related role attributes of the displayable content 140 based on the image 302. In an example, the layout extraction model 224 may determine the text-related role attributes of the displayable content 140 based on metadata found in the OCR text tree.

[0069] In an example, the order and / or arrangement of nodes in the accessibility tree 132 may be determined based on at least one of the order in which each node is encountered when traversing the OCR text tree, role attributes assigned to each node, and / or OCR model metadata associated with each node.

[0070] In an example, it may be determined that the role attribute includes an icon-related attribute. The icon-related attribute may include, for example, an envelope symbol for an email link, a back arrow for navigating back to a previous website, the letter "i" for a link to information, etc. In response to determining that the portion of the image includes an icon-related attribute, the layout extraction model 224 may execute the icon model 228 on the portion of the image 302 to identify the role attribute associated with the icon-related attribute. For example, Figure 1A and Figure 1B An example of an icon of an envelope that is a link to an email address is provided.

[0071] In an example, the layout extraction model 224 can determine that the role attribute includes a button. In an example, the layout extraction model 224 can identify text within the portion of the image. The text can be used to set the name attribute 308.

[0072] In an example, the layout extraction model 224 may determine that the attribute includes a navigation bar, a toolbar, or any other control or widget.

[0073] In an example, layout extraction model 224 can call other models 230 to identify other elements in image 302 and add those elements to the list of elements.

[0074] In an example, the role attribute may be a first role attribute, and the element list may further include a second role attribute. In an example, the layout extraction model 224 may generate a node 310. In an example, the node may be a first node associated with the first role attribute, and the accessibility tree may further include a second node associated with the second role attribute.

[0075] In an example, the layout extraction model 224 may not be able to determine the role attribute of the element. In an example where there is no role attribute for an element identified in the image 302, the layout extraction model 224 may not add the element to the element list 306 or to the accessibility tree 312.

[0076] After executing the layout extraction model 224, the accessibility tree enhancement module 232 receives the element list 306. In an example, the accessibility tree enhancement module 232 may further receive any combination of name attributes 308 and / or nodes 310. The accessibility tree enhancement module 232 adds a role attribute to at least one node in the accessibility tree 312. In an example, the accessibility tree enhancement module 232 may also add the node 310 generated by the layout extraction model 224 to the accessibility tree 312.

[0077] In an example, the accessibility tree enhancement module 232 may stitch the node 310 into the accessibility tree 312 based on the role attribute of the node 310. In an example, the position at which the node 310 is stitched into the accessibility tree 312 relative to other nodes may be determined based on one or more attributes of the other nodes. In an example, the position at which the node 310 is stitched into the accessibility tree 312 relative to other nodes may be determined based on metadata associated with the element list 306. In an example, the metadata may be received from the OCR model 226 or the layout extraction model 224. In an example, the metadata may indicate the position and / or size of the bounding box 130 or text associated with the element. In an example, the metadata may include font and formatting information associated with the element, or any other clues operable to provide semantic and / or layout information related to the relationship between elements in the image 302.

[0078] In an example, the layout extraction model 224 can evaluate the image 302 and generate a new accessibility tree 132 associated with the image 302. The new accessibility tree 132 can then be stitched into the previous accessibility tree at the image node associated with the image 302.

[0079] Figure 4A A method 400A according to an example is depicted. The method 400A is operable to determine an element list of an image, the element list comprising a bounding box defining a portion of the image, and a role attribute. The method 400A is further operable to add the role attribute to a node in an accessibility tree.

[0080] The method 400A begins at step 402. In step 402, an image 302 is received, the image 302 representing displayable content 110 for display by the application 210, as described above.

[0081] Method 400A continues with step 406. In step 406, layout extraction model 224 is executed using image 302 as input, generating element list 306 that includes at least bounding box 130 defining a portion of image 302, and role attributes, as described above.

[0082] Method 400A continues with step 408. In step 408, the role attribute is added to the node in accessibility tree 312, as described above.

[0083] In an example, method 400A may further include steps 404 and 410. In step 404, DOM 304 is received, as described above.

[0084] In step 410, the node 310 is added to the accessibility tree 312 based on the role attributes, as described above.

[0085] Figure 4B Depicting method 400B, Figure 4C A method 400C is depicted, and Figure 4D Method 400D is depicted. In an example, step 406 of method 400A may further include any of the steps of methods 400B, 400C, and 400D.

[0086] Method 400B begins at step 420. In step 420, it is determined whether the role attribute is associated with the text, as described above.

[0087] If the answer to step 420 is yes, then step 422 is performed. In step 422, the OCR model 226 is performed on the portion of the image 302, as described above.

[0088] Method 400B may continue with step 424. In step 424, the output of the OCR model is analyzed to generate text spans, as described above.

[0089] Method 400B may continue with step 426. At step 426, formatting of consecutive character groups in the output of the OCR model is determined, as described above.

[0090] Method 400B may proceed to step 428. In step 428, the text span may be set to the group of consecutive characters, as described above.

[0091] Method 400B may continue with step 430. In step 430, a name attribute may be set to the text span, as described above.

[0092] Method 400C begins at step 430. In step 430, it is determined whether the character attributes include icon-related attributes, as described above.

[0093] If the answer to step 430 is yes, then step 432 is performed. In step 432, the icon model 228 is executed on the portion of the image to generate character attributes associated with the icon-related attributes, as described above.

[0094] Method 400D begins at step 440. In step 440, it is determined whether the character attributes include buttons, as described above.

[0095] If step 440 evaluates yes, method 400D may proceed to step 442. In step 442, text within the portion of the image may be identified, as described above.

[0096] In an example, method 400D may proceed to step 444. In step 444, the text may be added to the name attribute of node 310 in accessibility tree 312, as described above.

[0097] In an example, the method described herein may include further steps. For example, after adding at least one of the role attributes or nodes associated with the element to the accessibility tree 312, the accessibility infrastructure 212 may execute user commands using the accessibility tree 312, such as finding the element in the page, selecting the element, and copying the element. In an example, the accessibility infrastructure 212 may use the enhanced accessibility tree 312 to perform any functions that accessibility software typically performs.

[0098] The methods described herein can allow previously unavailable displayable content to be included in an accessibility tree. The methods described do not rely on any particular user application to execute. In particular, when incorporated into a library accessible to a user application, these methods can allow any accessibility software to provide improved access to displayable content relative to prior rule-based models.

[0099] Various examples of the systems and techniques described herein may be implemented in digital electronic circuit systems, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various examples may include examples in one or more computer programs executable and / or interpretable on a programmable system, the programmable system including at least one programmable processor, the at least one programmable processor may be dedicated or general purpose, and may be coupled to receive data and instructions from a storage system, at least one input device, and at least one output device and transmit data and instructions to the storage system, at least one input device, and at least one output device. Various examples of the systems and techniques described herein may be implemented as and / or are generally referred to herein as circuits, modules, blocks, or systems that can be combined in software and hardware. For example, a module may include functions / actions / computer program instructions executed on a processor or some other programmable data processing device.

[0100] Some of the examples above are described as processes or methods depicted as flow charts. Although the flow charts describe the operations as sequential processes, many operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. Processes can terminate when their operations are completed, but can also have additional steps not included in the figure. A process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0101] The methods discussed above (some of which are illustrated by flow charts) may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored in a machine or computer readable medium, such as a storage medium. A processor may perform the necessary tasks.

[0102] The specific structural and functional details disclosed herein are merely representative for describing the examples. However, the examples are embodied in many alternative forms and should not be construed as being limited to only the implementations set forth herein.

[0103] It should be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the scope of the example implementation, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. As used herein, the term "and / or" includes any and all combinations of one or more items in the associated listed items.

[0104] The terms used herein are for the purpose of describing specific implementations only and are not intended to limit the example implementations. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should be further understood that the terms "includes", "comprising", "includes", and / or "including" when used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0105] It should also be noted that in some alternative examples, the functions / actions noted may not occur in the order noted in the figures. For example, depending on the functions / actions involved, two figures shown in succession may actually be executed simultaneously or may sometimes be executed in the reverse order.

[0106] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the example implementations belong. It should be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless explicitly so defined herein.

[0107] Portions of the above example implementations and corresponding detailed descriptions are presented in terms of software or algorithms and symbolic representations of operations on data bits in computer memory. These descriptions and representations are descriptions and representations used by those of ordinary skill in the art to effectively convey the substance of their work to other persons of ordinary skill in the art. As the term is used herein, and as generally used, an algorithm is considered to be a self-consistent sequence of steps that produces a desired result. The step is a step that requires physical manipulation of physical quantities. Although not necessary, these quantities are typically in the form of optical, electrical, or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated. It has been demonstrated that, primarily for common usage reasons, it is sometimes convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.

[0108] In the above exemplary implementations, references to symbolic representations of actions and operations that can be implemented as program modules or functional processes (e.g., in the form of flowcharts) include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types, and can be described and / or implemented using existing hardware at existing structural elements. Such existing hardware may include one or more central processing units (CPUs), digital signal processors (DSPs), application specific integrated circuits, field programmable gate arrays (FPGAs), computers, etc.

[0109] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specifically noted or as is apparent from the discussion, terms such as processing or computing or calculating or determining display refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical electronic quantities within the computer system's registers and memories and transforms that data into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage, transmission or display devices.

[0110] It is also noted that the software-implemented aspects of the example implementations are typically encoded on some form of non-transitory program storage medium or implemented on some type of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or hard disk) or optical (e.g., a compact disk read-only memory or CD ROM), and may be read-only or random access. Similarly, the transmission medium may be a twisted pair, a coaxial cable, an optical fiber, or some other suitable transmission medium known in the art. The examples are not limited by these aspects of any given example.

[0111] Finally, it should also be noted that although the appended claims set forth specific combinations of features described herein, the scope of the present disclosure is not limited to the specific combinations claimed below, but extends to cover any combination of features or examples disclosed herein, regardless of whether the specific combination is specifically listed in the appended claims.

[0112] In some aspects, the technology described herein relates to a method further comprising: determining that a character attribute is associated with text; and in response to determining that the portion of the image is associated with text, executing an OCR model on the portion of the image and analyzing an output of the OCR model to generate a text span, and setting a name attribute to the text span.

[0113] In some aspects, the technology described herein relates to a method in which analyzing the output of the OCR model to generate a text span includes: determining formatting applied to a consecutive character group in the output of the OCR model; and setting the text span to the consecutive character group.

[0114] In some aspects, the technology described herein relates to a method further comprising: determining that the role attribute includes an icon-related attribute; and in response to determining that the portion of the image includes the icon-related attribute, executing an icon model on the portion of the image to generate the role attribute associated with the icon-related attribute.

[0115] In some aspects, the technology described herein relates to a method further comprising: determining that a character attribute includes a button.

[0116] In some aspects, the technology described herein relates to a method further comprising: identifying text within the portion of the image; and adding the text to a name attribute of the node in the accessibility tree.

[0117] In some aspects, the technology described herein relates to a method, wherein the role attribute is a first role attribute and the list of elements further includes a second role attribute.

[0118] In some aspects, the technology described herein relates to a method further comprising: adding the node based on the role attribute.

[0119] In some aspects, the technology described herein relates to a method in which the image is received in response to determining that the accessibility tree of the displayable content displayed by the application has an image node or a document node with a missing attribute.

[0120] In some aspects, the technology described herein relates to a method in which the image comprises an image-based portable document format document.

[0121] In some aspects, the technology described herein relates to a method in which the image is received in response to a user command.

[0122] In some aspects, the techniques described herein relate to a method further comprising: receiving a document object model, and wherein executing the layout extraction model further comprises using the document object model as input.

[0123] In some aspects, the technology described herein relates to a method, wherein the method is performed on a browser extension or browser plug-in.

[0124] In some aspects, the technology described herein relates to a system wherein the layout extraction model is further configured to: determine that the role attribute is associated with text; and in response to determining that the portion of the image is associated with text, execute an OCR model on the portion of the image and analyze the output of the OCR model to generate a text span, and set the name attribute to the text span.

[0125] In some aspects, the technology described herein relates to a system in which the layout extraction model is further configured to: determine formatting applied to a consecutive character group in the output of the OCR model; and set the text span to the consecutive character group.

[0126] In some aspects, the technology described herein relates to a system wherein the layout extraction model is further configured to: determine that the role attribute includes an icon-related attribute; and in response to determining that the portion of the image includes the icon-related attribute, execute an icon model on the portion of the image to generate the role attribute associated with the icon-related attribute.

[0127] In some aspects, the techniques described herein relate to a system wherein the layout extraction model is further configured to determine that the role attribute comprises a button.

[0128] In some aspects, the techniques described herein relate to a system in which the layout extraction model is further configured to: identify text within the portion of the image; and add the text to a name attribute of the node in the accessibility tree.

[0129] In some aspects, the technology described herein relates to a system wherein the role attribute is a first role attribute and the list of elements further includes a second role attribute.

[0130] In some aspects, the technology described herein relates to a system wherein the accessibility tree enhancement module is further configured to add the node based on the role attribute.

[0131] In some aspects, the techniques described herein relate to a system in which the image is received in response to determining that the accessibility tree of the displayable content displayed by the application has an image node or a document node with a missing attribute.

[0132] In some aspects, the technology described herein relates to a system wherein the image comprises an image-based portable document format document.

[0133] In some aspects, the techniques described herein relate to a system in which the image is received in response to a user command.

[0134] In some aspects, the techniques described herein relate to a system, wherein the system further includes a document object model receiving module, and the layout extraction model further uses the document object model to determine the element list.

[0135] In some aspects, the techniques described herein relate to a system in which the layout extraction model is executed on a browser extension or browser plug-in.

[0136] In some aspects, the technology described herein relates to a computing device wherein the memory is further configured with instructions to: determine that the role attribute is associated with text; and in response to determining that the portion of the image is associated with text, execute an OCR model on the portion of the image and analyze the output of the OCR model to generate a text span, and set the name attribute to the text span.

[0137] In some aspects, the technology described herein relates to a computing device wherein analyzing the output of the OCR model to generate a text span includes determining formatting to apply to a consecutive group of characters in the output of the OCR model, and the memory is further configured with instructions to: set the text span to the consecutive group of characters.

[0138] In some aspects, the technology described herein relates to a computing device wherein the memory is further configured with instructions to: determine that the character attribute includes an icon-related attribute; and in response to determining that the portion of the image includes the icon-related attribute, execute an icon model on the portion of the image to generate the character attribute associated with the icon-related attribute.

[0139] In some aspects, the technology described herein relates to a computing device, wherein the memory is further configured with instructions to: determine that the character attribute includes a button.

[0140] In some aspects, the techniques described herein relate to a computing device wherein the memory is further configured with instructions to: identify text within the portion of the image; and add the text to a name attribute of the node in the accessibility tree.

[0141] In some aspects, the techniques described herein relate to a computing device in which the role attribute is a first role attribute and the list of elements further includes a second role attribute.

[0142] In some aspects, the techniques described herein relate to a computing device, wherein the memory is further configured with instructions to: add the node based on the role attribute.

[0143] In some aspects, the techniques described herein relate to a computing device in which the image is received in response to determining that the accessibility tree of the displayable content displayed by the application has an image node or a document node with a missing attribute.

[0144] In some aspects, the techniques described herein relate to a computing device in which the image comprises an image-based portable document format document.

[0145] In some aspects, the techniques described herein relate to a computing device in which the image is received in response to a user command.

[0146] In some aspects, the techniques described herein relate to a computing device, wherein the memory is further configured with instructions to: receive a document object model, and wherein executing the layout extraction model further includes using the document object model as input.

[0147] In some aspects, the techniques described herein relate to a computing device in which the instructions are executed on a browser extension or browser plug-in.

Claims

1. A method comprising: receiving an image, the image representing displayable content for display by an application; executing a layout extraction model that uses the image as input and generates as output a list of elements of the image, the list of elements comprising at least a bounding box defining a portion of the image, and role attributes; as well as Adds the role attribute to a node in the accessibility tree using the list of elements.

2. The method of claim 1, further comprising: Determining that the character attributes are associated with the text; as well as In response to determining that the portion of the image is associated with text, an OCR model is executed on the portion of the image and an output of the OCR model is analyzed to generate a text span, and a name attribute is set to the text span.

3. The method of claim 2, wherein analyzing the output of the OCR model to generate the text spans comprises determining formatting to apply to consecutive groups of characters in the output of the OCR model; and The text span is set to the consecutive character group.

4. The method of claim 1, further comprising: Determining that the role attributes include icon-related attributes; as well as In response to determining that the portion of the image includes the icon-related attribute, an icon model is executed on the portion of the image to generate the role attribute associated with the icon-related attribute.

5. The method of claim 1, further comprising: Determining that the role attributes include buttons; identifying text within the portion of the image; as well as The text is added to the name attribute of the node in the accessibility tree.

6. The method according to any one of claims 1 to 5, further comprising: The node is added based on the role attribute.

7. The method of any one of claims 1 to 6, wherein the image is received in response to determining that the accessibility tree of the displayable content displayed by the application has an image node or a document node with a missing attribute.

8. The method of claim 7, wherein the image comprises an image-based portable document format document.

9. The method of any one of claims 1 to 8, wherein the image is received in response to a user command.

10. The method according to any one of claims 1 to 9, further comprising: A document object model is received, and wherein executing the layout extraction model further comprises using the document object model as input.

11. The method of any one of claims 1 to 10, wherein the method is performed on a browser extension or a browser plug-in.

12. A system comprising: an image receiving module, the image receiving module being configured to receive an image, the image representing displayable content for display by an application; a layout extraction model configured to use the image as input and generate as output a list of elements of the image, the list of elements comprising at least a bounding box defining a portion of the image, and role attributes; an accessibility screen module configured to execute the layout extraction model; and An accessibility tree enhancement module is configured to add the role attribute to a node in an accessibility tree using the element list.

13. The system of claim 12, wherein the layout extraction model is further configured to: determine that the role attribute is associated with text, and in response to determining that the portion of the image is associated with text, execute an OCR model on the portion of the image and analyze the output of the OCR model to generate a text span, and set a name attribute to the text span.

14. The system of claim 13, wherein the layout extraction model is further configured to determine formatting applied to consecutive character groups in the output of the OCR model and to set the text span to the consecutive character groups.

15. A system as described in claim 12, wherein the layout extraction model is further configured to: determine that the role attribute includes an icon-related attribute, and in response to determining that the portion of the image includes the icon-related attribute, execute an icon model on the portion of the image to generate the role attribute associated with the icon-related attribute.

16. The system of claim 12, wherein the layout extraction model is further configured to: determine that the role attribute includes a button, identify text within the portion of the image, and add the text to a name attribute of the node in the accessibility tree.

17. The system of any one of claims 12 to 16, wherein the role attribute is a first role attribute and the element list further includes a second role attribute.

18. The system of any one of claims 12 to 17, wherein the accessibility tree enhancement module is further configured to add the node based on the role attribute.

19. The system of any one of claims 12 to 18, wherein the image is received in response to determining that the accessibility tree of the displayable content displayed by the application has an image node or a document node with a missing attribute.

20. The system of claim 19, wherein the image comprises an image-based portable document format document.

21. The system of any one of claims 12 to 20, wherein the image is received in response to a user command.

22. The system of any one of claims 12 to 21, wherein the system further comprises a document object model receiving module, and the layout extraction model further uses the document object model to determine the element list.

23. A computing device comprising: processor; as well as A memory configured with instructions for: receiving an image, the image representing displayable content for display by an application, executing a layout extraction model that uses the image as input and generates as output a list of elements of the image, the list of elements comprising at least a bounding box defining a portion of the image, and role attributes, and Adds the role attribute to a node in the accessibility tree using the list of elements.