Feature discovery layer

Through the operating system discovery mode, the content analysis engine is used to identify and highlight the operating system and third-party features in the desktop area, solving the problem that software applications fail to make full use of these features, and improving user productivity and computing resource usage efficiency.

CN120476381APending Publication Date: 2025-08-12MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480006368.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-29
Filing Date
2024-02-27
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Many software applications fail to take full advantage of the features provided by operating systems and third parties, resulting in low user productivity and inefficient use of computing resources.

Method used

Through the operating system discovery mode, the content analysis engine identifies and highlights the operating system and third-party features within the desktop area, providing visual prompts for users to call these features interactively.

Benefits of technology

It improves users' utilization of operating system and third-party features, enhances user productivity and optimizes the efficiency of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476381A_ABST
    Figure CN120476381A_ABST
Patent Text Reader

Abstract

An operating system (OS) discovery mode is disclosed that identifies and provides access to OS-provided features and / or third party-provided features within an applicable area of a desktop. In some configurations, upon activation of a discovery mode, content displayed by an application area is analyzed to identify content available to features provided by an OS and / or features provided by a third party. The visual cues appear in the applicable area in the vicinity of the identified content, highlighting the availability of the OS-provided features and / or the third-party-provided features. A user may interact with the visual cues to manipulate underlying content or invoke OS-provided features and / or third-party-provided features. Features provided by the OS and / or features provided by the third party may modify content displayed by the application, initiate an inline micro experience, cut or derive an image, etc. In the discovery mode, when it is found that a mouse cursor moves around on the desktop, a visual prompt is highlighted.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Software applications are programmed using the functionality provided by the operating system (OS), such as reading and writing files, sending and receiving data over a network, and the like. However, many applications do not fully utilize the full set of features provided by the OS. For example, an application may have been developed before the features provided by the OS are available. In other cases, application developers may not have the time, resources, or interest to utilize the features provided by a particular OS or the features provided by a third party. The features provided by the OS and / or the features provided by a third party may also not be used because the user does not know how to call them or because the user is not aware of the features. These features provided by the OS and / or the features provided by the third party increasingly utilize machine learning techniques and / or other artificial intelligence techniques. Among other shortcomings, the failure to fully utilize the features provided by the OS and / or the features provided by the third party limits the productivity of the user and makes the use of computing resources inefficient.

[0002] It is with these and other considerations that the disclosures made herein are made. Summary of the Invention

[0003] An operating system (OS) discovery mode is disclosed that identifies and provides access to OS-provided features and / or third-party-provided features within an applicable area of a desktop. In some configurations, once discovery mode is activated, content displayed by the applicable area is analyzed to identify content that can be used by the OS-provided features and / or third-party-provided features. Visual cues are rendered in the applicable area near the identified content, highlighting the availability of the OS-provided features and / or third-party-provided features. A user can interact with the visual cues to manipulate the underlying content or invoke the OS-provided features and / or third-party-provided features. The OS-provided features and / or third-party-provided features can modify the content displayed by the application, launch inline micro-experiences, crop or export images, etc. When in discovery mode, the visual cues are highlighted when the mouse cursor is discovered moving around on the desktop. In some configurations, discovery mode is automatically triggered so that the OS service automatically identifies and displays a set of visual cues across the applicable areas.

[0004] Features and technical advantages that are not explicitly described above will become clear by reading the following detailed description and examining the associated drawings. This summary is provided to introduce a selection of concepts in a simplified form that are further described in the following detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. For example, the term "technology" may refer to system(s), method(s), computer-readable instructions, module(s), algorithms, hardware logic, and / or operation(s) as allowed above, below, and throughout the document. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The detailed description is described with reference to the accompanying drawings. In the drawings, the leftmost digit(s) of a reference number identifies the drawing in which the reference number first appears. The same reference numbers in different drawings indicate similar or identical items. References to individual items of a plurality of items may use reference numbers with alphabetical letters to refer to each individual item. General references to items may use specific reference numbers without alphabetical order.

[0006] Figure 1A Active windows of applications running on the desktop are shown.

[0007] Figure 1B Shown is the discovery mode activated.

[0008] Figure 1C Shows a visual cue overlaid on an entity in the active window.

[0009] Figure 1D Shown is a visual cue that highlights an image entity in response to a discovery cursor hover.

[0010] Figure 1E A visual cue is shown highlighting text nested within an image entity.

[0011] Figure 1F An entity action list is shown displayed in response to selecting a nested text visual cue.

[0012] Figure 1G Deactivation of discovery mode is shown.

[0013] Figure 2A Shown is a paragraph entity within a highlighted text entity.

[0014] Figure 2B Shown is a list of entity actions displayed in response to selecting a highlighted paragraph entity.

[0015] Figure 2C Shown is selecting an action from an entity action list.

[0016] Figure 2D The selected paragraph is shown after it has been modified by the selected action.

[0017] Figure 3A The subject entity is shown highlighted in response to a cursor hovering over the subject entity.

[0018] Figure 3B A drag operation is shown to extract a subject entity from an image.

[0019] Figure 4A An image entity is shown highlighted in response to hovering over the image.

[0020] Figure 4B Shown is a list of entity actions displayed in response to receiving a selection of an image entity (outside of a subject entity).

[0021] Figure 4C The image is shown after the selected action has removed the background.

[0022] Figure 5A Shown is an entity action list displayed in response to receiving a selection of an address entity.

[0023] Figure 5B An inline micro-experience is shown displayed in response to receiving a selection of an address entity action.

[0024] Figures 6A to 6C A particular visual cue or a particular portion of a visual cue is shown highlighted proximate to a discovery cursor.

[0025] Figure 7 Shown is automatically triggering a discovery mode where an address entity visual cue is overlaid on top of an active window without requiring the user to explicitly activate the discovery mode.

[0026] Figure 8 is a flow chart of an example method for a feature discovery layer.

[0027] Figure 9 is a computer architecture diagram showing an illustrative computer hardware and software architecture for a computing system capable of implementing various aspects of the techniques and technologies presented herein.

[0028] Figure 10 is a diagram illustrating a distributed computing environment capable of implementing aspects of the techniques and technologies presented herein. DETAILED DESCRIPTION

[0029] Figure 1AAn active window 110 of an application 116 is shown. Application 116 runs within desktop 100, which is the user interface of operating system 106. Active window 110 partially obscures an inactive window 112, also running within desktop 100. Inactive window 112 may be part of application 116 or a different application. Desktop 100 includes menu 101, which contains, among other things, item icons for active and / or pinned applications, a search bar, and an activation button 102. In the Microsoft Windows operating system, menu 101 may also be a taskbar, title bar, or ribbon menu.

[0030] Activation button 102 is a graphical button found on menu 101. In some operating systems, such as MICROSOFT WINDOWS, selecting activation button 102 displays a menu that provides access to various applications, files, and settings on the computer. Typically, keyboards have a dedicated key or key combination that, when pressed, triggers activation button 102. Activation button 102 can also be selected by moving a cursor 120 over activation button 102 and selecting it, using voice activation, or by any other technique commonly used to select buttons.

[0031] The cursor 120 represents the current position of the mouse pointer on the desktop 100. As shown, the cursor 120 can be represented by the default operating system pointer icon. The cursor 120 is used to select, drag, and drop objects on the screen. The shape and appearance of the cursor 120 can change depending on the type of operation being performed. For example, moving the cursor over a link in a web browser can cause the cursor to change to a hand icon. The shape and appearance of the cursor can also change depending on whether the OS 106 has entered the discovery mode 118, as shown below in conjunction with Figure 1B discussed.

[0032] As mentioned herein, an active window is an application window with a graphical user interface (GUI) that is currently displayed in the foreground and has focus, which means that it is a window that receives keyboard, mouse, microphone, and / or other types of input. For example, when a cursor 120 hovers over or clicks on an active window 110, mouse input is provided to the active window 110. The active window is typically indicated by a highlighted and / or visually distinct title bar. Typically, only one window can be active at a time. An inactive window 112 is a window that is not currently in the foreground but is still open and accessible to the user. An inactive window 112 does not have OS focus, which means that it does not receive keyboard or mouse input.

[0033] Applications 116 are programs designed to perform a specific task or set of tasks for a user. Applications are designed to run on various operating systems, such as Windows, macOS, and Linux. Applications can be as simple as a calculator application or as complex as a video game. Some common examples of computer applications include word processors, web browsers, email clients, and media players.

[0034] As shown, active window 110 displays a block of text surrounded by two images. This content is illustrative and non-limiting, and any other type of content is similarly contemplated, including spreadsheets, video games, web pages, maps, computer-aided design drawings, videos, etc. Application 116 can draw the content in various ways using various technologies, such as 2D graphics, 3D graphics, vector graphics, raster graphics, animation, compositing, ray tracing, immediate mode, retained mode, etc.

[0035] The operating system 106 includes a content analysis engine 107. Although the following Figure 1B As described in more detail, the content analysis engine 107 analyzes the graphical output of the active window 116 and / or the underlying data associated with the active window 116 to identify content that can be used by the OS-provided features and / or third-party-provided features. In this way, even if the application 116 is not designed to utilize the OS-provided features and / or third-party-provided features, the OS-provided features and / or third-party-provided features can be made available to the user without the participation of the application 116.

[0036] As referred to herein, an applicable area 115 of a desktop 100 refers to one or more areas that are analyzed by the content analysis engine 107. The applicable area 115 can include portions of a window, such as an application-drawn portion of an active window 110 or a portion visible in an inactive window 112. The applicable area 115 can also include portions of a window that are not visible. Additionally or alternatively, the applicable area 115 can include a single active window, such as the active window 110, two or more windows associated with the same application 116, or a set of windows associated with two or more applications.

[0037] In some configurations, applicability region 115 may include windows that have been active within a defined time period, such as windows that have been active within the past five minutes. In some configurations, applicability region 115 may include windows from recently used applications, such as windows from the three most recently active applications. In some configurations, applicability region 115 may include windows from applications selected based on frequency of use. For example, if a user frequently uses an email application, applicability region 115 may be defined to include windows associated with the email application, as OS-provided and / or third-party-provided features are more likely to provide benefits to the user. In other embodiments, less frequently used applications may be selected for inclusion in applicability region 115 to highlight features that the user may not be aware of. Applicability region 115 may also be selected to include applications based on how frequently OS-provided and / or third-party-provided features are invoked from different applications. Applicability region 115 may also include portions of the desktop. Applicability region 115 may be determined based on a combination of these criteria, in addition to other criteria.

[0038] Figure 1B Activating discovery mode 118 is shown. In some configurations, discovery mode 118 is activated in response to a user pressing an activation key 113 on keyboard 111. Additionally or alternatively, discovery mode 118 can be activated by a mouse click, a mouse press and hold, a touch gesture such as a slide, voice activation, or the like. When pressed and released in a conventional manner (i.e., not held for a long time), activation key 113 can perform other functions, such as displaying a menu for launching applications. In some configurations, when discovery mode 118 is triggered by pressing and holding activation key 113, a mouse press and hold, or a touch press, discovery mode 118 remains active until the user releases the trigger.

[0039] Discovery mode 118 can also be entered by selecting an "enter discovery mode" action from a menu. For example, a user can move a cursor 120 over the activation button 102, right-click the mouse to activate a context menu, and select "enter discovery mode" from the context menu. The context menu can be similarly accessed by right-clicking or performing an equivalent operation on a window or on the desktop itself. When entered via the context menu, discovery mode 118 can continue without the user having to hold down a key. The user can exit discovery mode 118 by selecting an "exit discovery mode" action from the context menu or otherwise disengaging discovery mode 118. In some configurations, once discovery mode 118 has been entered, an activation indication 104 appears near the activation button 102.

[0040] As shown, activation indicator 104 highlights activation button 102. Activation indicator 104 can draw attention to activation button 102 by shading, bolding, increasing prominence, etc. Activation indicator 104 is one indication that discovery mode 118 has been entered. Other indications include a different mouse cursor icon, analysis indicator 101, color tint 114, etc.

[0041] In some configurations, upon entering the discovery mode 118, the content analysis engine 107 analyzes the graphics displayed in the applicable area 115. Figure 1B As shown, applicability region 115 coincides with active window 110 of application 116, but as discussed above, applicability region 115 may include only portions of active window 110, additional windows of application 116, and / or windows from other applications, etc. Active window 110 is the applicability region 115 for many of the examples discussed herein, but those skilled in the art will appreciate that the same techniques apply equally to the other types of applicability regions discussed above.

[0042] As shown, the content analysis engine 107 analyzes the active window 110 to identify content that can be used by features provided by the OS and / or features provided by third parties. The content analysis engine 107 can analyze a snapshot of the graphics displayed in the active window 110, or the content analysis engine 107 can observe the graphics displayed in the active window 110 over time.

[0043] The operating system 106 can provide OS-provided features and / or third-party-provided features. OS-provided features and / or third-party-provided features can also be provided by other providers via a plug-in mechanism. In some configurations, the plug-in mechanism allows for configuration or customization of how the applicable areas 115 are determined. The plug-in mechanism can also allow for identification of which high-level entities, how high-level entities are identified, identification of entities, which entities are highlighted with visual cues, which visual cues are used for which entities, which OS-provided features and / or third-party-provided features are available via visual cues, and / or configuration or customization of the implementation of specific OS-provided features and / or third-party-provided features.

[0044] The content analysis engine 107 can read the graphics displayed by the active window 110 from the display buffer 108. The display buffer 108 is a portion of the computer or GPU memory that is dedicated to storing image data to be displayed on the computer screen or monitor. The graphics displayed by the active window 110 are also referred to herein as content 109 stored in the display buffer 108.

[0045] The content 109 may be stored in the form of pixels and periodically updated by the application 116 to generate graphics displayed by the active window 110. Additionally or alternatively, the content 109 may be stored in a vector graphics format that mathematically defines the content. The content 109 may also be stored in the form of instructions for drawing the content, such as drawing primitives, a document object model (DOM), and the like.

[0046] Reading content 109 directly from display buffer 108 enables identification of portions of content 109 regardless of the technology used to generate the content. For example, an image displayed by active window 110 can be analyzed as an array of pixels regardless of any compression technology used to store the image on disk.

[0047] As another example, the text displayed in the active window 110 can be analyzed regardless of which library was used to generate the text. Optical character recognition (OCR) can be used to identify the original text from the pixel-based content. If the content 109 contains a series of drawing commands, the text can be inferred by analyzing the drawing commands that output the text to the display buffer 108.

[0048] Accessing raw display buffer data in this manner allows for analysis of more content, thereby increasing the likelihood that content that can be used by OS-provided features and / or third-party-provided features will be identified. For example, accessing content 109 from display buffer 108 enables text contained in an image to be identified and made selectable by the user. Without the ability to analyze content 109 from display buffer 108, this text would remain unselectable.

[0049] In some configurations, analysis indicator 101 visually indicates that content analysis engine 107 of OS 106 is analyzing the content of active window 110, which may be used by OS-provided features and / or third-party-provided features. As shown, analysis indicator 101 is a band that surrounds active window 110. In some configurations, analysis indicator 101 may appear as a tint or other overlay over some or all of applicable areas 115. Analysis indicator 101 may alert the user that active window 110 is being analyzed by changing color or shape. For example, analysis indicator 101 may flash, color cycle, or otherwise animate and / or indicate that content analysis engine 107 is in the process of identifying content within active window 110.

[0050] In some configurations, tint 114 is an example of a shading applied to portions of desktop 100 that are not currently being analyzed by content analysis engine 107. As shown, tint 114 obscures inactive windows 112 and portions of desktop 100 that do not contain any windows. In some configurations, tint 114 may also obscure menu 101. Tint 114 emphasizes that active window 110 is being analyzed, rather than inactive windows 112 or other portions of desktop 100.

[0051] The OS 106 may change one or more visualizations upon entering discovery mode 118. For example, the discovery cursor 122 may replace the default OS cursor icon 120. In other configurations, the default OS cursor 120 may change color. In some configurations, the discovery cursor 122 shares colors, animations, and other properties with the analysis indicator 101.

[0052] Analysis indicator 101, discovery cursor 122, and entity visual cues discussed below (e.g., in Figure 1C Together, these layers form a layer on top of active window 110, identifying and providing access to OS-provided features and / or third-party-provided features. In this context, a layer refers to a collection of user interface elements that appear between active window 110 and the user. Consistent with appearing above active window 110, user interface elements in this layer can intercept user interface commands that would otherwise go directly to active window 110.

[0053] Figure 1C 1. A visual cue is shown overlaid near an entity that the content analysis engine 107 has discovered in the active window 110. As referred to herein, an entity refers to a portion of the content displayed by the active window 110 that can be consumed by an OS-provided feature and / or a third-party-provided feature. For example, an OS-provided feature and / or a third-party-provided feature can show a location on a map. This OS-provided feature and / or a third-party-provided feature can use an address identified by the content analysis engine 107 to show the location of the address on a map.

[0054] In some configurations, the content analysis engine 107 first segments the content displayed by the active window 110 into images, text blocks, UI elements (such as buttons, scroll bars, and dialog boxes), and other high-level entities. As shown, the content displayed by the active window 110 has been divided into three high-level entities: image 140, text 150, and image 160. In some configurations, machine learning techniques and / or artificial intelligence techniques are used to identify the high-level entities.

[0055] As shown, the image 140 is highlighted by an image entity visual cue 148. The image entity visual cue 148 is depicted as a border around the image 140, but similarly consider the other means for highlighting the image 140 and other entities discussed herein. For example, the image entity visual cue 148 can cause a shadow to appear below the image 140. In addition, the image entity visual cue 148 can be animated. The image entity visual cue 148 can also highlight the image 140 by moving the image 140 around on the screen, such as by moving it back and forth. The image entity visual cue 148 can also highlight the image 140 by dimming the surrounding content, i.e., displaying it at a reduced brightness. Although the following is combined with Figures 6A to 6C As discussed in greater detail, the image entity visual cue 148 may appear or disappear, change color, or animate in response to increasing or decreasing distance from the discovery cursor 122 .

[0056] In some configurations, entities are nested within each other, forming a hierarchy of entities. For example, if the image 140 contains text, the content analysis engine 107 may identify text entities 149 within the image 140, such as the following: Figure 1D Depicted.

[0057] In some configurations, upon entering the discovery mode 118, visual cues for all entities are displayed. In other configurations, upon entering the discovery mode 118, visual cues for a portion of the entities are displayed. For example, upon entering the discovery mode 118, visual cues for top-level entities may be displayed, while visual cues for nested entities are not displayed until a parent entity, or an action associated with a parent entity, is selected. Delaying or otherwise staggering the display of visual cues for entities in this manner can, among other benefits, reduce screen clutter, reduce the number of choices a user faces at any given time, reduce and / or postpone consumption of computing resources, and improve human-computer interaction.

[0058] In some configurations, in addition to delaying the display of visual cues for nested entities, the content analysis engine 107 can also delay analyzing the content of a parent entity until it is selected. Child entities can then be identified during the analysis of the parent entity. Delaying analysis in this manner reduces the time and computational resources required to render the initial set of entity visual cues upon entering discovery mode. Delaying computation can also improve efficiency, including by completely skipping processing that would otherwise occur.

[0059] The text 150 has been processed by the content analysis engine 107 to identify multiple entities. A paragraph entity 152 is one of the multiple entities that represent a single paragraph in the text block. A date entity, such as "Tuesday, May 3, noon," can be identified by analyzing the text of the text 150 itself. In some configurations, machine learning techniques and / or artificial intelligence techniques are applied to distinguish text that can be used by features provided by the OS and / or features provided by third parties, such as dates, times, locations, etc. Other examples of text-based entities shown are the location entity "Willard Park" and the address entity 157 "1234 Main St., Lincoln, WI."

[0060] Each of the entities mentioned above is overlaid with a corresponding entity visual cue that can be used to identify and / or invoke OS-provided features and / or third-party-provided features related to the associated entity. For example, a date entity visual cue 154 appears as an underline beneath the text of the date entity "Tuesday, May 3, noon." Similarly, a location entity visual cue 156 appears as a line beneath the "Willard Park" location entity, and an address entity visual cue 158 appears as a line beneath the "1234 Main Street, Lincoln, Wisconsin" entity 157.

[0061] Images 140 and 160 may also be processed by content analysis engine 107 using one or more image recognition and / or processing techniques. For example, within image 160, content analysis engine 107 has identified the subject of the photo or image. The subject of the photo or image may be a person, an animal, or other prominently displayed object. The content that constitutes the subject may be customized using plugins that provide OS features and / or third-party features.

[0062] In some configurations, objects are identified based on the outline and content of the objects themselves. For example, edge detection and facial recognition can be used to identify people within a photo as subjects. Subjects can also be identified based on the location of the objects in the photo. For example, a person in the center of a photo can be identified as a subject, while people in the background may not be identified as subjects. The subject of a photo can also be determined based on whether the object is in focus, and / or one or more user-specified settings. As shown, the content analysis engine 107 identifies the subject of a photo or image by overlaying an extracted subject entity visual cue 162 on or near the subject (e.g., three people in the foreground of a photo or image).

[0063] Figure 1DThe highlighted image entity visual cue 148 is shown in response to the hovering of the discovery cursor 122. In this example, the image entity visual cue 148 has been shaded. Alternatively or additionally, the image entity visual cue 148 can be highlighted by changing the color or thickness of the border that appears around the image 140.

[0064] As briefly discussed above, the content analysis engine 107 can analyze the nested entities of the image 140 in response to discovering that the cursor 122 is hovering over the image entity visual cue 148. Nesting entities in this manner provides a number of benefits over identifying and / or displaying all entity visual cues at once. First, identifying a top-level entity such as the image 140 without having to further identify child entities reduces latency and lowers computational costs. Additionally, displaying a large number of entity visual cues (e.g., from all levels in an entity hierarchy), such as text contained in an image, can clutter the screen, thereby reducing the user experience.

[0065] As shown, the content analysis engine 107 has identified the text "Park 1.5 km" entity within the sign identified in the image 140. Therefore, the content analysis engine 107 has overlaid a visual cue 142 (hereinafter "visual cue 142") of the text entity in the image over the identified text. The visual cue 142 appears as a line below the text "Park 1.5 km."

[0066] Figure 1E 1. The text nested within the image 140 is shown highlighted. The discovery cursor 122 has been moved over the visual cue 142. In response, the discovery cursor 122 may change to an "I" cursor 124, indicating that the user may select and / or otherwise interact with the highlighted text. Moving the discovery cursor 122 over the visual cue 142 also causes the visual cue 142 to highlight the underlying text 146: "Park 1.5 km."

[0067] Figure 1F 146 . An entity action list 180 is shown displayed in response to receiving a selection of text highlighted by a text-in-image visual cue 142. The visual cue 142 can be selected by clicking a mouse or touchpad button while the "I" cursor 124 is over the visual cue 142. Other input methods are similarly contemplated, such as a touch press or swipe, voice command, gaze, and the like. In response to being selected, the shading style associated with the text-in-image entity visual cue 142 can change from highlighted text 146 to selected text 148. This change in shading style communicates to the user that the selection of the text 148 was successful and that the entity action list 180 is associated with the selected text 148.

[0068] As mentioned herein, the entity action list displays a list of actions that can be invoked for the underlying entity. In some configurations, a visual cue indicates which visual cue the entity action list is associated with, such as changing the shading of the corresponding entity. The entity action list 180 contains two entity actions selected by the content analysis engine 107: search network 182 and convert to mileage 184. Certain entity actions may be included in the entity action list 180 because they are generally applicable to any displayed content and / or specific types of content. For example, the content analysis engine 107 may select the search network 182 entity action because the action is available as long as text is selected. Other entity actions may be included in the entity action list 180 because they are related to the specific content selected, and / or attributes of the user and / or computing device. For example, the content analysis engine 107 may select the convert to mileage 184 action based on a measure of distance already identified in the text ("1.5 km") and / or the geographic location of the computing device.

[0069] The content analysis engine 107 may use artificial intelligence techniques, such as machine learning and natural language processing, to identify distances, such as "1.5 km". Additionally or alternatively, the content analysis engine 107 may use regular expressions, lexical analyzers, parsers, finite state machines, or other known techniques to identify text within text entities that may be used by features provided by the OS and / or features provided by third parties. Similar techniques may be used to identify dates, times, email addresses, URLs, etc. The same techniques may also be used to identify named entities, such as locations, people, businesses, countries, etc.

[0070] The content analysis engine 107 can provide a default set of entities to be identified, visual cues to be displayed, and entity actions to be included in the entity action list 180. The content analysis engine 107 can also call plug-ins to perform some or all of these operations. For example, a plug-in can define entity actions to be displayed in the entity action list 180. The content analysis engine 108 can query multiple plug-ins for entity actions to add to the entity action list 180.

[0071] Figure 1G The deactivation of discovery mode 118 is shown. The manner of exiting discovery mode 118 may depend on how discovery mode 118 was entered. For example, if discovery mode was entered by pressing activation key 113, releasing activation key 113 may result in deactivation of discovery mode. Discovery mode 118 may also be exited by pressing a designated key (such as an "Exit" key), by selecting "Exit Discovery Mode" from a top or bottom menu, or in any other suitable manner. Deactivation indicator 105 visually indicates that discovery mode 118 has been exited. For example, deactivation indicator 105 may restore activation button 102 to an unactivated state.

[0072] In some configurations, in response to deactivating the discovery mode 118, all entities, highlights, analysis indicators, and other UI elements generated by the discovery mode 118 are removed. In other configurations, the most recently selected or otherwise interacted entity may remain in place. Figure 1G This is illustrated by visual cue 142 remaining visible even though the other entity visual cues have been removed and “I” cursor 124 has returned to the default OS icon of cursor 120 .

[0073] Figure 2A A paragraph entity 152 within the highlighted text 150 is shown. The content analysis engine 107 can use machine learning-based techniques or traditional image segmentation techniques to distinguish between text, images, UI controls, and other sub-parts of application-generated images obtained from the display buffer 108. The content analysis engine 107 can also use a screen description framework (such as an accessibility framework) to distinguish and identify content, where the application provides metadata about the content 109 of the buffer 108. When the text entity is image-based, the content analysis engine 107 can use optical character recognition to extract the text. When the text entity is obtained from the screen description framework, the text can be obtained directly from the frame. In either case, the OS 106 is able to identify and utilize content generated by the application 116, which would otherwise be opaque and inaccessible.

[0074] The cursor 122 is found to have moved over the paragraph entity 152, becoming an "I" cursor 224. Moving the cursor 122 over the paragraph entity 152 also causes a hovering paragraph highlight 252 to be displayed over the paragraph entity 152. The hovering paragraph highlight 252 indicates to the user that the content generated by the active window 110 has been identified as text, and that the identified text can be manipulated by one or more OS-provided features and / or third-party-provided features.

[0075] Paragraph entity 152 is an example of an entity that is not highlighted by a visual cue until the discovery cursor 122 has moved over text 150 and / or paragraph entity 152. However, in other embodiments, a visual cue may be displayed near paragraph entity 152 when discovery mode 118 is entered. For example, a vertical bar visual cue may be displayed in the margin next to paragraph entity 152.

[0076] Changing the discovery cursor 122 to an "I" cursor 224 indicates that the user can insert a caret into the paragraph entity 152, select text within the paragraph entity 152, and perform other operations that typically require knowing what text has been drawn to the window 110. The caret refers to a vertical bar inserted between text characters that indicates where the addition, editing, or selection of text will occur. These text manipulation operations are enabled by the content analysis engine 107 to identify what text is drawn where in the active window 110. The shape and style of the discovery cursor 122 and the "I" cursors 124 and 224 are not limited, and other cursor icons and styles are also contemplated.

[0077] Figure 2B A paragraph entity action list 280 is shown displayed in response to selecting a highlighted paragraph entity 152. For example, when the "I" cursor 224 is over the paragraph entity 152, the user can select the highlighted paragraph entity 152 by clicking the right mouse button. In response to the selection, the highlight of the paragraph entity 152 changes to the selected paragraph 254, giving a visual indication of the entity with which the paragraph entity action list 280 is associated.

[0078] The paragraph entity action list 280 includes two entity actions: summarize 282 and translate 284. The content analysis engine 107 can select these entity actions based on an analysis of the text of the paragraph entity 152. These entity actions can also be based on the position of the "I" cursor 224 within the paragraph entity 152 when the selection is made.

[0079] In some configurations, when two or more entities are associated with the same text, only entity actions from one of the entities are displayed in the entity action list. For example, a priority order can determine which entity is selected and, therefore, which entity actions are available in the entity action list. In other configurations, when two or more entities are associated with the same text, entity actions from each entity are aggregated into a single entity action list. Similar techniques can be used to select entity actions for other types of entities, such as image-based entities. Figure 2A and Figure 2B In the example shown, the selected paragraph 254 can be determined to include two entities, as described above: (1) the paragraph entity 152, and (2) the date entity 154. For example, different entity actions can be exposed based on which entity the user is predicted to be interested in. For example, based on the position of the cursor 224 (e.g., over the paragraph entity 152, specifically Kimball's name, but not over the date entity 154) and / or other factors, a text-based paragraph entity action 280 such as Figure 2BIf the cursor 224 is alternatively positioned over the date entity 154, or it is otherwise predicted that the user is more interested in this entity, other paragraph entity actions may be revealed, such as a create calendar event action.

[0080] Figure 2C 28. A selection of an entity action 284 is shown received from the paragraph entity action list 280. As shown, the discovery cursor 222 selects the "Translate" entity action 284. In some examples, as shown, the discovery cursor 222 replaces the "I" cursor 224 when the cursor moves from the paragraph entity 152 to the paragraph entity action list 280.

[0081] Figure 2D 10. The selected paragraph 152 is shown after the selected action 284 has modified the paragraph. In the example shown, the text of the paragraph entity 152 has been translated into a different language in accordance with the selection of the "Translate" entity action. The translated text 262 has replaced the original text of the selected paragraph 152 in the active window 110. In some configurations, the translated text replaces the original text by writing the translated text to the display buffer 108. In other configurations, the content analysis engine 107 places the translated text into a surface that obscures the paragraph entity 152, effectively replacing it from the user's perspective.

[0082] In some configurations, the translated text 262 is highlighted as an action-completed paragraph 256, giving a visual indication that the selected entity action has been performed. In some configurations, once the entity action has been completed, the resulting effect is semi-permanent and will outlast the discovery mode 118. In other embodiments, when the discovery mode 118 is exited, the effect, such as the translated text, will be restored.

[0083] Figure 3A The subject entity is shown highlighted in response to finding that the cursor is hovering over the subject entity. Specifically, the content analysis engine 107 has identified the subject entity within the image 160. The subject entity is the portion of the photo or image that clearly appears in the foreground and is typically a person, animal, item for sale, or other important feature of the photo or image. Machine learning techniques and / or artificial intelligence techniques can be used to identify objects within the photo or image. The machine learning techniques and / or artificial intelligence techniques can further be used to determine the outline of the identified subject.

[0084] Once the subject has been identified, the content analysis engine 107 may display an extracted subject entity visual cue 162 within or adjacent to the subject entity, as described above. Figure 1C As shown. Figure 1CAs shown, this visual cue may be relatively small compared to the main entity itself. The visual cue can be a geometric shape, image, etc. that is located within the boundary of the main entity or near the main entity, for example, partially overlapping the main entity. Other visual cues can follow the outline of the main entity.

[0085] In some configurations, the extract subject entity 362 is highlighted in response to the discovery cursor 322 moving or hovering over the extract subject entity visual cue 162. In other configurations, and in contrast to the text entity visual cue 142 in the image being highlighted when the discovery cursor 322 hovers over the visual cue 142 itself, the extract subject entity 362 can be highlighted in response to the discovery cursor 322 moving over any portion of the subject entity. In either case, the extract subject entity visual cue 162 can be highlighted to indicate that it can be selected to execute an OS-provided feature and / or a third-party-provided feature.

[0086] Figure 3B A drag operation is shown to extract a subject entity from an image. In this example, an OS-provided and / or third-party-provided feature, Extract Subject Entity, is activated without a context menu. In some configurations, a copy 364 of the subject entity can be dragged to another application, where releasing the drag can initiate the copy operation.

[0087] Other OS provided features and / or third party provided features can be initiated in a similar manner. For example, the copy and cut keyboard shortcuts can be used to copy the subject entity 162 to the OS clipboard or to copy it to the clipboard while removing it from the image 160, respectively. When the highlighted entity contains text, you can use Figure 2A An "I" cursor 224 is depicted to select some or all of the text, and keyboard shortcuts may similarly be used to extract, remove, or replace the selected text.

[0088] Figure 4A The image 160 is shown highlighted in response to finding the cursor 422 hovering over the image background. Figure 3A As discussed, the content analysis engine 107 can distinguish the background of a photo or image from the subject of the photo or image.

[0089] The image 160 highlights the movement of the discovery cursor 422 over the image 160 by overlaying a shadow over the image 160. In some configurations, this shadow can be removed if the discovery cursor 422 moves over the extract subject entity visual cue 162 or the subject 362, in which case the subject 362 can be highlighted instead. The shadow can also be removed if the discovery cursor 422 moves away from the image 160.

[0090] Figure 4B 160 is displayed in response to receiving a selection of the image 160, the entity action list 480 (in addition to the subject entity 362). Figure 1F and Figure 2B The approach to discussing entity action lists selects entity action list 480. As shown, image entity action list 480 has a single entity action - remove background entity action 462.

[0091] Figure 4C 4. The image 160 is shown after the background has been removed by the remove background action 462. As shown, the discovery mode 118 has been exited as a result of completing the remove background action 462. However, in other configurations, the discovery mode 118 may continue as long as the activation key 113 remains pressed and the user does not otherwise exit the discovery mode 118.

[0092] One result of applying the remove background action 462 is that the cursor 422 is found to have been replaced by the OS or application default cursor 120. Additionally, the background of the image 160 has been removed. Figure 2C Introduced and Figure 2D In the translation entity action 284 applied in

[0064] , the background of the image 160 may be removed by writing directly to the display buffer 108 or by overwriting a copy of the image 160 with the background removed. Other techniques for applying these changes are similarly contemplated.

[0093] Figure 5A 58 is shown as an entity action list 580 displayed in response to selecting the address entity visual cue 558. The content analysis engine 107 has identified the text "1234 Main Street, Lincoln, Wisconsin" as an address and has generated the address entity visual cue 558 accordingly. As shown, the text has been generated in a manner similar to Figure 1D The address entity visual cue 558 is selected in the same manner as the text entity visual cue 142 in the image. Specifically, when the cursor moves over the address entity visual cue 558, an "I" cursor 524 appears. Once selected, such as by clicking with a mouse button, the highlight of the address entity visual cue 558 changes, indicating that a selection has been made. The entity action list 580 is also displayed in response to the selection of the address entity visual cue 558. The entity action list 580 includes a single entity action: direction 582.

[0094] Figure 5BAn inline micro-experience 502 is shown displayed in response to selecting an address entity action 582. The micro-experience exposes an application, widget, or other component that can provide functionality that is overlaid, adjacent to, or otherwise associated with the active window. The micro-experience enables features to be provided as part of the current user experience while maintaining focus on the current application.

[0095] The map micro-experience 502 can be a pop-up window, dialog box, or other UI control that presents OS functionality and / or third-party functionality based on the content of the address entity visual cue 558. In this example, the map micro-experience 502 presents a street map of the address associated with the address entity visual cue 558. The map micro-experience 502 is an example of a micro-experience. An email address entity can support a micro-experience that displays a list of recent communications with the underlying email address. An online encyclopedia micro-experience can display the definition of a selected word. In some configurations, the micro-experience can interact with something similar to a web page. At the same time, the micro-experience is not limited to displaying HTML and can actually leverage features of the OS 106 or other applications.

[0096] Figures 6A to 6C Specific visual cues that are near the discovery cursor are highlighted. In some configurations, specific entity visual cues can be selectively displayed based on proximity to the discovery cursor 622. Similar to the reasons for displaying nested entity visual cues, this technique can be used to reduce the number of entity visual cues overlaid on top of the active window 110. Specifically, reducing the number of visual cues reduces screen clutter and processing time.

[0097] like Figure 6A As shown, the discovery cursor 622 is closest to the entity visual cue 660 associated with the image 160 and the date entity visual cue 154, and therefore these entity visual cues are highlighted for display. The discovery cursor 622 is also close to the paragraph entity 152, but in this configuration, the paragraph entity 152 is not highlighted with a visual cue until the discovery cursor 622 moves over it.

[0098] In some configurations, the entity visual cue associated with a particular entity is not fully displayed. For example, the entity visual cue 660 may wrap around the image 160, similar to Figure 1C 140. However, to emphasize that the entity visual cue 660 is selected for display in response to the location of the discovery cursor 622, only the portion closest to the discovery cursor 622 is displayed. As the location of the discovery cursor 622 changes, the portion selected for display may be updated.

[0099] How close an entity visual cue must be to the discovery cursor 622 in order to be highlighted is configurable, such as by an administrator or end user. In addition, one or more filters may be applied to determine which entity visual cue is highlighted and how far from the discovery cursor 622 it is. These filters may be built into the content analysis engine 107 or provided by a plugin. Such filters may also be configurable.

[0100] Figure 6B It is shown that the discovery cursor 622 has moved, and the highlighted portion of the entity visual cue 660 has also moved. At the same time, the date entity visual cue 154 is no longer highlighted, and the address entity visual cue 158 is highlighted.

[0101] Similarly, Figure 6C The discovery cursor is shown having moved again, this time proximate the image entity visual cue 148. At the same time, the highlighting of the address entity visual cue 158 remains unchanged as it remains proximate the discovery cursor 622.

[0102] Figure 7 An ambient mode is shown in which discovery mode 118 is automatically triggered. Specifically, an ambient address entity visual cue 758 is overlaid on top of the active window 110 without the user having to explicitly activate discovery mode 118. Triggering discovery mode 118 automatically enables the availability of OS-provided features and / or third-party provided features to be revealed without the user having to manually enter discovery mode 118. In some configurations, discovery mode 118 is automatically triggered when the applicable area 115 changes, when a window is scrolled, when a new application is launched, or when previously unanalyzed content comes into view. While in an automatically triggered discovery mode, the user can still manually enter discovery mode 118, at which time additional visual cues may be revealed.

[0103] To reduce the chance of incorrect, invalid, or unwanted entity visual cues, the type, location, and / or number of entity visual cues may be limited when discovery mode 118 is automatically triggered. For example, address entities may be allowed, but subject entities may not be allowed. As another example, when discovery mode 118 has been automatically entered, entity visual cues may only be allowed within a specified distance of the discovery cursor 622.

[0104] In some configurations, when the discovery mode 118 is automatically triggered, the type, location, and number of entity visualization cues allowed can be customized based on personal usage history. For example, assuming a user has previously activated Figure 5A Based on this usage history, the same user may be more likely to be presented with an address entity visualization prompt when the discovery mode 118 is automatically triggered. This history is reflected in Figure 7 In, with Figure 1C Compared to the more visual cues shown in Figure 7 Only the ambient address entity visual cue 758 is shown.

[0105] refer to Figure 8 , the routine 800 begins at operation 802, where an activation 104 of the discovery mode 118 is received. The activation 104 can be in the form of pressing and holding an activation key 113 associated with an activation button 102 displayed by the OS 106.

[0106] Next, at operation 804, the display content extracted from the display buffer 108 of the active window 110 is segmented into high-level entities. For example, blocks of text, images, UI controls, and other identifiable types of content may be identified.

[0107] Next, at operation 806 , an entity is identified within one of the higher-level entities. For example, an email address entity may be identified from the text within the text 150 .

[0108] Next at operation 808, an entity visual cue is overlaid on top of the entity identified by operation 806. For example, the address entity visual cue 558 is displayed as an underline adjacent to the text of the associated entity.

[0109] Next at operation 810, an indication of a selection of the address entity visual cue 558 is received. This indication may be in the form of a mouse click.

[0110] Next, at operation 812, the entity associated with the selected entity visual cue is used to activate an OS-provided feature and / or a third-party-provided feature. Continuing with this example, the address of the selected address entity visual cue 558 can be used to initiate display of an OS-provided feature and / or a third-party-provided feature of the map microexperience 502.

[0111] The specific implementation of the technology disclosed herein is a problem of selection depending on the performance and other requirements of the computing device. Therefore, the logical operations described herein are variously referred to as states, operations, structural devices, actions, or modules. These states, operations, structural devices, actions, and modules can be implemented in hardware, software, firmware, with dedicated digital logic and any combination thereof. It should be understood that more or less operations can be performed than shown in the accompanying drawings and described herein. These operations can also be performed in an order different from those described herein.

[0112] It should also be understood that the methods shown can be terminated at any time and do not need to be performed in full. Some or all of the operations of these methods, and / or substantially equivalent operations, can be performed by executing computer-readable instructions, which are included on a computer storage medium, as defined below. As used in the specification and claims, the term "computer-readable instructions" and its variants are used broadly herein to include routines, applications, application modules, program modules, programs, components, data structures, algorithms, and the like. Computer-readable instructions can be implemented on various system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like.

[0113] Thus, it should be understood that the logical operations described herein are implemented as (1) a sequence of computer-implemented actions or program modules running on a computing system and / or (2) interconnected machine logic circuits or circuit modules within the computing system. Implementation is a matter of choice depending on the performance and other requirements of the computing system. Thus, the logical operations described herein are variously referred to as states, operations, structural devices, actions, or modules. These operations, structural devices, actions, and modules may be implemented in software, in firmware, in dedicated digital logic, or in any combination thereof.

[0114] For example, the operations of routine 800 are described herein as being implemented at least in part by modules that execute features disclosed herein, which may be dynamic link libraries (DLLs), static link libraries, functions generated by application programming interfaces (APIs), compilers, interpreters, scripts, or any other executable instruction sets. Data may be stored in data structures in one or more memory components. Data may be retrieved from a data structure by addressing a link or reference to the data structure.

[0115] Although the following diagrams relate to components of the accompanying drawings, it should be understood that the operations of routine 800 can also be implemented in many other ways. For example, routine 800 can be implemented at least in part by a processor or local circuit of another remote computer. In addition, one or more of the operations of routine 800 can alternatively or additionally be implemented at least in part by a chipset working alone or in combination with other software modules. In the examples described below, one or more modules of the computing system can receive and / or process the data disclosed herein. Any service, circuit, or application suitable for providing the technology disclosed herein can be used in the operations described herein.

[0116] Figure 9Additional details are shown of an example computer architecture 900 for a device, such as a computer or server configured as part of the systems described herein, capable of executing computer instructions (eg, modules or program components described herein). Figure 9 The illustrated computer architecture 900 includes a processing unit 902 , a system memory 904 including random access memory 906 (“RAM”) and read only memory (“ROM”) 908 , and a system bus 910 that couples the memory 904 to the processing unit 902 .

[0117] Processing unit(s), such as processing unit(s) 902, may represent, for example, a CPU-type processing unit, a GPU-type processing unit, a neural processing unit, a field programmable gate array (FPGA), another type of digital signal processor (DSP), or other hardware logic components that may be driven by a CPU in some cases. For example, but not limitation, exemplary types of hardware logic components that may be used include application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0118] A basic input / output system, containing the basic routines that help to transfer information between elements within the computer architecture 900, such as during start-up, is stored in ROM 908. The computer architecture 900 further includes a mass storage device 912 for storing an operating system 914, application(s) 916, modules 918, and other data described herein.

[0119] The mass storage device 912 is connected to the processing unit(s) 902 through a mass storage controller connected to the bus 910. The mass storage device 912 and its associated computer-readable media provide non-volatile storage for the computer architecture 900. Although the description of computer-readable media contained herein refers to mass storage devices, those skilled in the art will understand that computer-readable media can be any available computer-readable storage media or communication media that can be accessed by the computer architecture 900.

[0120] Computer-readable media may include computer-readable storage media and / or communication media. Computer-readable storage media may include one or more of volatile memory, non-volatile memory, and / or other persistent and / or secondary computer storage media, removable computer storage media, and non-removable computer storage media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Thus, computer storage media includes media in tangible and / or physical form included in a device and / or as part of a device or in a hardware component external to the device, including but not limited to random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), phase change memory (PCM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, compact disk read-only memory (CD-ROM), digital versatile disks (DVD), optical cards or other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage devices, magnetic cards or other magnetic storage devices or media, solid-state storage devices, storage arrays, network attached storage devices, storage area networks, hosted computer storage devices, or any other storage memory, storage devices, and / or storage media that can be used to store and maintain information for access by a computing device.

[0121] In contrast to computer-readable storage media, communication media can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal (such as a carrier wave) or other transmission mechanism. As defined herein, computer storage media does not include communication media. That is, computer-readable storage media itself does not include communication media consisting solely of a modulated data signal, carrier wave, or propagated signal.

[0122] According to various configurations, the computer architecture 900 can operate in a networked environment using logical connections to remote computers via a network 920. The computer architecture 900 can be connected to the network 920 via a network interface unit 922 connected to the bus 910. The computer architecture 900 can also include an input / output controller 924 for receiving and processing input from a number of other devices, including a keyboard, a mouse, a touch screen, or an electronic stylus or pen. Similarly, the input / output controller 924 can provide output to a display screen, a printer, or other types of output devices.

[0123] It should be understood that when loaded into (multiple) processing unit 902 and executed, the software components described herein can convert (multiple) processing unit 902 and the entire computer architecture 900 from a general-purpose computing system into a special-purpose computing system customized to facilitate the functionality presented herein. (Multiple) processing unit 902 can be composed of any number of transistors or other discrete circuit elements, which can individually or collectively assume any number of states. More specifically, (multiple) processing unit 902 can operate as a finite state machine in response to the executable instructions contained in the software modules disclosed herein. These computer-executable instructions can convert (multiple) processing unit 902 by specifying how (multiple) processing unit 902 should transition between various states, thereby converting the transistors or other discrete hardware elements that constitute (multiple) processing unit 902.

[0124] Figure 10 An exemplary distributed computing environment 1000 capable of executing the software components described herein is depicted. Figure 10 The distributed computing environment 1000 shown can be used to execute any aspects of the software components presented herein. For example, the distributed computing environment 1000 can be used to execute various aspects of the software components described herein.

[0125] Thus, the distributed computing environment 1000 can include a computing environment 1002 that operates on, communicates with, or is part of a network 1004. The network 1004 can include various access networks. One or more client devices 1006A through 1006N (hereinafter collectively and / or generally referred to as "clients 1006," and also referred to herein as computing devices 1006) can communicate with the computing environment 1002 via the network 1004. In one illustrated configuration, the clients 1006 include a computing device 1006A, such as a laptop, desktop computer, or other computing device; a slate or tablet computing device ("tablet computing device") 1006B; a mobile computing device 1006C, such as a mobile phone, smartphone, or other mobile computing device; a server computer 1006D; and / or other devices 1006N. It should be understood that any number of clients 1006 can communicate with the computing environment 1002.

[0126] In various examples, computing environment 1002 includes a server 1008, a data storage device 1010, and one or more network interfaces 1012. Server 1008 can host various services, virtual machines, portals, and / or other resources. In the configuration shown, server 1008 hosts virtual machines 1014, a web portal 1016, a mailbox service 1018, a storage service 1020, and / or a social networking service 1022. Figure 10 As shown, server 1008 may also host other services, applications, portals, and / or other resources (“other resources”) 1024 .

[0127] As mentioned above, the computing environment 1002 may include a data store 1010. According to various implementations, the functionality of the data store 1010 is provided by one or more databases operating on or in communication with the network 1004. The functionality of the data store 1010 may also be provided by one or more servers configured to host data for the computing environment 1002. The data store 1010 may include, host, or provide one or more real or virtual data repositories 1026A to 1026N (hereinafter collectively and / or generally referred to as "data repositories 1026"). The data repositories 1026 are configured to host data used or created by the servers 1008 and / or other data. That is, the data repositories 1026 may also host or store web documents, word documents, presentation documents, data structures, algorithms for execution by the recommendation engine, and / or other data used by any application. Aspects of the data repositories 1026 may be associated with a service for storing files.

[0128] The computing environment 1002 can communicate with or be accessed by a network interface 1012. The network interface 1012 can include various types of network hardware and software to support communication between two or more computing devices, including but not limited to computing devices and servers. It should be understood that the network interface 1012 can also be used to connect to other types of networks and / or computer systems.

[0129] It should be understood that the distributed computing environment 1000 described herein can provide any number of virtual computing resources and / or other distributed computing functions for any aspect of the software elements described herein, and these other distributed computing functions can be configured to perform any aspect of the software components disclosed herein. According to various implementations of the concepts and technologies disclosed herein, the distributed computing environment 1000 provides the software functions described herein as services to a computing device. It should be understood that a computing device can include a real machine or a virtual machine, including but not limited to a server computer, a network server, a personal computer, a mobile computing device, a smart phone, and / or other devices. Therefore, among other aspects, various configurations of the concepts and technologies disclosed herein enable any device configured to access the distributed computing environment 1000 to utilize the functions described herein to provide the technology disclosed herein.

[0130] This disclosure is supplemented by the following example clauses:

[0131] Example 1: A method comprising: receiving an indication that a discovery mode has been activated; retrieving content displayed by an active window; identifying an image within the retrieved content; identifying a text entity within the identified image; determining whether the text entity is applicable to an OS-provided feature or a third-party-provided feature; overlaying an entity visual cue onto the active window proximate to the text entity; receiving an indication of a selection of the entity visual cue; and invoking the OS-provided feature or the third-party-provided feature.

[0132] Example 2: The method of example 1, wherein invoking an OS feature or a third-party feature comprises providing a text entity to the OS-provided feature or the third-party-provided feature.

[0133] Example 3: The method of example 1, wherein the OS feature or the third-party feature modifies content displayed by the active window, launches an inline micro-experience, crops an image, or exports an image.

[0134] Example 4: The method of example 1, wherein the discovery mode comprises an ambient mode that analyzes content displayed by the active window when previously unanalyzed content comes into view.

[0135] Example 5: The method of example 4, wherein when in ambient mode, the number, location, or type of visual cues are limited.

[0136] Example 6: The method of example 5, wherein the number, location, and type of visual cues displayed when in ambient mode are customized based on personal usage history.

[0137] Example 7: The method of example 1, wherein the content displayed by the active window is retrieved from the display buffer after the active window has drawn the content to the display buffer.

[0138] Example 8: The method of example 1, wherein the content displayed by the active window is retrieved from a visible portion of the active window and a non-visible portion of the active window.

[0139] Example 9: A system comprising: a processing unit; and a computer-readable storage medium having computer-executable instructions stored on the computer-readable storage medium, which, when executed by the processing unit, cause the processing unit to: receive an indication that a discovery mode has been activated; retrieve content displayed by an active window; identify an image within the retrieved content; identify an image subject entity within the identified image; determine whether the image subject entity is applicable to an OS-provided feature or a third-party-provided feature; overlay an entity visual cue onto the active window proximate to the image subject entity; receive an indication of a selection of the entity visual cue; and invoke an OS-provided feature or a third-party-provided feature.

[0140] Example 10: The system of example 9, wherein the entity visualization cue comprises a shape displayed within or near the image subject entity.

[0141] Example 11: The system of example 10, wherein the computer-executable instructions further cause the processing unit to: in response to an indication that a system cursor moves over or near the image subject entity, highlight the entity visual cue.

[0142] Example 12: The system of example 10, wherein the computer-executable instructions further cause the processing unit to: highlight the image subject entity in response to an indication that a system cursor moves over or near the entity visual cue.

[0143] Example 13: The system of example 9, wherein the OS feature or the third-party feature extracts the image subject entity from the identified image.

[0144] Example 14: The system of example 9, wherein the computer-executable instructions further cause the processing unit to: identify a background portion of the image; and in response to determining that a system cursor moves over the background portion of the image, highlight the image.

[0145] Example 15: The system of Example 9, wherein the computer-executable instructions further cause the processing unit to: display an entity action list in response to a selection of the entity visualization prompt; and receive a selection from the entity action list to select an OS-provided feature or a third-party-provided feature.

[0146] Example 16: A computer-readable storage medium having computer-readable instructions encoded thereon that, when executed by a processing unit, cause a system to: receive an indication that a discovery mode has been activated; retrieve content displayed by an active window; identify a paragraph entity within the retrieved content; determine whether the paragraph entity is applicable to an OS-provided feature or a third-party-provided feature; in response to discovering that a cursor moves over or near the paragraph entity, overlay an entity visual cue onto the active window proximate to the paragraph entity; receive an indication of a selection of the entity visual cue; and invoke the OS-provided feature or the third-party-provided feature.

[0147] Example 17: The computer-readable storage medium of example 16, wherein the instructions further cause the processing unit to: extract text depicted in the paragraph entity from the retrieved content; and provide the extracted text to an OS-provided feature or a third-party-provided feature.

[0148] Example 18: The computer-readable storage medium of Example 16, wherein the system cursor changes to a discovery cursor in response to entering discovery mode, and wherein the discovery cursor changes to an "I" cursor in response to moving over or near a paragraph entity.

[0149] Example 19: Computer-readable storage medium according to Example 18, wherein the instructions further cause the processing unit to: display a caret near an "I" cursor within a paragraph entity; receive an indication of an editing command; apply the editing command at the location of the caret; and modify the content displayed by the active window to reflect the results of the editing command.

[0150] Example 20: The computer-readable storage medium of Example 18, wherein the instructions further cause the processing unit to: in response to a text selection command, select text from the paragraph entity at the "I" cursor position.

[0151] Although certain example embodiments have been described, these embodiments are presented as examples only and are not intended to limit the scope of the inventions disclosed herein. Therefore, nothing in the foregoing description is intended to suggest that any particular feature, characteristic, step, module, or block is necessary or indispensable. In fact, the novel methods and systems described herein may be embodied in a variety of other forms; furthermore, various omissions, substitutions, and changes in the form of the methods and systems described herein may be made without departing from the spirit of the inventions disclosed herein. The appended claims and their equivalents are intended to cover such forms or modifications as fall within the scope and spirit of certain inventions disclosed herein.

[0152] It should be understood that any reference to "first," "second," etc., elements in the Summary of the Invention and / or the Detailed Description are not intended to, and should not be interpreted as, necessarily corresponding to any reference to "first" and "second," etc., elements in the claims. On the contrary, any use of "first" and "second" in the Summary of the Invention, the Detailed Description, and / or the claims may be used to distinguish two different instances of the same element.

[0153] Finally, although various techniques have been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended descriptions is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed subject matter.

Claims

1. A method comprising: receiving an indication that discovery mode has been activated; Retrieve the content displayed by the active window; identifying an image within the retrieved content; identifying text entities within the identified image; Determining whether the text entity is applicable to features provided by the OS or features provided by a third party; Overlaying an entity visual prompt onto the active window close to the text entity; receiving an indication of a selection of the entity visual cue; and Call features provided by the OS or features provided by a third party.

2. The method according to claim 1, wherein calling the OS feature or the third-party feature comprises: The text entity is provided to the OS-provided feature or a third-party-provided feature. 3 . The method of claim 1 , wherein the OS feature or the third-party feature modifies the content displayed by the active window, launches an inline micro-experience, crops an image, or exports an image. 4 . The method of claim 1 , wherein the discovery mode comprises an ambient mode that analyzes content displayed by the active window when previously unanalyzed content comes into view. The method of claim 4 , wherein when in ambient mode, the number, location, or type of visual cues are limited. 6 . The method of claim 5 , wherein the number, the locations, and the types of visual cues displayed when in ambient mode are customized based on personal usage history.

7. The method of claim 1, wherein the content displayed by the active window is retrieved from a display buffer after the active window has drawn the content to the display buffer.

8. The method of claim 1, wherein the content displayed by the active window is retrieved from a visible portion of the active window and a non-visible portion of the active window.

9. A system comprising: processing unit; as well as A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by the processing unit, cause the processing unit to: receiving an indication that discovery mode has been activated; Retrieve the content displayed by the active window; identifying an image within the retrieved content; identifying an image subject entity within the identified image; Determining whether the image subject entity is applicable to features provided by the OS or features provided by a third party; Overlaying an entity visual prompt onto the active window close to the image main entity; receiving an indication of a selection of the entity visual cue; and Call features provided by the OS or features provided by a third party.

10. The system of claim 9, wherein the entity visualization cue comprises a shape displayed within or adjacent to the image subject entity.

11. The system of claim 10, wherein the computer-executable instructions further cause the processing unit to: In response to an indication that a system cursor is moved over or near the image subject entity, the entity visual cue is highlighted.

12. The system of claim 10, wherein the computer-executable instructions further cause the processing unit to: In response to an indication that a system cursor is moved over or near the entity visual cue, the image subject entity is highlighted.

13. The system of claim 9, wherein the OS feature or the third-party feature extracts the image subject entity from the identified image.

14. The system of claim 9, wherein the computer-executable instructions further cause the processing unit to: identifying a background portion of the image; and In response to determining that a system cursor is moving over the background portion of the image, the image is highlighted.

15. The system of claim 9, wherein the computer-executable instructions further cause the processing unit to: In response to the selection of the entity visual cue, displaying an entity action list; and A selection is received from the entity action list to select the OS-provided feature or the third-party-provided feature.

16. A computer-readable storage medium having computer-readable instructions encoded thereon that, when executed by a processing unit, cause a system to: receiving an indication that discovery mode has been activated; Retrieve the content displayed by the active window; Identify paragraph entities within the retrieved content; Determine whether the paragraph entity applies to features provided by the OS or features provided by a third party; In response to finding that the cursor moves on or near the paragraph entity, overlaying an entity visual prompt on the active window close to the paragraph entity; receiving an indication of a selection of the entity visual cue; and A feature provided by the OS or a feature provided by the third party is called.

17. The computer-readable storage medium of claim 16, wherein the instructions further cause the processing unit to: extracting text depicted in the paragraph entity from the retrieved content; and The extracted text is provided to a feature provided by the OS or a feature provided by the third party.

18. The computer-readable storage medium of claim 16, wherein a system cursor changes to the discovery cursor in response to entering the discovery mode, and wherein the discovery cursor changes to an "I" cursor in response to moving over or near the paragraph entity.

19. The computer-readable storage medium of claim 18, wherein the instructions further cause the processing unit to: Displaying a caret near the "I" cursor within the paragraph entity; receiving an instruction for an editing command; applying the editing command at the location of the caret; and The content displayed by the active window is modified to reflect the results of the editing command.

20. The computer-readable storage medium of claim 18, wherein the instructions further cause the processing unit to: In response to a text selection command, text is selected from the paragraph entity at the "I" cursor position.