Unknown word predictor and content integrated translator
By using personalized private vocabulary analysis storage (VAS) and machine learning technology, it automatically identifies and provides contextual explanations for unknown words, solving the problem of unknown words encountered by users in reading or communication, and improving comprehension and communication efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2017-05-16
- Publication Date
- 2026-05-12
AI Technical Summary
Users often struggle to understand the meaning of unknown words when reading or communicating, and current technologies cannot effectively provide immediate explanations or translations.
By establishing a personalized private vocabulary analysis store (VAS), user language usage patterns are captured. Using machine learning and natural language processing technologies, unknown words are automatically identified and context-related annotations or translations are provided, such as synonyms, definitions, and examples, enriching users' vocabulary understanding.
It improves users' ability to understand unknown words, enabling them to better comprehend complex content and enhance the reading and communication experience.
Smart Images

Figure CN115879445B_ABST
Abstract
Description
[0001] Related patent applications
[0002] This application is a divisional application of the invention patent application with international application number PCT / US2017 / 032780, international application date May 16, 2017, priority date May 31, 2016, entry into the Chinese national phase date November 27, 2018, and Chinese application number 201780032932.6. Background Technology
[0003] Users who want to know the definition of an unknown word can look it up in a copy or online dictionary. Search engines or personal assistant applications can also provide word definitions as direct search results. Additionally, automated computer translation of text from one language to another makes more content available to more people. Summary of the Invention
[0004] This summary provides a simplified overview of a series of concepts further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation as an aid in determining the scope of the claimed subject matter.
[0005] The technology described in this paper enables users to enrich their vocabulary through annotations and / or automatic translations of specific words predicted to be unknown to a particular user. Translations and / or annotations can be performed within the document the user is reading, based on context. For example, unknown words can be replaced with known synonyms or linked to definitions. The user experience is enriched by adaptive, intelligent, embedded annotations that explain the meaning of predicted unknown words. The original content can be a web page, book, paper, email, social media post, blog entry, or any other form of communication that includes linguistic content. Annotations can take the form of explanations of people, places, or things; definitions; translations; synonyms; one or more usage examples; another piece of real content using the same word; or some other form of comprehension aid. In one aspect, unknown words are replaced with known synonyms in the text. The use of this technology allows users to understand content that would otherwise be meaningless due to a limited vocabulary.
[0006] This technology is tailored to individual users by understanding their specific language vocabulary. The system captures language usage patterns when users consume content (news, entertainment, social media posts), perform document reading / creation / editing activities, or communicate through typical online channels (such as chat applications, email applications, voice / video conferencing systems). These patterns can be attached to and maintained in a rich, user-specific private vocabulary analysis store (VAS). A VAS is a knowledge base comprising one or more databases or data storage and analysis components that describes how a user uses one or more languages (and how that usage evolves / changes over time) in terms of patterns, statistics, specific word lists, frequencies, and contextual data. Attached Figure Description
[0007] The various aspects of the technology described in this application are described in detail below with reference to the accompanying drawings, wherein:
[0008] Figure 1 This is a block diagram of an exemplary computing environment suitable for implementing various aspects of the techniques described herein;
[0009] Figure 2 This is a diagram illustrating an exemplary computing environment that depicts aspects of the technology described herein, including an unknown word identification and annotation engine;
[0010] Figure 3 The diagram shows an interface with clearly emphasized unknown terms, based on the technical aspects described in this article.
[0011] Figure 4 It is a diagram illustrating an interface with explanatory annotations for unknown terms, based on the technical aspects described in this article;
[0012] Figure 5 It is a diagram of an interface with definitional annotations for unknown terms, based on the technical aspects described in this article;
[0013] Figure 6 This is a diagram illustrating an interface with translation and synonym annotation for unknown words, based on the technical aspects described in this article.
[0014] Figure 7 This is a diagram illustrating a method for generative language understanding aids based on aspects of the technology described in this paper;
[0015] Figure 8 This is a diagram illustrating a method for generative language understanding aids based on aspects of the technology described in this paper;
[0016] Figure 9 This is a diagram illustrating a method for generative language understanding aids based on aspects of the technology described in this paper; and
[0017] Figure 10 This is a block diagram of an exemplary computing environment suitable for implementing various aspects of the techniques described herein. Detailed Implementation
[0018] The technology of this application is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of this patent. Rather, the inventors have anticipated that the claimed subject matter may also be embodied in other ways to include steps or combinations of steps that differ from those described in this document, incorporating other existing or future technologies. Furthermore, although the terms “step” and / or “box” may be used herein to refer to different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and except where the order of individual steps is explicitly described.
[0019] The technology described in this paper enables users to enrich their vocabulary through annotations and / or automatic translations of specific words predicted to be unknown to a particular user. Translations and / or annotations can be performed within the context of the document the user is reading. For example, unknown words can be replaced with known synonyms or linked to definitions. The user experience leverages adaptive, intelligent, embedded annotations that interpret the meaning of predicted unknown words to enrich the original content. The original content can be a web page, book, document, scanned paper, email, social media post, blog entry, image including text, or any other form of communication including linguistic content. Annotations can take the form of explanations of people, places, or things; definitions; translations; synonyms; one or more usage examples; another piece of real content using the same word; or some other form of comprehension aid. The use of this technology allows users to understand content that would otherwise be meaningless due to a limited vocabulary.
[0020] This technology is tailored to individual users by understanding their specific language vocabulary. As users consume content (news, entertainment, social posts), perform document reading / creation / editing activities, or communicate through typical online channels, the system captures language use patterns maintained in a rich, private vocabulary analysis store (VAS) specific to each user. The VAS is a knowledge base describing how a user is using one or more languages (and how that usage has progressed / changed over time) in terms of patterns, statistics, specific word lists, frequencies, and contextual data. The VAS can also host records of ongoing translation requests from users to one or more connected dictionary / translation services. Requested words can be entered into a phrasebook, which can be part of or separate from the VAS. The phrasebook can also be populated with words identified as unknown to the user by a machine classifier. As used herein, the phrase machine classifier can include any statistical modeling technique suitable for determining whether a word is known or unknown to a person. The VAS can also host or connect to publicly available statistics and metadata regarding word usage rates and certain words for a specific audience. Audiences can be defined by various demographic characteristics such as geographic location. For example, the usage rate of certain words in the English language can differ significantly between people living in Australia and those living in the United States. Significant differences can exist between regions within a country. Audiences can be defined based on age, as word usage can vary across generations. Audiences can also be defined by profession. For example, doctors, lawyers, computer programmers, and psychologists may be familiar with individual words commonly used in their professions.
[0021] Audiences can also be defined by interests. For example, sports enthusiasts might have a deep understanding of words used in a sports context, while political enthusiasts might have a deep understanding of words used in a political context but a superficial understanding of words used in a sports context. User data, including reading logs, can be used to determine user interests. Words can be similarly associated with context. Even if a sports enthusiast has not been exposed to or used a particular word, such as "southpaw," they are more likely than a non-sports enthusiast to know the meaning of "southpaw" (e.g., left-handed pitcher). Therefore, a statistical classifier using context as input can assign lower confidence to the word "southpaw" being unknown to sports enthusiasts than to non-sports enthusiasts, even given the same levels of exposure and use for both.
[0022] A user's VAS can be maintained and kept in the background as a source of data and signals for the user to perform unknown word prediction: the wealth of information provided in the VAS can be used to predict, with a known confidence level, whether any given word is expected to be known to the user. In other words, the content of the VAS can be used as input to a classifier to calculate, with a certain confidence level, whether a given word is known or unknown. If a word is classified as unknown and the user is opening content containing that word (e.g., an article), the system enriches the original content within the article with embedded annotations about the unknown word. This happens smoothly without any request or action from the user. Aspects of the technology are not limited to using binary classifiers. For example, multi-class classifiers can be used to classify words into one of several different categories such as known, familiar, and unknown. In other cases, the classifier can assign a value between, for example, 0 and 1, indicating the level of familiarity the user has with a particular word at a given time.
[0023] In one aspect, a binary classification machine is used to categorize words as known or unknown. As an initial step, words with a commonality or frequency of use above a threshold within the language can be excluded from the analysis and designated as known. For example, words such as “the,” “you,” “him,” and “her” are widely used and known in the English language. A binary classification machine can be trained to analyze user data related to lexical knowledge in a VAS, such as the text of what is read or written, the time spent on the text on the page, and reading analytics from technologies such as gaze detection, where the time spent on individual words can be measured. Gaze detection uses cameras(s) to track eye movement on a page or screen. Returning to a specific word or appearing to pause on a word can indicate that the word is unknown to the user. The classification machine can also utilize audience data, as previously described. Users can be associated with or assigned to one or more audience profiles or specific demographic segments. These profiles or segments can provide additional input to the classifier and be used to determine whether a word is likely known or unknown.
[0024] User segments or profiles can be used to create an initial language profile for a user, which can be modified upon receiving additional activity. Even without content consumption data specific to a user, or with limited content consumption data specific to a user, knowledge of geolocation, occupation, age grouping, gender, social preferences, and activities (e.g., likes on sports pages, push notifications, etc.) can associate users with specific demographic segments, which are then used to initially set the expected vocabulary for that user. Segmentation and profile data can be combined with actual content consumption data, creative activity, communication history, and more for a user to update their expected vocabulary.
[0025] Looking up words in dictionaries, through translation services, personal assistant apps, search pages, or other technologies can strongly indicate that the word is unknown. In other words, technology can assume that recently searched words are unknown and provide annotations when the word appears in future text. The classification can change to known when a word is used by a user or when the user is exposed to it multiple times. In one aspect, words recently searched by a user are automatically added to a phrasebook. In another scenario, the system can use recently searched words as annotations for synonyms of unknown words. This helps the user learn recently searched words by associating them with synonyms the user already knows.
[0026] Other sources of information about a user's language knowledge can include spelling and grammar errors made when creating documents, emails, etc. In one aspect, technology collects instances where spelling or grammar errors are automatically corrected or manually corrected by the user.
[0027] In one aspect, the system scans the user's VAS (Visual Identity System) for language statistics and metadata captured and maintained in the user's phrasebook (which may include a list of known and unknown words with scores and metadata) of the content provided to the user (e.g., via a browser, application, OS extension, or pre-installed components that render online content, such as websites or social media posts). The system checks if any words in the content are listed as unknown in the phrasebook, and if so, generates annotations (e.g., injecting additional HTML elements to visualize synonyms or examples). The annotations enrich the content, thus helping the user grasp the meaning of unknown words in context, better understand the content, and use the words as part of their vocabulary to begin with.
[0028] In one aspect, this annotation involves using known synonyms to replace unknown words in the text. A given word can have many synonyms. Contextual synonyms share context with the unknown word. For example, sports could be the context of a baseball website, and politics could be the context of a political website or politically related social media posts. The topical closeness of various synonyms can be obtained from the definitions or usage examples of the words provided in a knowledge base. Topics can also be determined by analyzing the frequency of occurrence in content with the same context. For example, synonyms for an unknown word appearing in a sports article can be selected from many available synonyms based on their frequency of occurrence within a corpus of sports articles compared to the frequency of other synonyms used in the sports context.
[0029] When selecting contextual synonyms, the overall frequency of use can also be considered. Therefore, synonyms that appear infrequently overall but are used at a higher rate in sports articles can be selected in a sports context, especially if the viewer appears familiar with the article. In one aspect, only synonyms that are likely known to the user are evaluated for their closeness to the specific context.
[0030] Annotations can be provided in the same language as the content or in a different language. For example, the content could be in Spanish and the annotations in English. The language of the annotations can be selected by the user through a preference interface. Alternatively, the system can learn the languages understood by individual users, derive preferences from observing user events, and select the language of the annotations accordingly. For example, a user who typically looks up English words in an English language dictionary could receive English language annotations while reading English language content. Conversely, a user who typically seeks to translate English words into Spanish could receive Spanish annotations while reading English content. In one aspect, translations of unknown words from the first language to the user's native language (or any language the user is more fluent in than the content language) can be provided when all available synonyms in the first language may also be unknown to the user. The user's known languages can be explicitly provided by the user or learned by observing the language of the content consumed or created by the user.
[0031] Alternatively, footnotes can provide content in multiple languages. For example, footnotes can provide contextual synonyms for unknown English words in English and French translations of unknown English words. Footnotes can also provide content in languages the reader is very familiar with and in a second language the reader is less fluent in.
[0032] The system can track user reactions or receive feedback on predictions that can be automatically used to improve and optimize the unknown word prediction system. The system maintains this knowledge of the user's active vocabulary, along with any newly added extensions. As long as the user naturally consumes content (from virtually any website that provides content in the form of articles, videos, or audio), the system can identify recently searched words that are unknown or rarely used, and embed their best translations or annotations to help the user better understand and begin using them.
[0033] Having briefly described an overview of the various aspects of the technology described herein, the following describes an exemplary operating environment suitable for use in implementing this technology.
[0034] Now go to Figure 1A block diagram illustrating an example operating environment 100 in which some aspects of this disclosure can be adopted is provided. It should be understood that the arrangements and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, and groupings of functions, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be omitted entirely for clarity. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by one or more entities can be performed by hardware, firmware, and / or software. For example, some functions can be performed by a processor executing instructions stored in memory.
[0035] In addition to other components not shown, the example operating environment 100 includes a number of user devices, such as user devices 102a and 102b to 102n; a number of data sources, such as data sources 104a and 104b to 104n; a server 106; and a network 110. It should be understood that... Figure 1 The environment 100 shown is an example of a suitable operating environment. Figure 1 Each of the components shown can be transmitted via any type of computing device (such as, for example, a combination of...) Figure 10 The computing device 1000 described herein may be used for implementation. These components may communicate with each other via a network 110, which may include, but is not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). In an exemplary embodiment, network 110 includes the Internet and / or cellular networks, as well as any of a variety of possible public and / or private networks.
[0036] User devices 102a and 102b to 102n can be client devices on the client side of operating environment 100, while server 106 can be on the server side of operating environment 100. In one aspect, text content can be accessed via user devices. Various aspects of the technology can modify the text content to provide annotations that help users understand predicted unknown words. User actions (including text creation) via user devices can be monitored to build and continuously enrich user-specific VAS. User actions across multiple devices can be compiled into a session consisting of multiple events and entries including detailed word usage with frequency, context, phrases, and content / document attributes. For example, a user can read a digital book or web page through an application on a user device. This reading activity can be a session or part of a session, depending on the storage model being modeled as multiple records / entries / documents. The session is organized by all the words consumed by the user, as well as frequency, statistics, metadata, context, and timestamps, and is a key input for VAS post-processing and enrichment. Similarly, a user can compose emails or other documents, which can be used to determine the user's vocabulary and ultimately determine whether other words are known or unknown.
[0037] Server 106 may include server-side software designed to work in conjunction with client-side software on user devices 102a and 102b to 102n to implement any combination of the features and functionalities discussed in this disclosure. For example, server 106 may collect data from an individual or a group of individuals that can be used to make a decision about whether a word is known or unknown to the user. Additionally, the determination about whether a word is known or unknown may be made on server 106 or on the user device. In one aspect, a browser plugin or some other application feature may collect user data, pass content to an unknown word predictor, and receive a list of potentially unknown words in the content, along with probability and / or confidence levels and / or additional metadata. URLs and / or URIs that can be used by the client and / or exposed to the user for additional information, synonyms, or examples may also be provided. For example, server 106 may know about online dictionaries associated with a particular language, their APIs, etc. Metadata about the actual links or which online services should be queried and retrieved for synonyms, translations, or examples may be included. Querying internal resources or publicly available online resources for synonyms, examples, etc., may be the responsibility of the annotation engine 224 or a separate component invoked by it. Client-side applications can subsequently modify content, for example, by updating the document object model of the web page, and output / inject rich content utilizing hints, translations, contextual synonyms, examples, and links. This division of the operating environment 100 is provided to illustrate an example of a suitable environment, and for each implementation, there is no requirement for server 106 and any combination of user devices 102a and 102b to 102n to remain as separate entities.
[0038] User devices 102a and 102b to 102n may include any type of computing device capable of being used by a user. For example, in one aspect, user devices 102a to 102n may be, as described herein, [the following is a separate, unrelated section:] Figure 10 The type of computing device described. By way of example and not limitation, a user device can be implemented as a personal computer (PC), a laptop computer, a mobile or mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a virtual reality headset, augmented reality glasses, a holographic system, a personal digital assistant (PDA), an MP3 player, a global positioning system (GPS) or device, a video player, a handheld communication device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an electrical appliance, a consumer electronics device, a workstation, or any combination of the devices described herein, or any other suitable device.
[0039] Data sources 104a and 104b to 104n may include data sources and / or data systems that provide content to operating environment 100 or in conjunction with it. Figure 2 The system 200 described herein comprises the following components (e.g., in one aspect, one or more data sources 104a to 104n provide (or are available for access to) content 210). Data sources 104a and 104b to 104n may be separate from user devices 102a and 102b to 102n and server 106, or may be incorporated into and / or integrated into at least one of those components. Data sources may include knowledge bases storing information, statistics, and metadata about words and language. For example, data sources 104a to 104n may include online dictionaries, knowledge bases, translation services, contextual dictionaries, spell-checking engines, search engines, or other services that can be used to generate annotations. Additionally, user interactions with services provided by data sources 104a to 104n may be collected in the VAS and used to determine whether a word is unknown to the user. For example, words looked up in a dictionary may be added to a phrasebook as part of the VAS.
[0040] Operating environment 100 can be used for implementation Figure 2 One or more of the components of the system 200 described herein include components for collecting user data, monitoring read and write events, generating annotations, and identifying unknown words.
[0041] Now combine Figure 1 refer to Figure 2This document provides block diagrams illustrating aspects suitable for implementing the techniques described herein and generally designated as System 200, representing various aspects of an example computing system architecture. System 200 analyzes content, identifies words in the content that a particular user may not know at a given time, and generates explanatory annotations for those unknown words. System 200 represents only one example of a suitable computing system architecture. Other arrangements and elements may be used in addition to or in place of those shown, and some elements may be omitted entirely for clarity. Furthermore, regarding the operating environment 100, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. For example, aspects of the lexical enrichment component 220 may be located on different computing devices or may all be located on the same device.
[0042] Example system 200 includes network 110, which combines Figure 1 The components of system 200 described and communicatively coupled thereto include a vocabulary enrichment component 220. The vocabulary enrichment component 220 (including its components 222, 224, 226, 227, 228, and 230) can be implemented as a set of compiled computer instructions or functions, program modules, data models, databases, data storage, computer software services, or in one or more computer systems (such as, for example, in conjunction with...). Figure 10 The arrangement of the process executed on the computing device 1000 is described.
[0043] In one aspect, the functions performed by the components of system 200 are associated with one or more applications, services, or routines. Specifically, such applications, services, or routines may operate on one or more user devices (such as user device 102a), servers (such as server 106), may be distributed across one or more user devices and servers, or may be implemented in the cloud. Furthermore, in some aspects, these components of system 200 may be distributed across a network including one or more servers (such as server 106) and client devices (such as user device 102a), may be distributed in the cloud, or may reside on a user device such as user device 102a. Additionally, these components, the functions performed by these components, or the services performed by these components may be implemented at suitable abstraction layers, such as the operating system layer, application layer, hardware layer, etc., of a computing system. Alternatively or additionally, the functions of these components and / or aspects of the techniques described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that can be used include Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Products (ASSPs), Systems-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on. Furthermore, although the functionality of the specific components shown in Example System 200 is described herein, it is anticipated that in some respects the functionality of these components may be shared or distributed across other components.
[0044] The vocabulary enrichment component 220 can receive or retrieve content 210 used to identify unknown words when it receives an indication that the user is accessing content. For example, the vocabulary enrichment component 220 can access website content when it receives an indication that the user has navigated to a website and the website's URL or URI. Similarly, the vocabulary enrichment component 220 can access text in a book opened in an e-reader application. The content can be web pages, social media posts, books, audio recordings including voice, video recordings including voice, presentations, documents, user interfaces, and so on.
[0045] In the case of audio or video recordings, the natural language processing component can be used to perform speech-to-text conversion. The vocabulary enrichment component 220 can then analyze the text to generate a list of possible unknown words and provide annotations or other forms of comprehension assistance. In the case of video content, annotations can take the form of overlay text or rich text precisely time-aligned with the appearance of unknown words in the video. In the case of audio content, annotations can take the form of subtitles precisely time-aligned with the appearance of unknown words in the audio.
[0046] The vocabulary enrichment component 220 includes an unknown word predictor 222, an annotation engine 224, a user language data store 226, a common language data store 228, and a user activity monitor 230. These components work together to identify words the user may not know and generate annotations within the content that allow the user to better understand unknown words in a specific context. Annotations can be displayed within the annotation content 240. See below for reference. Figure 3-7 Examples of the content provided in the comments.
[0047] Now go to Figure 3 The illustrations, based on aspects of the technology described herein, depict a content display including visually distinguishable unknown words. In this case, content 300, including the opening paragraphs of the novel *Moby Dick*, includes several underlined words to provide visual indications that aid comprehension and can be accessed through word selection. The visually distinguishable words in content 300 include the words “spleen” 310, “flourish” 312, “Cato” 314, “insular” 316, “Manhattoes” 318, “circumambulate” 320, and “sentinels” 322. While underlining is used in this example, other methods of visually distinguishing words are possible. For example, words can be visually distinguished by their font, color, size, underline, double underline, and other visible characteristics to indicate that a word is selectable.
[0048] Figure 3 The visually distinguishable words in the examples were not randomly selected or chosen for a general audience. Instead, the visually distinguishable words were chosen based on the individual viewer's vocabulary, potentially using knowledge unknown to the specific user. As used in this article, unknown means that the user is unaware of the definition or meaning of the word or phrase. The user may have seen the word before, but the word may still be unknown.
[0049] The state of a word (e.g., whether it is known or unknown to the user) can change over time. An unknown word can become known as the user encounters or uses it. For example, a given user consuming the exact same content (e.g., an article) at different times might receive different predicted unknown words. The state of a word can change because the user may have become familiar with some of the initially predicted unknown words (through consuming additional content, using them when creating documents, etc.). In such cases, the system has already tracked the fact that the user has used these initially unknown words and has already represented the words as known at the current time. Therefore, the same user consuming the same article several months later might receive different predictions for potentially unknown words.
[0050] Any two users consuming the exact same content (e.g., articles) can obtain different predictions of unknown words, and possibly different synonyms, links, and visual aids, based on each user's unique use of the exposed data via the annotation engine 224. Figure 4 , Figure 5 and Figure 6 The illustration shows the different understandings or invitations that can be accessed by selecting visually distinguishable words.
[0051] Now go to Figure 4 Explanatory notes 330 are illustrated according to various aspects of the technology described herein. Explanatory notes provide information about people, places, entities, or things. When the user selects the visually distinguishable word "Manhattoes" 318, an explanation 332 for Manhattan is provided in explanatory notes 330. In this case, the explanation is "Manhattan: A borough of New York City." Additionally, underlining Manhattan 334 indicates that the word Manhattan is optional and can lead to (e.g., a link / URL to a specific document or online resource, such as Wikipedia) further details about Manhattan. In one aspect, the explanatory notes include a link back to the source content where further information can be accessed. For example, the link could be to a Wikipedia entry or a knowledge base entry that provides additional information about Manhattan.
[0052] Turning now to 5, a defining annotation 340 is provided according to the aspect of the technology described herein. The defining annotation provides a definition of the word. When a word has multiple definitions, the context in which the word is used within content 300 can be considered. For example, the use of a word as a noun, verb, adjective, or some other part of speech can be determined to select the definition most likely reflecting the author's intended use of the word. In this example, defining annotation 340 may emerge when the user selects the word “circumambulate” 320. The definition 342 of circumambulate 320 is “to walk all the way around something.” Definitions can be accessed from dictionaries to generate the annotation. In one aspect, defining annotation 340 includes links to source content such as online dictionaries.
[0053] As an alternative to the definition, example annotations (not shown) can be provided. The system can retrieve real-world examples of unknown words being used from content consumed across users, geographic locations, and contexts. The system can select the most relevant and / or most popular examples. Examples can be complete sentences using a specific word. Examples can also be derived from knowledge bases such as those using wizards or dictionaries.
[0054] Turning to point 6, translation annotations and synonym annotations are provided according to aspects of the technology described herein. In this case, a single annotation 350 includes a translation in a different language and a synonym for the unknown word in the language of the unknown word. Annotation 350 may appear when the user selects the unknown word “sentinels” 322. Annotation 350 includes a translation of “sentinels” in Greek 352 and a synonym 354 in English. Translations may be provided when aspects of the technology determine that a particular user speaks more than one language. For example, the technology may determine that a user is familiar with more than one language when the user is observed reading text in a different language or writing in a different language. For example, a particular user viewing content 300 may speak English as a second language and Greek as a first language. In this case, providing a translation 352 of the unknown word in the second language in the user's first language may be helpful for understanding. Aspects of the technology are not limited to any particular language or alphabet. For example, if the viewer is determined to understand Spanish, “sentinels” may have been translated into Spanish as “centinela” (sentinel). Synonym 354 is provided in the same language as the base content 300, which is English in this example. Both translation 352 and synonym 354 can be linked to the source content, or alternative translations or additional synonyms can be accessed.
[0055] Annotations can be provided in the same language as the content or in a different language. For example, the content could be in Spanish and the annotations in English. The language of the annotations can be selected by the user through a preference interface. Alternatively, the system can learn an individual's language understanding, derive preferences from observing user events, and select the annotation language accordingly. For example, a user who typically looks up English words in an English language dictionary could receive English language annotations while reading English language content. Conversely, a user who typically seeks to translate English words into Spanish could receive Spanish language annotations while reading English content. In one aspect, translations of unknown words from the first language to the user's native language (or any language the user is more fluent in than the content language) can be provided when all available synonyms in the user's first language may also be unknown to the user. The user's known languages can be explicitly provided by the user or learned by observing the language of content consumed or created by the user.
[0056] Alternatively, footnotes can provide content in multiple languages. For example, footnotes can provide contextual synonyms for unknown English words in English and French translations of unknown English words. Footnotes can also provide content in languages the reader is very familiar with and in a second language the reader is less fluent in.
[0057] Return to Figure 2 As described, the unknown word predictor 222 identifies words within content 210 that may be unknown to a given user at a given time. As a first step, the unknown word predictor 222 may use broad filters to identify potentially unknown words. For example, words with a frequency higher than a threshold for use within the language may be excluded as potentially unknown words. In one aspect, the filter is user-specific. For example, the filter may be generally applied based on the user's associated reading level. The reading level may be calculated based on the user's reading content, reading speed, and ongoing analysis of vocabulary use when composing communications.
[0058] In one approach, potentially unknown words are then fed into a machine classifier, which calculates a confidence score indicating that the word is unknown to a specific user at a given time. The machine classifier can receive multiple signals related to the user and the public as input to determine if a word is unknown. The machine classifier can utilize one or more of the following: neural networks, support vector machines, generalized linear modeling, data mining, text mining, natural language processing algorithms, regression, or decision trees. These can be used in isolation or in combination as part of a machine classifier in some configurations to calculate confidence scores or otherwise classify words on a known spectrum / unknown spectrum.
[0059] In one aspect, the signal includes reading data. Reading data can include text from content read by a user, categorization of content read by the user, reading analytics, and other data related to the user's reading habits. Content categorization can be based on reading level, topic, source type, or other factors. Reading level can be determined ergonomically, for example, by a knowledge base describing reading levels for various books. Reading level can also be determined analytically by analyzing sentence length, sentence complexity, and the complexity of the vocabulary used. Topic can be determined by analyzing a knowledge base of known works (such as books). Topic can also be determined by evaluating metadata associated with various types of content, such as web pages. Topic can also be determined directly by considering the content of the text and the categorization method of the source. Topic can be used to define a user's interests. For example, a user can be categorized as having one or more interests, such as sports, religious beliefs, politics, music, etc. Different levels of granularity are possible. Instead of sports, a user can be categorized as having an interest in baseball, football, rugby, basketball, track and field, or hockey.
[0060] On the other hand, signal data includes creation or writing data. Creation or writing data can include email text, social media posts, documents, presentations, and other communications and documents generated by the user.
[0061] On the other hand, signal data includes searched words. Searched words are those words that a user searches for in order to obtain additional information about or a translation of. For example, a user may have already looked up the word in an online dictionary or translation service. A user may have already searched for the word through a search engine, online dictionary, or personal assistant application.
[0062] Signals can also include user feedback. User feedback can be explicit or implicit. Implicit feedback can include user interaction with notes previously presented to the user. For example, when a word is identified as unknown and associated with a note that has been ignored by the user (e.g., unused, unconsumed, no interaction), no interaction can be implicit feedback suggesting that the user is familiar with the word. Similarly, access to an note for interaction with an unknown word suggests that the word has been correctly identified as unknown. Explicit feedback can be provided through mechanisms available in the note. For example, an optional "remove" command can be provided within the note, which, when selected, removes the associated unknown word from the phrasebook and classifies it as known. In this respect, words explicitly identified as known can be used as signals.
[0063] Using all the above signals, the frequency and relevance of exposure or use of various words can be determined and used as input to a machine classifier. Other signals can be obtained. These signals are calculated by performing operations on the original signal data.
[0064] The machine classifier portion of the unknown word predictor 222 generates a list of unknown words. The machine classifier can use any type of technique or combination of techniques capable of determining whether a word is known or unknown. The determination can be binary or represented as a confidence factor describing the probability that a word is unknown. For example, neural networks, support vector machines, generalized linear modeling, data mining, text mining, natural language processing algorithms, regression, or decision trees can be used in isolation or in combination as part of the machine classifier in some configurations. Unknown words from the content can be added to a phrasebook that includes both known and unknown words. Words determined to be unknown can be input for subsequent determinations. Unknown words can be passed to the annotation engine 224.
[0065] As previously referenced Figure 3-6As illustrated, annotation engine 224 can change the appearance of unknown words within the displayed content. Different methods can be used to change the appearance of words. For example, when various aspects of the techniques described herein interact with a web browser, the techniques can update the document object model to change the appearance of the content, thereby making unknown words underlined or otherwise clearly distinguishable from the rest of the text.
[0066] Annotation engine 224 can also construct annotations, which are displayed in response to selection or other interaction with unknown words. Annotations may include: explanations of people, places, entities, or things; definitions of words; usage of words in sentences; translations of words in different languages; synonyms; contextual synonyms; or some other form of comprehension aid. Annotations can be constructed from external sources such as online dictionaries, web pages, and knowledge bases. Annotations may also be supported by internal knowledge bases / content analytics stores that provide real-world examples of the use of specific words in real-world content. Annotations may include links to online sources used to construct the annotations or to other online sources that may provide additional information about unknown words or phrases. The output of annotation engine 224 is the content 240 of the annotations.
[0067] In one variant, annotation engine 224 uses other words with the same or similar meanings to replace words within the content. In one aspect, unknown words are replaced with known words. Known words can be explicitly designated as optional. Upon selection, annotations can be displayed, showing the original unknown word. The reverse approach is also possible. Known words that are synonyms of unknown words can be used to replace the unknown words. This provides an opportunity for the user to be exposed to unknown words. In this case, a general list of unknown words for a specific user, such as those found in a phrasebook associated with the user, can be used to generate a list of synonyms for those words. Synonyms can be compared with words within the content.
[0068] In another variation, a user can request unknown synonyms for a known word. In one aspect, words with available unknown synonyms can be made selectable. In another aspect, a user can identify a word and request unknown synonyms, even if the word is not selectable or is otherwise visually distinguishable.
[0069] In one aspect, contextual synonyms can be selected. Contextual synonyms share context with unknown words. For example, sports could be the context of a baseball website, and politics could be the context of a political website or a politically related social media post. The context of an unknown word can be determined through words in the content, the source of the content, metadata associated with the content, and other information. Context can be determined at various levels of granularity. For example, the context could be sports, or it could be baseball, rugby, basketball, etc.
[0070] The context of various synonyms can be obtained from word definitions or usage examples provided in a knowledge base. Topics can also be determined by analyzing the frequency of occurrence in content with the same context. For example, synonyms for unknown words appearing in sports articles can be selected from many available synonyms based on their frequency of occurrence within a corpus of sports articles compared to the frequency of other synonyms used in sports contexts. In addition to synonyms, various aspects of the technology can utilize contextual synonyms or other understanding aids to replace acronyms or slang.
[0071] When selecting contextual synonyms, the overall frequency of use can also be considered. Therefore, synonyms that appear infrequently overall but are used at a higher rate in sports articles can be selected in a sports context, especially if the viewer appears familiar with the article. In one aspect, only synonyms that are likely known to the user are evaluated for their closeness to the specific context.
[0072] User language data storage 226 may include raw language data collected by user activity monitor 230. The raw language data may include user reading data and reading patterns. The raw language data may include knowledge about individual users, such as the languages used by the user and the user's proficiency in each language. User language data storage 226 may also include a phrasebook listing recently searched words. In one aspect, the phrasebook may be part of lexical analysis storage (VAS) 227.
[0073] VAS 227 may include a subset of user-specific language data. This subset may include all signals from the previously described signals as input to a machine classifier or other statistical modeling techniques.
[0074] The public language data store 228 may include usage patterns within a general group or a specific subset of a general group. As mentioned, subsets can be categorized by audience data defined by demographic information such as age, income, geographic location, country, gender, occupation, education level, lifestyle-related attributes, or other factors. Audiences can also be defined by interests across various subject categories. The public language data store may also include dictionaries and other knowledge bases providing information on the relationships between grammar and words. Words can be grouped by meaning, common usage patterns, or user interests associated with the context.
[0075] User monitor 230 can exist on multiple user devices and generate user reading data to determine whether words are known or unknown. User monitor 230 can generate a series of records or entries for each reading or creation event (including words, statistics, or metadata). User monitor 230 can generate implicit feedback on user interactions with provided annotations. User monitor 230 can rely on information provided by other applications to generate reading statistics or implement techniques to generate such statistics. For example, user monitor 230 can rely on gaze detection technology to determine processing time on various words and phrases and subsequently record words or phrases that appear to slow down the user's reading. User monitor 230 can access application programming interfaces provided by the reader to determine what the user is reading and how fast the user is reading a given book or other content.
[0076] Now go to Figure 7 A method 700 for generating language understanding aids is described in this paper, based on aspects of the technology described herein. Method 700 can be implemented by a system similar to the lexical enrichment component 220.
[0077] At step 710, a user-specific phrasebook is constructed. The words are in the user's first language and may be unknown to a particular user. Words in the phrasebook may be unknown to the user at any given time and may change over time. In one aspect, words searched by the user are added to the phrasebook. For example, words looked up in a dictionary or submitted to a translation service can be included in the phrasebook. The phrasebook may also include references, etc. Figure 2 The terms explained are those identified as unknown to the user by a machine classifier. For example, content can be evaluated by a machine classifier to identify unknown words. Some or all of these words can be annotated in the content and added to a phrasebook. In one aspect, the phrasebook may include words expected to be unknown based on demographic information about the user.
[0078] At step 720, content including text being accessed by a specific user is received. This content can be a web page, social media post, book, audio recording including speech, video recording including speech, presentation, document, user interface, etc. In the case of audio or video recordings, a natural language processing component can be used to perform speech-to-text conversion. In one aspect, the content is intercepted by a browser plugin for analysis when the user navigates to a web page. The plugin can perform analysis or pass the content to another component installed on the client device or server for analysis. Document applications, presentation applications, e-readers, or other content applications can similarly retrieve content for analysis.
[0079] At step 730, unknown words from the user-specific phrasebook are determined within the text by comparing words in the text with words in the user-specific phrasebook. In one aspect, each word in the text is compared with a word in the phrasebook to determine if the word might be unknown. Unknown words can then be marked as unknown.
[0080] At step 740, the appearance of the unknown word within the content is changed to provide an indication that the unknown word is optional. Various ways to change the appearance of an unknown word include changing its color, underlining it, double-lined it, and so on. In one aspect, the appearance is changed by adding a link to the unknown word, and the appearance changes automatically based on display attributes within the application displaying the content. For example, a document application might change the color of each linked word from black to blue.
[0081] At step 750, an instruction from the user to interact with the unknown word is received. The user can interact with the word on a touchscreen, using voice commands or gestures, using a mouse or touchpad, or through some other method. Word selection is one example of interaction, and hovering over a word is another.
[0082] At step 760, annotations providing additional information about the unknown word are output for display. Annotations may include: explanations of people, places, entities, or things; definitions of the word; usage of the word in a sentence; translations of the word in different languages; synonyms; or some other form of comprehension aid. Annotations can be built from external sources such as online dictionaries, web pages, and knowledge bases. Annotations may include links to online sources used to build the annotations or other online sources that may be relevant to providing additional information about the unknown word or phrase.
[0083] In one variation, the technique described herein uses other words with the same or similar meanings to replace words within the content. In one aspect, unknown words are replaced using known words. Known words can be explicitly designated as selectable. Upon selection, a note can be displayed showing the original unknown word. The reverse approach is also possible. Known words that are synonyms of unknown words can be used to replace the unknown words. This provides an opportunity to expose the user to unknown words. In this case, a general list of unknown words for a specific user, such as those found in a phrasebook associated with the user, can be used to generate a list of synonyms for those words. Synonyms can be compared with words within the content.
[0084] In another variation, a user can request unknown synonyms for a known word. In one aspect, words with available unknown synonyms can be made selectable. In another aspect, a user can identify a word and request unknown synonyms, even if the word is not selectable or is otherwise visually distinguishable.
[0085] In one aspect, contextual synonyms can be selected. Contextual synonyms share context with unknown words. For example, sports could be the context of a baseball website, and politics could be the context of a political website or a politically related social media post. The context of an unknown word can be determined by words in the content, the source of the content, metadata associated with the content, and other information. Content can be determined at various levels of granularity. For example, the context could be sports, or it could be baseball, rugby, basketball, etc.
[0086] The context of various synonyms can be obtained from word definitions or usage examples provided in a knowledge base. Topics can also be determined by analyzing the frequency of occurrence in content with the same context. For example, synonyms for unknown words appearing in sports articles can be selected from many available synonyms based on their frequency of occurrence within a corpus of sports articles compared to the frequency of other synonyms used in sports contexts. In addition to synonyms, various aspects of the technology can utilize contextual synonyms or other understanding aids to replace acronyms or slang.
[0087] When selecting contextual synonyms, the overall frequency of use can also be considered. Therefore, synonyms that appear infrequently overall but are used at a higher rate in sports articles can be selected in a sports context, especially if the viewer appears familiar with the article. In one aspect, only synonyms that are likely known to the user are evaluated for their closeness to the specific context.
[0088] Now go to Figure 8 A method 800 for generating language understanding aids is described in this paper, based on aspects of the technology described herein. Method 800 can be implemented by a system similar to the lexical enrichment component 220.
[0089] At step 810, text content in a first language is received. The content can be a web page, social media post, book, audio recording including speech, video recording including speech, presentation, document, user interface, etc. In the case of audio or video recordings, a natural language processing component can be used to perform speech-to-text conversion. In one aspect, the content is intercepted by a browser plugin or extension for analysis when the user navigates to the web page. The plugin can perform analysis or pass the content to another component installed on the client device or server for analysis. Document applications, presentation applications, e-readers, or other content applications can similarly retrieve content for analysis.
[0090] At step 820, the text content is scanned to identify words with a confidence level higher than a threshold of user-unknown words determined by a machine classifier using text previously read or written by the user as input. In one implementation, this can be done in real time. In another implementation, unknown word estimation occurs periodically or is triggered by specific events. For example, words can be checked for unknown word estimation when new content is being consumed.
[0091] At step 830, text content is output for display to the user, where words have a visual appearance different from other words in the text content to indicate that a word is selectable. Various ways to change the appearance of an unknown word include changing its color, underlining it, double-underlining it, and so on. In one aspect, the appearance is changed by adding links to the unknown word, and the appearance changes automatically based on display attributes within the application displaying the content. For example, a document application might change the color of each linked word from black to blue. The user can select a word or otherwise interact with it to receive annotations, such as previous references. Figure 2-6 Those described. Annotations can help with understanding, such as translations, definitions, usage examples, or contextual synonyms for unknown words.
[0092] Now go to Figure 9 A method 900 for generating language understanding aids is described in this paper, based on aspects of the technology described herein. Method 900 can be implemented by a system similar to a lexical enrichment component 220.
[0093] In step 910, a user-specific vocabulary analysis store (VAS) is constructed by collecting reading and writing data specific to the user. User actions (including text creation) on the user's device can be monitored to construct the user-specific VAS. User actions across multiple devices can be compiled into a single record. For example, a user might read a book or web page through an app on their device. This reading activity can form part of a reading record. Similarly, a user can compose emails or other documents that can be used to identify the user's vocabulary and ultimately determine whether other words are known or unknown.
[0094] A Virtual Aid (VAS) is a knowledge base describing how a user uses one or more languages (and how that usage progresses / changes over time) in terms of patterns, statistics, specific word lists, frequencies, and contextual data. A VAS can also host a record of ongoing translation requests from users to one or more connected dictionary / translation services. Requested words can be entered into a phrasebook, which can be part of or separate from the VAS. The phrasebook can also be populated with words identified as unknown to the user by a machine classifier. As used herein, a phrase machine classifier can include any statistical modeling technique suitable for determining whether a word is known or unknown to a person. A VAS can also host or connect to publicly available statistics and metadata about word usage by a specific audience and certain words. Audiences can be defined by various demographic characteristics such as geographic location. For example, the usage of certain words in the English language can differ significantly between people living in Australia and those living in the United States. Significant differences can exist between regions within a country. Audiences can be defined based on age, as there can be generational differences in word usage. Audiences can also be defined by profession. For example, doctors, lawyers, computer programmers, and psychologists may be familiar with individual words common to their professions.
[0095] Audiences can also be defined by interests. For example, a sports enthusiast might have a deep understanding of words used in a sports context, while a political enthusiast might have a deep understanding of words used in a political context but a superficial understanding of words used in a sports context. User data, including reading logs, can be used to determine user interests. Words can be similarly associated with context. Even if a sports enthusiast has not been exposed to or used a particular word, such as "southpaw," they are more likely than a non-sports enthusiast to know the meaning of "southpaw" (e.g., left-handed pitcher). Therefore, a statistical classifier using context as input can assign lower confidence to the word "southpaw" being unknown to a sports enthusiast than to a non-sports enthusiast, even given the same levels of exposure and use for both.
[0096] A user's VAS can be maintained and kept in the background as a source of data and signals for the user to perform unknown word prediction: the wealth of information provided in the VAS can be used to predict, with a known confidence level, whether any given word is expected to be known to the user. In other words, the content of the VAS can be used as input to a classifier to calculate, with a certain confidence level, whether a given word is known or unknown. If a word is classified as unknown and the user is opening content containing that word (e.g., an article), the system enriches the original content within the article with embedded annotations about the unknown word. This happens smoothly without any request or action from the user. Aspects of the technology are not limited to using binary classifiers. For example, multi-class classifiers can be used to classify words into one of several different categories such as known, familiar, and unknown. In other cases, the classifier can assign a value between, for example, 0 and 1, indicating the level of familiarity the user has with a particular word at a given time.
[0097] In one aspect, a binary classification machine is used to classify words as known or unknown. As an initial step, words with a commonality or frequency of use above a threshold within the language can be excluded from the analysis and designated as known. For example, words such as “the,” “you,” “him,” and “her” are widely used in the English language and are known. A binary classification machine can be trained to analyze user data related to lexical knowledge in a VAS, such as the text of what is read or written, the time spent on the text on the page, and reading analytics from technologies such as gaze detection, where the time spent on individual words can be measured. Gaze detection uses cameras(s) to track eye movement on a page or screen. Returning to a specific word or appearing to pause on a word can indicate that the word is unknown to the user. The classification machine can also utilize audience data, as previously described. Users can be associated with or assigned to one or more audience profiles or specific demographic segments. These profiles or segments can provide additional input to the classifier and be used to determine whether a word is likely to be known or unknown.
[0098] User segments or profiles can be used to create an initial language profile for a user, which can be modified upon receiving additional activity. Even without content consumption data specific to a user, or with limited content consumption data specific to a user, knowledge of geolocation, occupation, age grouping, gender, social preferences, and activities (e.g., likes on sports pages, push notifications, etc.) can associate users with specific demographic segments, which are then used to initially set the expected vocabulary for that user. Segmentation and profile data can be combined with actual content consumption data, creative activity, communication history, and more for a user to update their expected vocabulary.
[0099] Looking up words in dictionaries, through translation services, personal assistant apps, search pages, or other technologies can strongly indicate that the word is unknown. In other words, technology can assume that recently searched words are unknown and provide annotations when the word appears in future text. The classification can change to known when a word is used by a user or when the user is exposed to it multiple times. In one aspect, words recently searched by a user are automatically added to a phrasebook. In another scenario, the system can use recently searched words as annotations for synonyms of unknown words. This helps the user learn recently searched words by associating them with synonyms the user already knows.
[0100] Other sources of information about a user's language knowledge can include spelling and grammar errors made when composing documents, emails, etc. In one aspect, technology collects instances where spelling or grammar errors are automatically corrected or manually corrected by the user.
[0101] At step 920, content including text being accessed by a specific user is received. This content can be a web page, social media post, book, audio recording including speech, video recording including speech, presentation, document, user interface, etc. In the case of audio or video recordings, a natural language processing component can be used to perform speech-to-text conversion. In one aspect, the content is intercepted by a browser plugin for analysis when the user navigates to a web page. The plugin can perform analysis or pass the content to another component installed on the client device or server for analysis. Document applications, presentation applications, e-readers, or other content applications can similarly retrieve content for analysis.
[0102] At step 930, a subset of words in the text is identified as potentially unknown to the user. The techniques described herein can use broad filters to identify subsets of potentially unknown words. For example, words with a frequency higher than a threshold for use within the language can be excluded as potentially unknown words. In one aspect, the filter is user-specific. For example, the filter can be generally applied based on a reading level associated with the user. The reading level can be calculated based on an analysis of the user's reading content, reading speed, and vocabulary usage when composing communications.
[0103] At step 940, data from the user-specific VAS and a subset of potentially unknown words are input into a machine classifier, which generates a confidence score for each word in the subset. The confidence score indicates that the word is potentially unknown to the specific user. The machine classifier may receive multiple signals related to the user and the public as input to determine that a word is unknown. The machine classifier may utilize one or more of the following: neural networks, support vector machines, generalized linear modeling, data mining, text mining, natural language processing algorithms, regression, or decision trees. These may be used in isolation or in combination as part of a machine classifier in some configurations to compute confidence scores or otherwise classify words on a known / unknown spectrum.
[0104] In one aspect, the signal includes reading data. Reading data can include text from content read by a user, categorization of content read by the user, reading analytics, and other data related to the user's reading habits. Content categorization can be based on reading level, topic, source type, or other factors. Reading level can be determined ergonomically, for example, by a knowledge base describing reading levels for various books. Reading level can also be determined analytically by analyzing sentence length, sentence complexity, and the complexity of the vocabulary used. Topic can be determined by analyzing a knowledge base of known works (such as books). Topic can also be determined by evaluating metadata associated with various types of content, such as web pages. Topic can also be determined directly by considering the content of the text and the categorization method of the source. Topic can be used to define a user's interests. For example, a user can be categorized as having one or more interests, such as sports, religious beliefs, politics, music, etc. Different levels of granularity are possible. Instead of sports, a user can be categorized as having an interest in baseball, football, rugby, basketball, track and field, or hockey.
[0105] On the other hand, signal data includes creation or writing data. Creation or writing data can include email text, social media posts, documents, presentations, and other communications and documents generated by the user.
[0106] On the other hand, signal data includes searched words. Searched words are those words that a user searches for in order to obtain additional information about or a translation of. For example, a user may have already looked up the word in an online dictionary or translation service. A user may have already searched for the word through a search engine, online dictionary, or personal assistant application.
[0107] Signals can also include user feedback. User feedback can be explicit or implicit. Implicit feedback can include user interaction with notes previously presented to the user. For example, when a word is identified as unknown and associated with a note that has been ignored by the user (e.g., unused, unconsumed, no interaction), no interaction can be implicit feedback suggesting that the user is familiar with the word. Similarly, access to an note for interaction with an unknown word suggests that the word has been correctly identified as unknown. Explicit feedback can be provided through mechanisms available in the note. For example, an optional "remove" command can be provided within the note, which, when selected, removes the associated unknown word from the phrasebook and classifies it as known. In this respect, words explicitly identified as known can be used as signals.
[0108] Using all the above signals, the frequency and relevance of exposure or use of various words can be determined and used as input to a machine classifier. Other signals can be obtained. These signals are calculated by performing operations on the original signal data.
[0109] The machine classifier generates a confidence score for each word being evaluated. Words with confidence scores higher than a certain threshold can be designated as unknown. The machine classifier can use any type of technique or combination of techniques capable of making classifications. For example, neural networks, support vector machines, data mining, or decision trees can be used in isolation or in combination as part of a machine classifier in some configurations. Unknown words from the content can be added to a phrasebook that includes unknown words. Words identified as unknown can be input for subsequent identification.
[0110] At step 950, an unknown word is selected from the subset because it has a higher confidence score than the confidence scores calculated for the other words in the subset. In one aspect, the word with the highest confidence score is selected as long as its confidence score is above a threshold that distinguishes the boundary between known and unknown words. In another aspect, all words above the threshold are selected. In yet another aspect, a finite number of words above the threshold are selected.
[0111] At step 960, the appearance of the unknown words within the content is changed to provide an indication that the unknown words are selectable. Various ways to change the appearance of unknown words include changing their color, underlining them, double-lined them, and so on. In one aspect, the appearance is changed by adding links to the unknown words, and the appearance changes automatically based on display attributes within the application displaying the content. For example, a document application might change the color of each linked word from black to blue.
[0112] At step 970, an instruction from the user to interact with the unknown word is received. The user can interact with the word on a touchscreen, using voice commands or gestures, using a mouse or touchpad, or through some other method. Word selection is one example of interaction, and hovering over a word is another.
[0113] At step 980, annotations providing additional information about the unknown word are output for display. Annotations may include: explanations of people, places, or things; definitions of the word; usage of the word in a sentence; translations of the word in different languages; synonyms; or some other form of comprehension aid. Annotations can be built from external sources such as online dictionaries, web pages, and knowledge bases. Annotations may include links to online sources used to build the annotations or other online sources that may be relevant to providing additional information about the unknown word or phrase.
[0114] In one variation, the technique described herein uses other words with the same or similar meanings to replace words within the content. In one aspect, unknown words are replaced using known words. Known words can be explicitly designated as selectable. Upon selection, a note can be displayed showing the original unknown word. The reverse approach is also possible. Known words that are synonyms of unknown words can be used to replace the unknown words. This provides an opportunity to expose the user to unknown words. In this case, a general list of unknown words for a specific user, such as those found in a phrasebook associated with the user, can be used to generate a list of synonyms for those words. Synonyms can be compared with words within the content.
[0115] In another variation, a user can request unknown synonyms for a known word. In one aspect, words with available unknown synonyms can be made selectable. In another aspect, a user can identify a word and request unknown synonyms, even if the word is not selectable or is otherwise visually distinguishable.
[0116] In one aspect, contextual synonyms can be selected. Contextual synonyms share context with unknown words. For example, sports could be the context of a baseball website, and politics could be the context of a political website or a politically related social media post. The context of an unknown word can be determined by words in the content, the source of the content, metadata associated with the content, and other information. Content can be determined at various levels of granularity. For example, the context could be sports, or it could be baseball, rugby, basketball, etc.
[0117] The context of various synonyms can be obtained from word definitions or usage examples provided in a knowledge base. Topics can also be determined by analyzing the frequency of occurrence in content with the same context. For example, synonyms for unknown words appearing in sports articles can be selected from many available synonyms based on their frequency of occurrence within a corpus of sports articles compared to the frequency of other synonyms used in sports contexts. In addition to synonyms, various aspects of the technology can utilize contextual synonyms or other understanding aids to replace acronyms or slang.
[0118] When selecting contextual synonyms, the overall frequency of use can also be considered. Therefore, synonyms that appear infrequently overall but are used at a higher rate in sports articles can be selected in a sports context, especially if the viewer appears familiar with the article. In one aspect, only synonyms that are likely known to the user are evaluated for their closeness to the specific context.
[0119] Annotations can be provided in the same language as the content or in a different language. For example, the content could be in Spanish and the annotations in English. The language of the annotations can be selected by the user through a preference interface. Alternatively, the system can learn an individual's language understanding, derive preferences from observing user events, and select the annotation language accordingly. For example, a user who typically looks up English words in an English language dictionary could receive English language annotations while reading English language content. Conversely, a user who typically seeks to translate English words into Spanish could receive Spanish language annotations while reading English content. In one aspect, translations of unknown words from the first language to the user's native language (or any language the user is more fluent in than the content language) can be provided when all available synonyms in the user's first language may also be unknown to the user. The user's known languages can be explicitly provided by the user or learned by observing the language of content consumed or created by the user.
[0120] Alternatively, footnotes can provide content in multiple languages. For example, footnotes can provide contextual synonyms for unknown English words in English and French translations of unknown English words. Footnotes can also provide content in languages the reader is very familiar with and in a second language the reader is less fluent in.
[0121] Exemplary operating environment
[0122] Refer generally to the accompanying drawings and first, specifically refer to... Figure 10An exemplary operating environment for implementing various aspects of the techniques described herein is shown and generally designated as computing device 1000. Computing device 1000 is merely one example of a suitable computing environment and is not intended to imply any limitation on the scope of use of the techniques described herein. Computing device 1000 should also not be construed as having any dependency or requirement relating to any of the components or combinations of components illustrated.
[0123] The techniques described herein can be described within the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program components, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Generally, a program component, including routines, programs, objects, components, data structures, and the like, refers to code that performs a specific task or implements a specific abstract data type. The techniques described herein can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, special-purpose computing devices, and so on. Aspects of the techniques described herein can also be practiced in distributed computing environments, where tasks are performed by remote processing devices linked via a communication network.
[0124] Continue to refer to Figure 10 The computing device 1000 includes a bus 1010 directly or indirectly coupled to: a memory 1012, one or more processors 1014, one or more presentation components 1016, an input / output (I / O) port 1018, an I / O component 1020, and an illustrative power supply 1022. The bus 1010 represents something that can be one or more buses (e.g., an address bus, a data bus, or a combination thereof). Although lines are used for clarity... Figure 10 The various boxes are depicted, but in reality, the individual components are not so clearly defined, and metaphorically, the lines will be more accurately described as gray and blurred. For example, presentation components such as display devices can be considered as I / O components. Furthermore, processors have memory. The inventors of this paper recognize this as a matter of domain and reiterate... Figure 10 The illustrations are merely illustrative of exemplary computing devices that can be used in conjunction with one or more aspects of the techniques described herein. No distinction is made between categories such as “workstation,” “server,” “laptop,” “handheld device,” etc., as all are contemplated within… Figure 10 Within the scope of this term, it refers to "computer" or "computing device." The computing device 1000 can be a PC, tablet computer, smartphone, virtual reality headset, augmented reality headset, game console, and the like.
[0125] Computing device 1000 typically includes various computer-readable media. Computer-readable media can be any available medium accessible by computing device 1000 and includes both volatile and non-volatile media, removable media, and non-removable media. By way of example, and not limitation, computer-readable media can include computer storage media and communication media. Computer storage media includes both volatile and non-volatile media, removable media, and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data.
[0126] Computer storage media include RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital multi-disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices. Computer storage media do not include transmitted data signals.
[0127] Communication media typically embody computer-readable media, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission media, and include any information delivery medium. The term "modulated data signal" means a signal having one or more of its characteristics set or altered in a manner that encodes information in the signal. By way of example, and not limitation, communication media includes wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0128] Memory 1012 includes computer storage media in the form of volatile memory and / or non-volatile memory. Memory 1012 can be removable, non-removable, or a combination thereof. Exemplary memory includes solid-state memory, hard disk drive, optical disk drive, etc. Computing device 1000 includes one or more processors 1014 that read data from various entities such as bus 1010, memory 1012, or I / O component 1020. Multiple presentation components 1016 present data instructions to a user or other device. Exemplary presentation components 1016 include display devices, speakers, printing components, oscillation components, etc. I / O ports 1018 allow computing device 1000 to be logically coupled to other devices, including I / O components 1020, some of which may be built-in.
[0129] Illustrative I / O components include microphones, joysticks, gaming pads, satellite antennas, scanners, printers, display devices, wireless devices, controllers (such as styluses, keyboards, and mice), natural user interfaces (NUIs), and so on. In various aspects, a pen digitizer (not shown) and accompanying input devices (also not shown, but which may include a pen or stylus by way of example only) are provided to digitally capture freehand user input. The connection between the pen digitizer and the processor(s) 1014 can be direct or coupled via a serial port, parallel port, and / or other interfaces and / or system buses known in the art. Additionally, the digitizer input component can be a component separate from output components such as those of a display device, or in some aspects, the available input area of the digitizer can coexist with the display area of the display device, be integrated with the display device, or exist as a separate device superimposed or otherwise attached to the display device. Any and all such variations and any combinations thereof are contemplated to be within the scope of the various aspects of the techniques described herein.
[0130] NUI processes user-generated air gestures, voice, or other physiological input. Appropriate NUI input can be interpreted as inking strokes presented by the computing device 1000. These requests can be sent to appropriate network elements for further processing. NUI implements any combination of voice recognition, touch and stylus recognition, face recognition, biometric recognition, gesture recognition on and near the screen, air gestures, head and eye tracking, and touch recognition associated with the display on the computing device 1000. The computing device 1000 can be equipped with a depth camera, such as a stereo camera system, an infrared camera system, an RGB camera system, and combinations thereof for gesture detection and recognition. Additionally, the computing device 1000 can be equipped with an accelerometer or gyroscope enabling motion detection. The output of the accelerometer or gyroscope can be provided to the display of the computing device 1000 to render immersive augmented reality or virtual reality.
[0131] Computing device 1000 may include radio device 1024. The radio device transmits and receives radio communications. Computing device 1000 may be a wireless terminal suitable for receiving communications and media via various wireless networks. Computing device 1000 may communicate with other devices via wireless protocols such as Code Division Multiple Access (“CDMA”), Global System for Mobile Communications (“GSM”), or Time Division Multiple Access (“TDMA”). Radio communications may be short-range connections, long-range connections, or a combination of both. When we refer to “short” and “long” type connections, we do not intend to refer to a spatial relationship between two devices. Rather, we generally use “short” and “long” to refer to different kinds or types of connections (i.e., primary and secondary connections). Short-range connections may include access to devices (e.g., mobile hotspots) that provide access to wireless communication networks. Connections, such as WLAN connections using the 802.11 protocol, are examples of short-range connections. Bluetooth connections to another computing device are a second example of short-range connections. Long-range connections can include connections using one or more of the following protocols: CDMA, GPRS, GSM, TDMA, and 802.16.
[0132] The various aspects of the technology have been described in an illustrative rather than restrictive manner. It will be understood that certain features and sub-combinations are useful and can be employed without reference to other features and sub-combinations. This is foreseeable by and within the scope of the claims.
Claims
1. A method for generating language understanding aids, comprising: A user-specific phrasebook of first-language updated words is used for a specific user. The user-specific phrasebook includes words unknown to the specific user. The words unknown to the specific user are predicted to be unknown to the specific user based at least in part on observed user interactions with known words, and the observed user interactions indicate the reading level associated with the specific user. Receive content including text that is being accessed by the specific user; The text is determined to contain words unknown to the specific user by comparing the words in the text with words in the user-specific phrasebook. Changing the appearance of the word within the content that is unknown to the specific user to provide an optional indication of the word that is unknown to the specific user; Receive instructions from the user to interact with the word that is unknown to the specific user; and The output includes annotations that provide additional information about the word that is unknown to the specific user for display purposes.
2. The method of claim 1, wherein the additional information is a translation of the word unknown to the particular user in a second language, wherein the word unknown to the particular user is further determined based on the public availability of the word unknown to the particular user among a particular audience.
3. The method according to claim 1, wherein the method further comprises: Determine that the specific user has searched for information about the second unknown word or has written or spoken the second unknown word at least twice; as well as Add the second unknown word to the user-specific phrasebook as a known word.
4. The method of claim 1, wherein the user-specific phrasebook is populated by a machine classifier that classifies words as known or unknown to the specific user based on reading data for that specific user, the reading data including the speed at which the specific user reads the classified words.
5. The method of claim 4, wherein the method includes updating a user-specific vocabulary analysis store, the user-specific vocabulary analysis store including the frequency of words appearing in text written by the specific user and text read by the specific user, wherein information from the user-specific vocabulary analysis store is used as input to the machine classifier for classifying words.
6. The method of claim 5, wherein the method further comprises updating the user-specific vocabulary analysis storage to include the interaction between the specific user and the annotation.
7. The method of claim 5, wherein the method further comprises updating the user-specific vocabulary analysis storage to include user interests, which can be used by the machine classifier to calculate a confidence score for a given word that is unknown.
8. The method according to claim 1, wherein the content is a web page.
9. A method for generating language understanding aids, comprising: Receive text content in a specific language that is not the user's first language; The text content is scanned to identify words with a confidence level higher than a threshold unknown to the user, determined by a machine classifier that uses observed user interactions with known words in text previously read or written by the user as input, wherein the observed user interactions indicate a reading level associated with the user, and the known words are in the specific language. as well as The text content is output for display to the user, wherein the words have a different visual appearance from other words in the text content to indicate that the words are selectable.
10. The method of claim 9, further comprising: Receive user interaction with the word; as well as The output of neighboring words helps the user understand the annotations of the words.
11. The method of claim 10, wherein the annotation is a translation of the word in a second language different from the particular language, wherein the user profile for the user indicates that the user is familiar with the second language.
12. The method of claim 10, wherein the annotation includes an explanation of the word from a knowledge base, wherein the word is one of a person, place, or thing.
13. The method of claim 10, wherein the annotation includes a definition of the word or an example of the word being used in a sentence.
14. The method of claim 10, wherein the annotation includes a contextual synonym of the word.
15. The method of claim 9, wherein the machine classifier further uses a record of words looked up by the user in an online dictionary as input.
16. One or more computer storage media having computer-executable instructions thereon, the computer-executable instructions, when executed by a computing device, performing a method for generating language understanding aids, the method comprising: Update user-specific vocabulary analysis storage by collecting user reading and writing data for specific users; Receive content including text accessed by the specific user; The user-specific vocabulary analysis storage is used to determine a subset of words in the text that are potentially unknown to the user, wherein the determination of the subset of words is based at least in part on user interactions with the subset of words that have not been previously observed, the user interactions indicating a reading level associated with the user, and the subset of potentially unknown words in the text is in a language other than the user's first language; Data from the user-specific vocabulary analysis storage and a subset of potentially unknown words are input into a machine classifier, which generates a confidence score for each word in the subset of words, the confidence score indicating whether the word is unknown to the specific user; Select an unknown word from a subset of the words, the unknown word having a confidence score higher than other confidence scores calculated for other words in the subset of the words; Change the appearance of the unknown word within the content to provide an indication that the unknown word is selectable; Receive instructions from the user regarding interaction with the unknown word; and Output annotations that provide additional information about the unknown word for display purposes.
17. The medium of claim 16 further includes collecting user interaction data, the user interaction data describing user interactions with the annotation, and recording the user interaction data into the user-specific vocabulary analysis storage.
18. The medium according to claim 16, further comprising: Replace the unknown word with a synonym that was not classified as unknown by the machine classifier.
19. The medium of claim 16, wherein the method further comprises inputting general language statistics into the machine classifier, the general language statistics including the frequency of use of different words within a group.
20. The medium of claim 16, wherein the method further comprises entering a geographic region associated with the particular user into the machine classifier.